{"id":1238,"date":"2026-09-17T16:56:51","date_gmt":"2026-09-17T07:56:51","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/engine-llama-cpp-en\/"},"modified":"2026-09-18T11:58:11","modified_gmt":"2026-09-18T02:58:11","slug":"engine-llama-cpp-en","status":"publish","type":"page","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/","title":{"rendered":"llama.cpp: Releases and Overview"},"content":{"rendered":"<p>LLM inference engine written in C\/C++. Runs GGUF models on CPU and GPU (CUDA \/ Metal \/ Vulkan \/ ROCm) and underpins much of the local-LLM ecosystem, including Ollama, LM Studio and KoboldCpp.<\/p>\n<p><em>Category: Inference engines and runtimes. Part of our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engines-en\/\">engine and tool tracker<\/a>.<\/em><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">ggml-org\/llama.cpp<\/a><\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>MIT<\/td>\n<\/tr>\n<tr>\n<td>Main language<\/td>\n<td>C++<\/td>\n<\/tr>\n<tr>\n<td>GitHub stars<\/td>\n<td>128,500 (as of 2026-09-17)<\/td>\n<\/tr>\n<tr>\n<td>Latest release<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\/tag\/v0.4.1\">v0.4.1<\/a> (2026-09-15)<\/td>\n<\/tr>\n<tr>\n<td>Install \/ docs<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/blob\/master\/docs\/build.md\">official documentation<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The project describes itself as: \u201cLLM inference in C\/C++\u201d<\/p>\n<h2>Release History<\/h2>\n<p><em>Compiled by Local Model Watch from the project&#8217;s GitHub releases. Pre-releases are omitted. Where we wrote an article about a release, it is linked in the last column; smaller releases are tracked here only.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Version<\/th>\n<th>Released<\/th>\n<th>Release notes<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>v0.4.1<\/td>\n<td>2026-09-15<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\/tag\/v0.4.1\">GitHub<\/a><\/td>\n<td>\u2014<\/td>\n<\/tr>\n<tr>\n<td>v0.4.0<\/td>\n<td>2026-09-05<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\/tag\/v0.4.0\">GitHub<\/a><\/td>\n<td>\u2014<\/td>\n<\/tr>\n<tr>\n<td>v0.3.0<\/td>\n<td>2026-08-25<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\/tag\/v0.3.0\">GitHub<\/a><\/td>\n<td>\u2014<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Articles on Local Model Watch<\/h2>\n<p><em>No articles yet.<\/em><\/p>\n<p><em>Last updated 2026-09-18 (JST). Facts above come from the GitHub API; the one-paragraph summary is written by the site&#8217;s editors.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>LLM inference engine written in C\/C++. Runs GGUF models on CPU and GPU (CUDA \/ Metal \/ Vulkan \/ ROCm) and underpins much of the local-LLM ecosystem, including Ollama, LM Studio and KoboldCpp. Category: Inference engines and runtimes. Part of our engine and tool tracker. At a Glance Item Value Repository ggml-org\/llama.cpp License MIT Main [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-1238","page","type-page","status-publish","hentry"],"lang":"en","translations":{"en":1238,"ja":1236},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/1238","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=1238"}],"version-history":[{"count":1,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/1238\/revisions"}],"predecessor-version":[{"id":1560,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/1238\/revisions\/1560"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=1238"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}