{"id":10175,"date":"2026-10-06T02:15:11","date_gmt":"2026-10-05T17:15:11","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/06\/llamacpp-v0-6-0-released\/"},"modified":"2026-10-06T02:15:11","modified_gmt":"2026-10-05T17:15:11","slug":"llamacpp-v0-6-0-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/06\/llamacpp-v0-6-0-released\/","title":{"rendered":"llama.cpp v0.6.0 Released with llama_batch_ext and Metal Boosts"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">ggml-org\/llama.cpp<\/a><\/td>\n<\/tr>\n<tr>\n<td>Version<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\/tag\/v0.6.0\">v0.6.0<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-06<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>MIT<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>The latest version v0.6.0 of llama.cpp, the C++ LLM inference engine, has been released. llama.cpp is an open-source project developed under the MIT license, aimed at running large language models efficiently on consumer-grade hardware.<\/p>\n<p>The most impactful change in this update is the introduction of the new batch processing API, <code>llama_batch_ext<\/code>. This enables the mixed processing of tokens and embeddings within a single batch, strengthening support for advanced inference techniques such as MTP (Multi-Token Prediction). For users developing custom tools using the API or those wanting to leverage the latest speculative decoding features, this is an important release to consider for migration.<\/p>\n<h2>Breaking Changes and Deprecations<\/h2>\n<p>Internal session formats and API specifications have been updated, meaning previous configurations and some code will no longer work as-is.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Old Element<\/th>\n<th style=\"text-align: left;\">New Element<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\"><code>llama_batch<\/code><\/td>\n<td style=\"text-align: left;\"><code>llama_batch_ext<\/code> (recommended)<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><code>LLAMA_SESSION_VERSION<\/code> 10 or lower<\/td>\n<td style=\"text-align: left;\"><code>LLAMA_SESSION_VERSION<\/code> 11<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><code>LLAMA_STATE_SEQ_VERSION<\/code> 3 or lower<\/td>\n<td style=\"text-align: left;\"><code>LLAMA_STATE_SEQ_VERSION<\/code> 4<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Because session and sequence state versions have been updated, session files and state data saved with previous versions may no longer be readable in this release. Additionally, developers are encouraged to migrate from the legacy <code>llama_batch<\/code> API to the more flexible <code>llama_batch_ext<\/code>.<\/p>\n<h2>Key Changes<\/h2>\n<h3>Introduction of the Extended Batch API <code>llama_batch_ext<\/code><\/h3>\n<p>An improved API, <code>llama_batch_ext<\/code>, has been implemented to increase inference flexibility. This API allows tokens and embeddings to be included in the same batch, as well as attaching &#8220;state&#8221; embeddings per token. This makes it easier to support specialized model structures like MTP and Deepstack. Developers embedding llama.cpp as a library will need to rewrite code for the new API, but gain the ability to build more complex inference pipelines.<\/p>\n<h3>Major Speedups for the Metal Backend<\/h3>\n<p>Inference performance in Apple Silicon environments (Mac) has significantly improved. A new Flash Attention kernel for F16 KV caches has been added, alongside new matrix multiplication (mat-mul) kernels optimized for speculative decoding and batch processing. According to reports, matrix operations on Apple GPUs are accelerated by up to approximately 3x. Users running local LLMs on a Mac can expect to experience a noticeable speed boost after updating.<\/p>\n<h3>Web UI Overhaul and Hugging Face Hub Integration<\/h3>\n<p>The UI provided by the built-in web server has been greatly enhanced. A Hugging Face Hub data layer has been integrated, implementing a pipeline to search for and download models directly from the browser. A feature to estimate whether a selected model fits into the current PC&#8217;s memory has also been added. Users who previously managed models via the command line can now experiment with models much more intuitively.<\/p>\n<h3>Support for Qwen4Exp and MTP Speculative Decoding<\/h3>\n<p>Speculative decoding using MTP (Multi-Token Prediction) is now available for Qwen4Exp models. Tests in a DGX Spark environment reportedly show a decoding speed increase of about 1.5x. Alongside this, optimizations to halve the memory used for indexer score calculation and improvements to mask generation efficiency have been implemented. Using these features requires re-downloading corresponding models and checking configuration settings.<\/p>\n<h3>Addition of Decision Model API <code>\/v1\/systemone<\/code><\/h3>\n<p>A new endpoint, <code>\/v1\/systemone<\/code>, has been added to the server features to handle Decision Models. Models such as Laya, Julia-1, Lev, OpenJev, and Kev are supported; these extend traditional embedding models tailored for specific decision-making tasks. OpenJev also supports image inputs, allowing multimodal decision-making tasks to be executed via the API.<\/p>\n<h3>Update to Core Library ggml v0.26.0<\/h3>\n<p>The underlying numerical computation library ggml has been updated to v0.26.0. This update brings many low-level improvements, including added support for BF16 operations on the CPU backend, implementations of sparse Flash Attention in Vulkan and Metal, and build support for Windows ARM64 environments. Furthermore, GGUF format size validation has been made stricter, making the detection of corrupted model files more reliable.<\/p>\n<h2>Supported Models and Hardware<\/h2>\n<h3>Newly Supported Models<\/h3>\n<p>This update brings support for numerous new model architectures, allowing users to convert these state-of-the-art models into the GGUF format and run local inference.<\/p>\n<ul>\n<li><strong>GLM-5.3-Flash (GLM5-Next)<\/strong>: A newly supported 320B parameter multimodal MoE (Mixture of Experts) model with KDA\/DSA hybrid architecture, supporting both text and vision (images). It supports distinctive structures such as mHC and MoE.<\/li>\n<li><strong>Clef<\/strong>: A new decision model with full support for both text and vision is supported. Server-side support has also been added to handle Clef image inputs.<\/li>\n<li><strong>Ling 3.0 VL<\/strong>: Newly supported via integration into the BailingMoeV3 architecture.<\/li>\n<li><strong>Nimble<\/strong>: Added as a new decision model, usable via the <code>\/v1\/systemone<\/code> server API.<\/li>\n<li><strong>LFM2.5-Encoder-230M \/ LFM2.5-Encoder-350M<\/strong>: Newly registered as <code>Lfm2BidirectionalForMaskedLM<\/code>, enabling use as encoder models.<\/li>\n<li><strong>Classifier Pooling for Re-rankers (<code>classifier_pooling<\/code>)<\/strong>: Support for classifier pooling has been added in re-ranker models based on Causal LLMs.<\/li>\n<\/ul>\n<h3>Enhanced Hardware and Accelerator Support<\/h3>\n<p>Alongside the upgrade of underlying ggml to v0.26.0, an extensive number of optimizations and new compute kernels have been added across various hardware backends.<\/p>\n<ul>\n<li><strong>CPU Backend<\/strong>: BF16 (Bfloat16) operations have been added, and tiled k-quant matrix multiplication (<code>mul_mat<\/code>) is now supported.<\/li>\n<li><strong>CUDA Backend<\/strong>: A model-driven W4A4 (NVFP4\/MXFP4) matrix multiplication path has been added, and accumulation operations in MMQ (Min-Max Quantization) were optimized for NVFP4 types. Shared-expert fusion into MMVQ has also been implemented. Meanwhile, two bugs occurring in Volta-generation Flash Attention were fixed.<\/li>\n<li><strong>Vulkan Backend<\/strong>: Sparse Flash Attention kernels for quantized K\/V (key-value) caches were added. Out-of-bounds access bugs in Flash Attention shared memory writes were also fixed.<\/li>\n<li><strong>Metal Backend<\/strong>: Flash Attention kernels utilizing the new tensor API were implemented for F16 K\/V caches. Additionally, few-row MMA matrix multiplication kernels to accelerate speculative decoding and batch decoding were added, yielding up to roughly 3x speedups on Apple GPUs.<\/li>\n<li><strong>SYCL Backend<\/strong>: Sparse Flash Attention kernels were added, alongside support for Q8_0 DMMV ESIMD and wide loads in MMVQ. Larger register files were added for D=512 Flash Attention vector kernels, avoiding slow oneDNN reference matrix multiplications and Flash Attention fallbacks.<\/li>\n<li><strong>WebGPU Backend<\/strong>: MMVQ support was added for Q1_0, Q5_0, Q5_1, Q3_K, Q5_K, Q6_K, and MXFP4. F16 support in <code>fill<\/code> and <code>set_rows<\/code>, as well as BFloat16 support in <code>MUL_MAT<\/code>, <code>MUL_MAT_ID<\/code>, and <code>GET_ROWS<\/code> were added.<\/li>\n<li><strong>Hexagon Backend<\/strong>: A sampler was added, and <code>q2_k<\/code> and <code>q3_k<\/code> quantization types are now supported. ALLREDUCE optimizations for safe scatter mode and improvements to DMA copy\/concat processing were made.<\/li>\n<li><strong>Windows ARM64 MSVC Build<\/strong>: Builds using MSVC (Microsoft Visual C++) are now officially enabled in Windows ARM64 environments.<\/li>\n<\/ul>\n<h2>How to Update<\/h2>\n<p>Specific installation commands for adopting version v0.6.0 are not explicitly stated in the provided materials.<\/p>\n<p>Therefore, please refer to the instructions on the official GitHub releases page and README regarding build procedures from the latest source code or downloading precompiled binaries tailored to your environment.<\/p>\n<p>Generally, updates can be performed by cloning the repository and executing build commands appropriate for your system, or by acquiring the necessary assets from the releases section.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/llama-cpp-v0-5-0-released\/\">llama.cpp v0.5.0 Released with Backend and Server Upgrades<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/06\/ggml-v0260-released\/\">ggml v0.26.0 Released with Sparse Flash Attention &amp; New Backends<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/llama-cpp-v041-released\/\">llama.cpp v0.4.1 Released with Breaking Changes<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/ggml-v0-24-0-released-2\/\">ggml v0.24.0 Released with Backend Improvements and API Updates<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Follow this tool<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp overview and release history (5 releases tracked)<\/a><\/li>\n<li><strong>Other inference engines and runtimes<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-sglang-en\/\">SGLang<\/a><\/li>\n<li><strong>Engines mentioned in this article<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ggml-en\/\">ggml<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\/tag\/v0.6.0\">ggml-org\/llama.cpp v0.6.0<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/24669\">PR #24669<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/27773\">PR #27773<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29831\">PR #29831<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29969\">PR #29969<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29151\">PR #29151<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29844\">PR #29844<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29862\">PR #29862<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29627\">PR #29627<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29869\">PR #29869<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29570\">PR #29570<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29639\">PR #29639<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29358\">PR #29358<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29483\">PR #29483<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29897\">PR #29897<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29062\">PR #29062<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/28985\">PR #28985<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29803\">PR #29803<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29753\">PR #29753<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>llama.cpp v0.6.0 is out, featuring the new llama_batch_ext API, major Metal backend speedups, web UI updates, and new model support.<\/p>\n","protected":false},"author":1,"featured_media":10174,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[105],"tags":[505,1187,161,163,1073,1547],"class_list":["post-10175","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engines-and-tools","tag-apple-silicon-en","tag-cuda-en","tag-ggml-en","tag-gguf-en","tag-llama-cpp-en","tag-verified"],"lang":"en","translations":{"en":10175,"ja":10173},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10175","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=10175"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10175\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/10174"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=10175"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=10175"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=10175"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}