llama.cpp v0.6.0 Released with llama_batch_ext and Metal Boosts

llama.cpp v0.6.0 Released with llama_batch_ext and Metal Boosts

At a Glance

Item Value
Repository ggml-org/llama.cpp
Version v0.6.0
Published 2026-10-06
License MIT
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

The latest version v0.6.0 of llama.cpp, the C++ LLM inference engine, has been released. llama.cpp is an open-source project developed under the MIT license, aimed at running large language models efficiently on consumer-grade hardware.

The most impactful change in this update is the introduction of the new batch processing API, llama_batch_ext. This enables the mixed processing of tokens and embeddings within a single batch, strengthening support for advanced inference techniques such as MTP (Multi-Token Prediction). For users developing custom tools using the API or those wanting to leverage the latest speculative decoding features, this is an important release to consider for migration.

Breaking Changes and Deprecations

Internal session formats and API specifications have been updated, meaning previous configurations and some code will no longer work as-is.

Old Element New Element
llama_batch llama_batch_ext (recommended)
LLAMA_SESSION_VERSION 10 or lower LLAMA_SESSION_VERSION 11
LLAMA_STATE_SEQ_VERSION 3 or lower LLAMA_STATE_SEQ_VERSION 4

Because session and sequence state versions have been updated, session files and state data saved with previous versions may no longer be readable in this release. Additionally, developers are encouraged to migrate from the legacy llama_batch API to the more flexible llama_batch_ext.

Key Changes

Introduction of the Extended Batch API llama_batch_ext

An improved API, llama_batch_ext, has been implemented to increase inference flexibility. This API allows tokens and embeddings to be included in the same batch, as well as attaching “state" embeddings per token. This makes it easier to support specialized model structures like MTP and Deepstack. Developers embedding llama.cpp as a library will need to rewrite code for the new API, but gain the ability to build more complex inference pipelines.

Major Speedups for the Metal Backend

Inference performance in Apple Silicon environments (Mac) has significantly improved. A new Flash Attention kernel for F16 KV caches has been added, alongside new matrix multiplication (mat-mul) kernels optimized for speculative decoding and batch processing. According to reports, matrix operations on Apple GPUs are accelerated by up to approximately 3x. Users running local LLMs on a Mac can expect to experience a noticeable speed boost after updating.

Web UI Overhaul and Hugging Face Hub Integration

The UI provided by the built-in web server has been greatly enhanced. A Hugging Face Hub data layer has been integrated, implementing a pipeline to search for and download models directly from the browser. A feature to estimate whether a selected model fits into the current PC’s memory has also been added. Users who previously managed models via the command line can now experiment with models much more intuitively.

Support for Qwen4Exp and MTP Speculative Decoding

Speculative decoding using MTP (Multi-Token Prediction) is now available for Qwen4Exp models. Tests in a DGX Spark environment reportedly show a decoding speed increase of about 1.5x. Alongside this, optimizations to halve the memory used for indexer score calculation and improvements to mask generation efficiency have been implemented. Using these features requires re-downloading corresponding models and checking configuration settings.

Addition of Decision Model API /v1/systemone

A new endpoint, /v1/systemone, has been added to the server features to handle Decision Models. Models such as Laya, Julia-1, Lev, OpenJev, and Kev are supported; these extend traditional embedding models tailored for specific decision-making tasks. OpenJev also supports image inputs, allowing multimodal decision-making tasks to be executed via the API.

Update to Core Library ggml v0.26.0

The underlying numerical computation library ggml has been updated to v0.26.0. This update brings many low-level improvements, including added support for BF16 operations on the CPU backend, implementations of sparse Flash Attention in Vulkan and Metal, and build support for Windows ARM64 environments. Furthermore, GGUF format size validation has been made stricter, making the detection of corrupted model files more reliable.

Supported Models and Hardware

Newly Supported Models

This update brings support for numerous new model architectures, allowing users to convert these state-of-the-art models into the GGUF format and run local inference.

  • GLM-5.3-Flash (GLM5-Next): A newly supported 320B parameter multimodal MoE (Mixture of Experts) model with KDA/DSA hybrid architecture, supporting both text and vision (images). It supports distinctive structures such as mHC and MoE.
  • Clef: A new decision model with full support for both text and vision is supported. Server-side support has also been added to handle Clef image inputs.
  • Ling 3.0 VL: Newly supported via integration into the BailingMoeV3 architecture.
  • Nimble: Added as a new decision model, usable via the /v1/systemone server API.
  • LFM2.5-Encoder-230M / LFM2.5-Encoder-350M: Newly registered as Lfm2BidirectionalForMaskedLM, enabling use as encoder models.
  • Classifier Pooling for Re-rankers (classifier_pooling): Support for classifier pooling has been added in re-ranker models based on Causal LLMs.

Enhanced Hardware and Accelerator Support

Alongside the upgrade of underlying ggml to v0.26.0, an extensive number of optimizations and new compute kernels have been added across various hardware backends.

  • CPU Backend: BF16 (Bfloat16) operations have been added, and tiled k-quant matrix multiplication (mul_mat) is now supported.
  • CUDA Backend: A model-driven W4A4 (NVFP4/MXFP4) matrix multiplication path has been added, and accumulation operations in MMQ (Min-Max Quantization) were optimized for NVFP4 types. Shared-expert fusion into MMVQ has also been implemented. Meanwhile, two bugs occurring in Volta-generation Flash Attention were fixed.
  • Vulkan Backend: Sparse Flash Attention kernels for quantized K/V (key-value) caches were added. Out-of-bounds access bugs in Flash Attention shared memory writes were also fixed.
  • Metal Backend: Flash Attention kernels utilizing the new tensor API were implemented for F16 K/V caches. Additionally, few-row MMA matrix multiplication kernels to accelerate speculative decoding and batch decoding were added, yielding up to roughly 3x speedups on Apple GPUs.
  • SYCL Backend: Sparse Flash Attention kernels were added, alongside support for Q8_0 DMMV ESIMD and wide loads in MMVQ. Larger register files were added for D=512 Flash Attention vector kernels, avoiding slow oneDNN reference matrix multiplications and Flash Attention fallbacks.
  • WebGPU Backend: MMVQ support was added for Q1_0, Q5_0, Q5_1, Q3_K, Q5_K, Q6_K, and MXFP4. F16 support in fill and set_rows, as well as BFloat16 support in MUL_MAT, MUL_MAT_ID, and GET_ROWS were added.
  • Hexagon Backend: A sampler was added, and q2_k and q3_k quantization types are now supported. ALLREDUCE optimizations for safe scatter mode and improvements to DMA copy/concat processing were made.
  • Windows ARM64 MSVC Build: Builds using MSVC (Microsoft Visual C++) are now officially enabled in Windows ARM64 environments.

How to Update

Specific installation commands for adopting version v0.6.0 are not explicitly stated in the provided materials.

Therefore, please refer to the instructions on the official GitHub releases page and README regarding build procedures from the latest source code or downloading precompiled binaries tailored to your environment.

Generally, updates can be performed by cloning the repository and executing build commands appropriate for your system, or by acquiring the necessary assets from the releases section.

Related Articles

What to Read Next

Sources