ggml v0.26.0 Released with Sparse Flash Attention & New Backends

ggml v0.26.0 Released with Sparse Flash Attention & New Backends

At a Glance

Item Value
Repository ggml-org/ggml
Version v0.26.0
Published 2026-10-05
License MIT
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

ggml-org/ggml is a C++ tensor library (MIT license) that serves as the foundation for inference engines like llama.cpp. It handles computation graph construction and execution for machine learning models as well as backend abstraction, forming the base for llama.cpp and many tools that use it. The newly released v0.26.0 compiles changes made since v0.25.3.

The most widespread change in this release is the addition and optimization of sparse Flash Attention kernels for the SYCL, Vulkan, and Metal backends. Sparse FA for quantized KV caches has been newly implemented, which relates to memory usage and speed when handling long contexts. Users running llama.cpp-based tools on Intel GPUs (SYCL), AMD/Intel Vulkan environments, or Apple Silicon (Metal) are likely to benefit from this update.

Key Changes

Lightning Indexer Memory Halving and Tiling Optimization

On CUDA, Metal, and Vulkan, optimizations have been added to halve the score memory for the lightning indexer and perform tiling for keys and tokens. In addition, the MUSA (Moore Threads) backend now uses vector-version lightning indexer kernels. Improvements in memory usage and speed can be expected when running models that utilize the indexer.

CPU Backend BF16 Support and Tile mul_mat for k-quants

BF16 unary/GLU/binary/scale operations have been added to the CPU side, and BF16 is now accepted as src1 for mul_mat. Furthermore, tile mul_mat for k-quants using int8 unpack tiles and 16×16 micro-kernels has been introduced. This change is relevant for users who do not have a GPU and run BF16 models or k-quant (such as Q3_K/Q4_K/Q5_K/Q6_K) quantized models using only the CPU.

CUDA W4A4 Paths for NVFP4/MXFP4 and Shared-Expert Fusion in MMVQ

A mul_mat path for W4A4 (NVFP4/MXFP4), controllable from the model side via llama_prec_policy, has been added. Calculation type handling in the NVFP4 MMQ side and cuBLAS paths has also been optimized. Additionally, optimizations to fuse MoE model shared experts into MMVQ have been included. This is relevant when running NVFP4/MXFP4 format models or MoE models on NVIDIA GPUs.

Hexagon Backend Quantization Support and Sampler Addition

Sampler functionality has been added to the Hexagon (Snapdragon NPU) backend, and q2_k, q3_k, and q5_k have been added to the supported quantization types. Dynamic quantizer improvements have also been made. This is relevant for users running models using the Hexagon backend on Snapdragon-equipped devices.

Faster Model Loading and Enhanced GGUF Validation

In addition to faster model loading, a bug where loading invalid GGUF files with extremely large KV dimensions caused a hang has been fixed. Furthermore, GGUF size validation has been tightened to detect and reject integer overflows and tensor sizes that wrap after padding. This change relates to stability and safety when handling GGUF files from unknown sources.

Windows ARM64 (MSVC) Build Enabled

Building with MSVC’s cl.exe has been enabled in Windows ARM64 environments. This is relevant for users who want to perform native builds on Snapdragon-powered Windows PCs and other similar devices.

Supported Models and Hardware

This version advances support for new architectures and quantization formats across a wide variety of backends. It includes changes that directly translate to performance improvements and expanded runnable models for users operating in specific hardware environments.

Expanded GPU and Accelerator Support

In the Vulkan backend, matrix multiplication (matmul) using int8 coopmat1 has been implemented for AMD’s RDNA3 and RDNA4 architectures. This is expected to improve inference efficiency on the latest Radeon series and Ryzen integrated GPUs. Additionally, fine-grained adjustments have been made per device, such as tuning GDN kernels for Intel GPUs, adjusting tile sizes for Samsung GPUs with 32KB shared memory, and optimizing argmax kernel selection for Adreno GPUs.

In Apple Silicon (Metal) environments, Flash Attention kernels for F16 format KV caches have been added. Furthermore, calculations at BF16 precision are now supported in MXFP4 format matrix operations, aiming to balance precision retention and speed. Optimizations for sparse Flash Attention are also progressing.

For the SYCL backend targeting Intel GPUs, in addition to sparse Flash Attention support, wide loads for DMMV ESIMD and MMVQ in Q8_0 format are now supported. This improves execution speed for quantized models on Intel Arc and data center GPUs. Moreover, tensor AllReduce synchronization overhead has been reduced by utilizing pinned host buffers.

Enhanced WebGPU, Mobile, and Edge Environments

The WebGPU backend, running inside browsers, has significantly increased its supported quantization formats. Q1_0, Q5_0, Q5_1, Q3_K, Q5_K, Q6_K, and MXFP4 MMVQ (Matrix-Matrix Vector Quantization) are now supported. Furthermore, bfloat16 format matrix operations and row fetching (GET_ROWS) are now possible, greatly expanding local inference options in web browser environments.

For Qualcomm’s Snapdragon Hexagon NPU, new quantization types such as q2_k, q3_k, and q5_k have been added. Combined with sampler feature support and dynamic quantizer improvements, power-efficient and fast inference on mobile devices becomes available for a wider range of models. The OpenCL backend also received fixes for the Q5_K gemm_nonshuffle kernel targeting Adreno.

Enterprise and Special Architectures

The OpenVINO backend has been updated to version 2026.4.1. This brings performance optimizations, an expanded set of supported operations (ops), and improved device listing displays. It also includes optimizations for handling GET_ROWS from weight views.

In IBM’s z Systems (zDNN) backend, buffer reset implementations, memory leak fixes, and crash fixes related to zero-row tensors have been implemented, improving stability in mainframe environments. Furthermore, support for a wide range of architectures from edge to enterprise has been strengthened, such as adding the Q8_0 IME1 matrix kernel for SpacemiT X60 (RISC-V). The RPC backend received improvements using RDMA completion channels to avoid spinning.

How to Get It

If you are building and using ggml from source, you can compile the latest version by following these steps according to the quick start guide:

git clone https://github.com/ggml-org/ggml
cd ggml

mkdir build && cd build
cmake..
cmake --build. --config Release -j 8

To update an existing environment, run git pull within the repository and then re-build using the steps above. When enabling specific backends (CUDA, Vulkan, Metal, etc.), appropriate flags must be passed during cmake... Check the project documentation for specific build flags for each backend.

Releases Since Our Last Article

Compiled by Local Model Watch from the project’s GitHub releases: the versions between this release and the last one we covered, which did not get separate articles. Full history: release tracker.

Version Released Release notes
v0.25.3 2026-09-25 GitHub

Related Articles

What to Read Next

Sources