{"id":10033,"date":"2026-10-06T00:29:53","date_gmt":"2026-10-05T15:29:53","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/06\/ggml-v0260-released\/"},"modified":"2026-10-06T02:16:23","modified_gmt":"2026-10-05T17:16:23","slug":"ggml-v0260-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/06\/ggml-v0260-released\/","title":{"rendered":"ggml v0.26.0 Released with Sparse Flash Attention &#038; New Backends"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/ggml\">ggml-org\/ggml<\/a><\/td>\n<\/tr>\n<tr>\n<td>Version<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/ggml\/releases\/tag\/v0.26.0\">v0.26.0<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-05<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>MIT<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>ggml-org\/ggml is a C++ tensor library (MIT license) that serves as the foundation for inference engines like llama.cpp. It handles computation graph construction and execution for machine learning models as well as backend abstraction, forming the base for llama.cpp and many tools that use it. The newly released v0.26.0 compiles changes made since v0.25.3.<\/p>\n<p>The most widespread change in this release is the addition and optimization of sparse Flash Attention kernels for the SYCL, Vulkan, and Metal backends. Sparse FA for quantized KV caches has been newly implemented, which relates to memory usage and speed when handling long contexts. Users running llama.cpp-based tools on Intel GPUs (SYCL), AMD\/Intel Vulkan environments, or Apple Silicon (Metal) are likely to benefit from this update.<\/p>\n<h2>Key Changes<\/h2>\n<p><strong>Lightning Indexer Memory Halving and Tiling Optimization<\/strong><\/p>\n<p>On CUDA, Metal, and Vulkan, optimizations have been added to halve the score memory for the lightning indexer and perform tiling for keys and tokens. In addition, the MUSA (Moore Threads) backend now uses vector-version lightning indexer kernels. Improvements in memory usage and speed can be expected when running models that utilize the indexer.<\/p>\n<p><strong>CPU Backend BF16 Support and Tile mul_mat for k-quants<\/strong><\/p>\n<p>BF16 unary\/GLU\/binary\/scale operations have been added to the CPU side, and BF16 is now accepted as <code>src1<\/code> for <code>mul_mat<\/code>. Furthermore, tile <code>mul_mat<\/code> for k-quants using int8 unpack tiles and 16&#215;16 micro-kernels has been introduced. This change is relevant for users who do not have a GPU and run BF16 models or k-quant (such as Q3_K\/Q4_K\/Q5_K\/Q6_K) quantized models using only the CPU.<\/p>\n<p><strong>CUDA W4A4 Paths for NVFP4\/MXFP4 and Shared-Expert Fusion in MMVQ<\/strong><\/p>\n<p>A <code>mul_mat<\/code> path for W4A4 (NVFP4\/MXFP4), controllable from the model side via <code>llama_prec_policy<\/code>, has been added. Calculation type handling in the NVFP4 MMQ side and cuBLAS paths has also been optimized. Additionally, optimizations to fuse MoE model shared experts into MMVQ have been included. This is relevant when running NVFP4\/MXFP4 format models or MoE models on NVIDIA GPUs.<\/p>\n<p><strong>Hexagon Backend Quantization Support and Sampler Addition<\/strong><\/p>\n<p>Sampler functionality has been added to the Hexagon (Snapdragon NPU) backend, and q2_k, q3_k, and q5_k have been added to the supported quantization types. Dynamic quantizer improvements have also been made. This is relevant for users running models using the Hexagon backend on Snapdragon-equipped devices.<\/p>\n<p><strong>Faster Model Loading and Enhanced GGUF Validation<\/strong><\/p>\n<p>In addition to faster model loading, a bug where loading invalid GGUF files with extremely large KV dimensions caused a hang has been fixed. Furthermore, GGUF size validation has been tightened to detect and reject integer overflows and tensor sizes that wrap after padding. This change relates to stability and safety when handling GGUF files from unknown sources.<\/p>\n<p><strong>Windows ARM64 (MSVC) Build Enabled<\/strong><\/p>\n<p>Building with MSVC&#8217;s <code>cl.exe<\/code> has been enabled in Windows ARM64 environments. This is relevant for users who want to perform native builds on Snapdragon-powered Windows PCs and other similar devices.<\/p>\n<h2>Supported Models and Hardware<\/h2>\n<p>This version advances support for new architectures and quantization formats across a wide variety of backends. It includes changes that directly translate to performance improvements and expanded runnable models for users operating in specific hardware environments.<\/p>\n<h3>Expanded GPU and Accelerator Support<\/h3>\n<p>In the Vulkan backend, matrix multiplication (matmul) using <code>int8 coopmat1<\/code> has been implemented for AMD&#8217;s RDNA3 and RDNA4 architectures. This is expected to improve inference efficiency on the latest Radeon series and Ryzen integrated GPUs. Additionally, fine-grained adjustments have been made per device, such as tuning GDN kernels for Intel GPUs, adjusting tile sizes for Samsung GPUs with 32KB shared memory, and optimizing argmax kernel selection for Adreno GPUs.<\/p>\n<p>In Apple Silicon (Metal) environments, Flash Attention kernels for F16 format KV caches have been added. Furthermore, calculations at BF16 precision are now supported in MXFP4 format matrix operations, aiming to balance precision retention and speed. Optimizations for sparse Flash Attention are also progressing.<\/p>\n<p>For the SYCL backend targeting Intel GPUs, in addition to sparse Flash Attention support, wide loads for DMMV ESIMD and MMVQ in Q8_0 format are now supported. This improves execution speed for quantized models on Intel Arc and data center GPUs. Moreover, tensor AllReduce synchronization overhead has been reduced by utilizing pinned host buffers.<\/p>\n<h3>Enhanced WebGPU, Mobile, and Edge Environments<\/h3>\n<p>The WebGPU backend, running inside browsers, has significantly increased its supported quantization formats. Q1_0, Q5_0, Q5_1, Q3_K, Q5_K, Q6_K, and MXFP4 MMVQ (Matrix-Matrix Vector Quantization) are now supported. Furthermore, bfloat16 format matrix operations and row fetching (<code>GET_ROWS<\/code>) are now possible, greatly expanding local inference options in web browser environments.<\/p>\n<p>For Qualcomm&#8217;s Snapdragon Hexagon NPU, new quantization types such as q2_k, q3_k, and q5_k have been added. Combined with sampler feature support and dynamic quantizer improvements, power-efficient and fast inference on mobile devices becomes available for a wider range of models. The OpenCL backend also received fixes for the Q5_K <code>gemm_nonshuffle<\/code> kernel targeting Adreno.<\/p>\n<h3>Enterprise and Special Architectures<\/h3>\n<p>The OpenVINO backend has been updated to version 2026.4.1. This brings performance optimizations, an expanded set of supported operations (ops), and improved device listing displays. It also includes optimizations for handling <code>GET_ROWS<\/code> from weight views.<\/p>\n<p>In IBM&#8217;s z Systems (zDNN) backend, buffer reset implementations, memory leak fixes, and crash fixes related to zero-row tensors have been implemented, improving stability in mainframe environments. Furthermore, support for a wide range of architectures from edge to enterprise has been strengthened, such as adding the Q8_0 IME1 matrix kernel for SpacemiT X60 (RISC-V). The RPC backend received improvements using RDMA completion channels to avoid spinning.<\/p>\n<h2>How to Get It<\/h2>\n<p>If you are building and using ggml from source, you can compile the latest version by following these steps according to the quick start guide:<\/p>\n<pre><code class=\"language-bash\">git clone https:\/\/github.com\/ggml-org\/ggml\ncd ggml\n\nmkdir build &amp;&amp; cd build\ncmake..\ncmake --build. --config Release -j 8\n<\/code><\/pre>\n<p>To update an existing environment, run <code>git pull<\/code> within the repository and then re-build using the steps above. When enabling specific backends (CUDA, Vulkan, Metal, etc.), appropriate flags must be passed during <code>cmake..<\/code>. Check the project documentation for specific build flags for each backend.<\/p>\n<p><!-- lmw:releases --><\/p>\n<h2>Releases Since Our Last Article<\/h2>\n<p><em>Compiled by Local Model Watch from the project&#8217;s GitHub releases: the versions between this release and the last one we covered, which did not get separate articles.<\/em> <em>Full history: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ggml-en\/\">release tracker<\/a>.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Version<\/th>\n<th>Released<\/th>\n<th>Release notes<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>v0.25.3<\/td>\n<td>2026-09-25<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/ggml\/releases\/tag\/v0.25.3\">GitHub<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:releases --><\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/ggml-v0-25-2-released\/\">ggml v0.25.2 Released with Hardware Backend Optimizations<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/06\/llamacpp-v0-6-0-released\/\">llama.cpp v0.6.0 Released with llama_batch_ext and Metal Boosts<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/llama-cpp-v0-5-0-released\/\">llama.cpp v0.5.0 Released with Backend and Server Upgrades<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ggml-v025-0-released\/\">ggml v0.25.0 Released: FlashAttention and MoE Optimizations<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Follow this tool<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ggml-en\/\">ggml overview and release history (41 releases tracked)<\/a><\/li>\n<li><strong>Other quantization, model formats and fine-tuning<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-exllamav3-en\/\">ExLlamaV3<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-unsloth-en\/\">Unsloth<\/a><\/li>\n<li><strong>Engines mentioned in this article<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/github.com\/ggml-org\/ggml\/releases\/tag\/v0.26.0\">ggml-org\/ggml v0.26.0 Release Notes<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/23671\">PR #23671: Add alloc_buffer_n to buffer type interface<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/27952\">PR #27952: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29483\">PR #29483: WebGPU: add MMVQ support for Q1_0\/Q5_0\/Q5_1\/Q3_K\/Q5_K\/Q6_K\/MXFP4<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29852\">PR #29852: ggml-openvino: update to 2026.4.1<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/pull\/29541\">PR #29541: Add new public header include\/ggml-zdnn.h<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>ggml v0.26.0 is released, featuring sparse Flash Attention optimizations, expanded quantization support for WebGPU, and updates across multiple hardware ba<\/p>\n","protected":false},"author":1,"featured_media":10032,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[105],"tags":[505,161,163,1073,1547,511,3018],"class_list":["post-10033","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engines-and-tools","tag-apple-silicon-en","tag-ggml-en","tag-gguf-en","tag-llama-cpp-en","tag-verified","tag-vulkan-en","tag-webgpu-en"],"lang":"en","translations":{"en":10033,"ja":10031},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10033","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=10033"}],"version-history":[{"count":1,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10033\/revisions"}],"predecessor-version":[{"id":10198,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10033\/revisions\/10198"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/10032"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=10033"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=10033"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=10033"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}