{"id":10268,"date":"2026-10-06T07:14:59","date_gmt":"2026-10-05T22:14:59","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/06\/magnitudedev-magnitude-026-released\/"},"modified":"2026-10-06T07:14:59","modified_gmt":"2026-10-05T22:14:59","slug":"magnitudedev-magnitude-026-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/06\/magnitudedev-magnitude-026-released\/","title":{"rendered":"magnitudedev\/magnitude v0.2.6 Released: Major Mac Speedups"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\">magnitudedev\/magnitude<\/a><\/td>\n<\/tr>\n<tr>\n<td>Version<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/@magnitudedev\/cli@0.2.6\">@magnitudedev\/cli@0.2.6<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-06<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>Apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>The latest version of <code>magnitudedev\/magnitude<\/code>, <code>@magnitudedev\/cli@0.2.6<\/code>, has been released. <code>magnitudedev\/magnitude<\/code> is an open-source inference engine written in Rust (licensed under Apache-2.0) that automatically compiles, tunes, and optimizes kernels to match the execution hardware. It supports Apple Silicon, NVIDIA, AMD, and CPU, and features the ability to run open models up to 2x faster compared to llama.cpp.<\/p>\n<p>The most significant changes in this update are a dramatic improvement in prompt processing speed on Mac environments (especially M5 and later chips, and Apple Silicon generations such as M1\/M2\/M4 Pro) and enhanced stability during long-prompt processing. For users running local LLMs on Macs, this is a highly valuable update in terms of both performance and stability.<\/p>\n<h2>Key Changes<\/h2>\n<h3>Dramatically Faster Prompt Processing on Mac Environments<\/h3>\n<p>Prompt processing speed has improved by approximately 50% on M5 and later Macs. For example, when running Qwen3.5-4B with a 64K context length, prompt processing speed has increased from the previous 652 tok\/s to approximately 970 tok\/s. This is due to prompt attention processing directly reading key-value tensors through GPU tensor operations without staging, as well as improved kernel tuning that avoids maintaining unstable and slow default settings.<\/p>\n<p>Additionally, for all Macs, weight block decoding has been changed from every 64 rows to once every up to 512 rows. Without altering the output results, this speeds up matrix multiplication by 7\u201311% on M4 Pro and 17\u201325% on M1. On M4 Pro with Qwen3.5-4B (64K context), prompt processing improves from the previous 513 tok\/s to 534 tok\/s.<\/p>\n<h3>Fixed Crashes During Long-Context Processing on M1\/M2 Macs<\/h3>\n<p>A bug has been fixed where M1 and M2 Macs encountered a &#8220;device lost: \u2026 Impacting Interactivity&#8221; error while processing long prompts, leaving the model unloaded.<\/p>\n<p>In this version, similar to llama.cpp, processing has been added to relax macOS GPU watchdog limits at startup. As a result, when using Gemma 4 26B, long prompt processing of 35k tokens and 69k tokens\u2014which previously failed around the 20k token mark\u2014can now complete without failing midway.<\/p>\n<h3>Resolved Metal Compilation Errors on Certain Macs<\/h3>\n<p>An issue where loading models such as Qwen 3.6 on certain Mac environments like M5 Max failed with a Metal shader compilation error reading &#8220;no template named &#8216;extents&#8217; in namespace &#8216;metal'&#8221; has been fixed. Through the fix in <a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/pull\/166\">PR #166<\/a>, compiling Metal kernels now uses the exact same language version confirmed during device detection, allowing kernels utilizing tensor operations to build successfully on supported devices.<\/p>\n<h3>Prevention of 502 Timeouts on Long-Running Requests<\/h3>\n<p>An issue where long-running processing requests for local models failed with a 502 error after approximately 5 minutes has been fixed. For non-streaming generation and long prompt processing, data is not transmitted until generation is complete, but improvements were made so that connections to the engine do not time out while waiting.<\/p>\n<h3>Prevention of Force Quits Due to JSON Errors on Windows<\/h3>\n<p>An issue has been fixed where the engine abnormally aborted on Windows environments when chat templates or tool calls generated invalid JSON. Enabling C++ exception handling during the MSVC build allows the engine to properly report JSON errors without crashing.<\/p>\n<h3>Various Bug Fixes for Tool Calling and Session Processing<\/h3>\n<p>Multiple bugs related to agent features and API integrations have been fixed.<\/p>\n<ul>\n<li><strong>Fixed Codex Tool Result Transmission Error<\/strong>: Fixed an issue where Codex failed to send tool execution results to the local model. Reasoning items where clients set <code>content<\/code> or <code>summary<\/code> to null upon resending are now accepted as empty.<\/li>\n<li><strong>Skipping Empty Assistant Turns<\/strong>: Fixed an issue where OpenCode sessions were rejected with an &#8220;assistant content is required unless tool_calls are present&#8221; error when a step failed before generating output. Both Chat Completions and Anthropic API now skip empty assistant turns in the history.<\/li>\n<li><strong>Stabilized Tool Calling in Qwen Models<\/strong>: Fixed an issue where Oh My Pi or OpenClaw failed when Qwen models included free-form JSON in tool optional properties. Even if parallel tool calling grammars cannot be efficiently compiled, it now falls back to allow a single tool call per turn if the single-call grammar is compilable.<\/li>\n<li><strong>Improved Error Reporting on Headless Startup<\/strong>: Fixed an issue where <code>magnitude serve<\/code> running in the background (headless) crashed with an EBADF error instead of reporting startup failures, ensuring that errors causing termination are now properly reported.<\/li>\n<\/ul>\n<h2>Supported Models and Hardware<\/h2>\n<p>No newly supported model architectures are listed in the documentation for this release, but stability and performance on existing supported hardware have been significantly improved. Supported platforms remain Apple Silicon, NVIDIA, AMD, and CPU.<\/p>\n<p>Particular mention is made of the following environments:<\/p>\n<ul>\n<li>Apple Silicon (M1\/M2): Models such as Gemma 4 26B can now process long prompts of 35k and 69k tokens.<\/li>\n<li>Apple Silicon (M4 Pro): Prompt processing for Qwen3.5-4B (64K context) has been accelerated.<\/li>\n<li>Apple Silicon (M5 and later \/ M5 Max): In addition to faster Qwen3.5-4B processing speed, models like Qwen 3.6 can now be loaded without Metal shader compilation errors.<\/li>\n<\/ul>\n<p>Note that these are improvements to the behavior of already-supported models and do not represent an expansion of support for new model families.<\/p>\n<h2>How to Get It<\/h2>\n<p>Since the documentation does not contain specific installation commands, please refer to the release page for instructions on updating to the latest version.<\/p>\n<p>If Mac users have encountered failures during long prompt processing or Metal compilation errors, updating to this version may resolve them, so prioritizing an update is recommended.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/magnitude-cli-v0-2-5-released\/\">Magnitude CLI v0.2.5 Released: M5 Mac &amp; MoE Speedups<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/magnitude-0-2-4-released\/\">Magnitude 0.2.4 Released: Faster Inference and Lower Memory<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/01\/magnitude-self-optimizing-inference-engine-for-agents\/\">Magnitude: Up to 2x Faster Than llama.cpp Officially, 21x Slower on Our CPU<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Follow this tool<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-magnitude-en\/\">Magnitude overview and release history (5 releases tracked)<\/a><\/li>\n<li><strong>Other inference engines and runtimes<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li>magnitudedev\/magnitude @magnitudedev\/cli@0.2.6 Release Notes<\/li>\n<li>PR #166: Fix Metal kernel compilation version<\/li>\n<\/ul>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.6\">https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.6<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>magnitudedev\/magnitude v0.2.6 is out, bringing massive prompt processing speedups and stability fixes for Apple Silicon Mac users.<\/p>\n","protected":false},"author":1,"featured_media":10267,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[105],"tags":[505,2791,3057,2793,1547],"class_list":["post-10268","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engines-and-tools","tag-apple-silicon-en","tag-magnitude-en","tag-magnitudedev-en","tag-rust-en","tag-verified"],"lang":"en","translations":{"en":10268,"ja":10266},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10268","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=10268"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/10268\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/10267"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=10268"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=10268"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=10268"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}