{"id":9202,"date":"2026-10-03T00:19:12","date_gmt":"2026-10-02T15:19:12","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/03\/magnitude-0-2-4-released\/"},"modified":"2026-10-03T00:20:34","modified_gmt":"2026-10-02T15:20:34","slug":"magnitude-0-2-4-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/magnitude-0-2-4-released\/","title":{"rendered":"Magnitude 0.2.4 Released: Faster Inference and Lower Memory"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\">magnitudedev\/magnitude<\/a><\/td>\n<\/tr>\n<tr>\n<td>Version<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/@magnitudedev\/cli@0.2.4\">@magnitudedev\/cli@0.2.4<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-02<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>Apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>Version 0.2.4 (@magnitudedev\/cli@0.2.4) of &#8220;<a href=\"https:\/\/github.com\/magnitudedev\/magnitude\">magnitudedev\/magnitude<\/a>&#8221; has been <a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.4\">released<\/a>. magnitudedev\/magnitude is an open-source inference engine written in Rust and designed for agents, licensed under Apache-2.0. By compiling and tuning kernels directly on your device and optimizing them for the hardware, it can run open models up to 2x faster compared to llama.cpp. It supports Apple Silicon, NVIDIA, AMD, and CPU-only operation.<\/p>\n<p>The most significant impacts of this update for users are the drastic reduction in memory required to run models and the reduction of initial load kernel tuning times to a maximum of about 1 minute across all devices. This makes it more comfortable to handle larger models and longer contexts while dramatically eliminating startup wait time stress.<\/p>\n<h2>Key Changes<\/h2>\n<h3>Significant Reduction in Memory Usage<\/h3>\n<p>Memory required by models beyond their weights has been reduced, enabling the operation of larger models and longer contexts. As a specific example, the working memory required when running Qwen3.5-4B has been reduced from 3.7 GB to 120 MB, and the reserved memory for execution has decreased from 10.0 GB to 5.9 GB. There is no drop in inference speed resulting from this memory reduction. Additionally, image processing memory is now allocated when the first image is processed and released when idling.<\/p>\n<h3>Limiting Initial Tuning Time to a Maximum of About 1 Minute<\/h3>\n<p>An issue where initial model loading took about 20 to 50 minutes on CPU-only machines and appeared to hang has been fixed. Previously, the lower the processing capability of the device, the longer the tuning took, but with this change, tuning completes in a maximum of about 1 minute on all devices, accompanied by progress display. For the initial load of Qwen3.5-4B on an M4 Max environment, the tuning time was reduced to about 50 seconds (formerly 149 seconds on GPU and 279 seconds on CPU), while maintaining subsequent inference speeds. Even for large models like Gemma 4 26B, test data preparation is consolidated into a single step, and kernels that do not finish within 1 minute retain their default settings, keeping the tuning time within 1 minute.<\/p>\n<h3>Speeding Up Q6_K Models on x86 CPUs<\/h3>\n<p>In x86 architecture CPU environments, inference speeds for models with Q6_K weights have been improved. This is because the processing was changed to convert half-precision scales without routing through slow processor paths. Users running Q6_K quantized models on x86 CPU environments will directly benefit from this.<\/p>\n<h3>Improved Error Handling on Request Failures<\/h3>\n<p>A bug where requests failed midway\u2014such as running out of memory during a long conversation\u2014yet returned an empty response and appeared to succeed has been fixed. After the update, it returns a 503 error code, a <code>Retry-After<\/code> header, and an error message. This allows agent harnesses like Pi, OpenCode, and Claude Code to automatically retry. Furthermore, memory pressure and request failures from the inference engine that previously failed silently are now logged.<\/p>\n<h3>Fix for Load Failures on M1\/M2 Macs<\/h3>\n<p>On Macs equipped with M1 and M2 chips, a bug where the default kernel settings exceeded the chip&#8217;s allowed thread count and caused model loading to fail has been fixed. With this fix, tuning now begins from the closest workable setting. This is an effective fix for users with M1 or M2 Macs who previously experienced startup failures.<\/p>\n<h3>Fix for Startup Errors via SSH<\/h3>\n<p>A problem where the <code>magnitude serve<\/code> command failed to start when accessing a Mac via SSH with no console user logged in has been resolved. This improves convenience for developers who start and use servers on Macs remotely.<\/p>\n<h2>Supported Models and Hardware<\/h2>\n<p>In this version 0.2.4, operational stability and performance in specific model architectures and hardware environments have been substantially improved.<\/p>\n<h3>Supported Models<\/h3>\n<p>Operation and optimization for the following models have been confirmed in this version:<\/p>\n<ul>\n<li>\n<p><strong>Gemma 4 26B<\/strong><br \/>\n  For large models like Gemma 4 26B, kernel tuning during the initial load has been optimized to finish within 1 minute. The test data preparation process has been streamlined, and a mechanism has been introduced to maintain default settings for kernels that do not complete processing within the timeframe, allowing large models to start smoothly.<\/p>\n<\/li>\n<li>\n<p><strong>Qwen3.5-4B<\/strong><br \/>\n  Memory efficiency for Qwen3.5-4B has been dramatically improved. Working memory and reserved execution memory required for operation have been significantly reduced, enabling memory-efficient execution without sacrificing inference speed.<\/p>\n<\/li>\n<li>\n<p><strong>Q6_K Quantized Models<\/strong><br \/>\n  For models with Q6_K weights, inference speeds on x86 architecture CPUs have been improved.<\/p>\n<\/li>\n<\/ul>\n<h3>Supported Hardware<\/h3>\n<p>magnitudedev\/magnitude supports Apple Silicon, NVIDIA, AMD, and CPU-only environments, but this update specifically brings fixes and optimizations to the following hardware environments:<\/p>\n<ul>\n<li>\n<p><strong>Apple Silicon (M1, M2, M4 Max)<\/strong><br \/>\n  On Macs with M1 and M2 chips, the issue where default kernel settings exceeded the chip&#8217;s allowed thread count and caused load failures has been fixed. Additionally, in M4 Max environments, optimizations for newer chips have progressed, reducing the initial load tuning time for Qwen3.5-4B to about 50 seconds.<\/p>\n<\/li>\n<li>\n<p><strong>x86 CPU<\/strong><br \/>\n  In x86 architecture CPU environments, half-precision scale conversion processing for Q6_K quantized models has been accelerated, boosting CPU-only inference performance.<\/p>\n<\/li>\n<li>\n<p><strong>Headless Macs (via SSH)<\/strong><br \/>\n  It now supports use cases where <code>magnitude serve<\/code> is launched by remotely accessing a Mac via SSH when no user is logged into the console. This enhances convenience when operating Macs remotely for server purposes.<\/p>\n<\/li>\n<\/ul>\n<h2>How to Get It<\/h2>\n<p>Specific installation steps and command-line instructions to introduce this release (@magnitudedev\/cli@0.2.4) are not directly listed in the provided materials.<\/p>\n<p>Therefore, when performing an update, please follow the instructions on the project&#8217;s official release page or GitHub repository.<\/p>\n<p>Generally, the following command is used when starting this tool:<\/p>\n<pre><code class=\"language-bash\">magnitude serve\n<\/code><\/pre>\n<p>For detailed installation steps and how to obtain the latest CLI tool, please check the <a href=\"https:\/\/github.com\/magnitudedev\/magnitude\">GitHub Repository<\/a> and the <a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.4\">Official Release Page<\/a>.<\/p>\n<p><!-- lmw:releases --><\/p>\n<h2>Releases Since Our Last Article<\/h2>\n<p><em>Compiled by Local Model Watch from the project&#8217;s GitHub releases: the versions between this release and the last one we covered, which did not get separate articles.<\/em> <em>Full history: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-magnitude-en\/\">release tracker<\/a>.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Version<\/th>\n<th>Released<\/th>\n<th>Release notes<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>@magnitudedev\/cli@0.2.3<\/td>\n<td>2026-10-01<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.3\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>@magnitudedev\/cli@0.2.2<\/td>\n<td>2026-10-01<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.2\">GitHub<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:releases --><\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/01\/magnitude-self-optimizing-inference-engine-for-agents\/\">Magnitude: Up to 2x Faster Than llama.cpp Officially, 21x Slower on Our CPU<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Follow this tool<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-magnitude-en\/\">Magnitude overview and release history (3 releases tracked)<\/a><\/li>\n<li><strong>Other inference engines and runtimes<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<p>This article is based on the following official materials and commit information:<\/p>\n<ul>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.4\">magnitudedev\/magnitude @magnitudedev\/cli@0.2.4 Release Notes<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\">magnitudedev\/magnitude Official Repository<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/commit\/f0498ce285e4e97815ba16dd74044ce5efebb63a\">Commit f0498ce (Improved initial load kernel tuning for large models like Gemma 4 26B)<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/commit\/e38af0e9a232293dca12c5bad3a91b96e58ca12e\">Commit e38af0e (Fixed load failures on M1\/M2 Macs)<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/commit\/cef7f3bb9c2dd06d96aeb59dde8039b6f83386a5\">Commit cef7f3b (Improved error handling on request failure due to out of memory, etc.)<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/commit\/4713d97b4b913a3ed2194f53826ae01658600839\">Commit 4713d97 (Fixed startup errors for magnitude serve on Macs via SSH)<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/commit\/7eabf993ebbd96a4ab2c067d2c7f2e7d3e4ff8f8\">Commit 7eabf99 (Reduced memory usage, shortened initial tuning time, accelerated Q6_K models on x86 CPU)<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Magnitude 0.2.4 is released with significantly reduced memory usage, faster kernel tuning, and improved error handling for local open models.<\/p>\n","protected":false},"author":1,"featured_media":9201,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[105],"tags":[505,163,2791,2793,1547,1193],"class_list":["post-9202","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engines-and-tools","tag-apple-silicon-en","tag-gguf-en","tag-magnitude-en","tag-rust-en","tag-verified","tag--en"],"lang":"en","translations":{"en":9202,"ja":9200},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9202","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9202"}],"version-history":[{"count":1,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9202\/revisions"}],"predecessor-version":[{"id":9302,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9202\/revisions\/9302"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9201"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9202"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9202"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9202"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}