Magnitude CLI v0.2.5 Released: M5 Mac & MoE Speedups

Magnitude CLI v0.2.5 Released: M5 Mac & MoE Speedups

At a Glance

Item Value
Repository magnitudedev/magnitude
Version @magnitudedev/cli@0.2.5
Published 2026-10-03
License Apache-2.0
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

Version 0.2.5 of @magnitudedev/cli, the CLI tool for the agent-oriented inference engine magnitudedev/magnitude, has been released. Magnitude is an open-source inference engine written in Rust that automatically compiles and optimizes kernels to match the hardware of the execution environment, provided under the Apache-2.0 license. It supports a wide range of environments including Apple Silicon, NVIDIA GPUs, AMD GPUs, and CPUs.

The biggest changes in this version are a significant speedup in prompt processing (prefill) on Mac environments with M5 and later, and improved generation speed for Mixture-of-Experts (MoE) models combined with Multi-token prediction. On M5 environments, prompt processing speed reaches approximately 2x, and generation speed for supported MoE models and long contexts is also improved. Users running target hardware or applicable models locally can significantly reduce inference wait times, making this update well worth considering.

Key Changes

Prompt Processing on M5 and Later Macs Speeded Up by Approx. 2x

Optimizations have been introduced to execute matrix multiplication and attention processing using GPU Tensor operations. This speeds up prompt processing on M5 and later Macs by approximately 2x. For example, when running Qwen3.5-4B with a 64K context, the prompt processing speed doubles from the conventional 308 tok/s to 649 tok/s, and the Time to First Token is halved from 213 seconds to 101 seconds. There is no difference in the output content resulting from this change. Note that processing speeds remain unchanged for Mac generations prior to M5. No configuration changes are required, and you can benefit immediately after updating.

Generation Speed Boost for Multi-token Prediction in MoE Models

For MoE models equipped with Multi-token prediction, the specification has been changed to draft (pre-generate) up to 3 tokens ahead instead of the conventional 1 token ahead. This significantly improves generation speed. Specifically, when running Qwen3.6-35B-A3B, speeds on the GB10 environment increase from the conventional 97–105 tok/s to 106–138 tok/s, and on the M4 Pro environment from 103–111 tok/s to 109–138 tok/s. This directly impacts environments operating applicable MoE models.

Generation Speed Improvements in 16K Context

For generation processing in 16K long contexts, optimizations were made to reduce the read volume of the output layer and consolidate each step’s processing into fewer GPU launches. This improvement increases generation speed while maintaining output quality by approximately 8% on Mac environments (from 67.0 tok/s to 72.4 tok/s) and approximately 11% on NVIDIA GPU environments (from 66.3 tok/s to 73.4 tok/s). General users running models in long contexts will benefit from this.

Reduction in Verification Time for Multi-token Prediction on NVIDIA GPUs

When executing Multi-token prediction on NVIDIA GPUs, efficiency has been improved so that if multiple draft tokens select the same expert, the expert weights are read only once in each step. This reduces verification processing time by up to 11%. Processing speeds are improved for environments utilizing the relevant speculative decoding feature on NVIDIA GPUs.

Fix for Model Load Failure Bugs on M1 and M2 Macs

An issue where models failed to load on M1 and M2 generation Macs due to a threadgroup limit error stating “requests N threads per threadgroup; the pipeline allows M" has been fixed. Metal kernels are now built to accept the thread count at startup, resolving load failures such as Qwen3.8 27B on M1 Max environments. Additionally, an issue where models with 16 or more query heads per key (Gemma 4 12B, Muse Glimmer 30B, Nemotron 3.5 Lightning, Qwen3.5 122B, Nemotron 3 Super) failed to work on M1/M2 Macs has also been fixed. By splitting key query heads into groups during prefill attention, processing is kept within the thread limits of each Mac. There is no impact on speed or output in other environments.

Fix for Shader Compilation Errors on Vulkan GPUs

In Vulkan environments, a bug causing shader compilation errors when loading models using DFlash2 speculative decoding (such as Qwen3.8 27B) or Nemotron models has been fixed. The system has been revised so that all GPU kernels—including all kernels loaded by models listed in the catalog—are pre-compiled for Vulkan, CUDA, and Metal prior to release. Target models will now launch correctly in environments performing inference via Vulkan.

Improved Error Display When Deleting Running or Loading Models

Behavior has been fixed where attempting to delete a model currently running or in the process of loading caused a misleading error. Going forward, a stop process will automatically be interposed before deleting a model, and that fact will be clearly indicated in the confirmation message.

Supported Models and Hardware

This version significantly improves model compatibility in specific hardware environments, newly enabling models that previously had operational limitations.

Enhanced Support for Apple Silicon (M1 / M2 / M5 and Later)

Errors caused by Metal kernel threadgroup limits have been resolved on M1 and M2 generation Macs. As a result, the following models now run newly or stably in M1 and M2 Mac environments:

  • Qwen3.8 27B (Fixed load failures on M1 Max, etc.)
  • Gemma 4 12B
  • Muse Glimmer 30B
  • Nemotron 3.5 Lightning
  • Qwen3.5 122B
  • Nemotron 3 Super

These models are characterized by having 16 or more query heads per key, and in previous versions, processing limits were exceeded on M1 and M2 Macs, causing errors. With this update, a mechanism to split key query heads into groups during prefill attention has been introduced, making it possible to run them within the limits of all Mac environments without sacrificing speed or output.

Furthermore, on the latest Mac environments with M5 and later, new support has been added for matrix multiplication and attention processing using GPU tensor operations. This maximizes the performance of the latest hardware, such as approximately doubling prompt processing speeds for models like Qwen3.5-4B.

Expansion of Supported Models in Vulkan GPU Environments

In Vulkan-compatible GPU environments, models that previously failed to launch due to shader compilation errors can now be loaded normally. Specifically, the following models and configurations are newly supported:

  • Models using DFlash2 speculative decoding (Qwen3.8 27B, etc.)
  • Various models in the Nemotron series

Because all GPU kernels are now pre-compiled for Vulkan, CUDA, and Metal environments prior to release, errors can be avoided and stable inference executed even in Vulkan environments.

Other Supported Hardware and Models

Optimizations in this version also bring benefits in the form of improved generation speeds for the following hardware and model combinations:

  • GB10 (NVIDIA GPU) and M4 Pro (Mac): Generation speed is improved for the MoE model “Qwen3.6-35B-A3B" with Multi-token prediction enabled.
  • General Macs and NVIDIA GPUs: Generation speed is improved across all models utilizing 16K long contexts.

How to Get It

The provided materials do not contain specific installation or update commands for this version. Therefore, for details on update procedures in your environment, please check the release page or the official repository information directly.

Normally, CLI tool updates are performed via package managers, but for optimal installation procedures and dependency checks for each environment, we recommend referring to the official documentation.

Related Articles

What to Read Next

Sources