Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory

Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory

At a Glance

Item Value
Repository Qwen/Qwen3.8-Flash-Next
Publisher guide Alibaba (Qwen): models and licenses
Published 2026-08-24
License qwen-community-1.0
Formats safetensors
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

The Qwen team has released “Qwen3.8-Flash-Next", an open-weight model adopting a next-generation architecture. This model is an experimental preview release that provides an early look at the design forming the basis for the future Qwen4, and is a multimodal causal language model with an integrated vision encoder.

It introduces hybrid attention combining Qwen Sparse Attention (QSA), which performs sparse processing in micro-block units, and Gated DeltaNet, as well as N-gram Embedding, which scales parameters while suppressing computational load. Through these features, it has been developed to dramatically improve inference efficiency and reduce latency in environments where models are becoming increasingly large and long-context.

Specifications

  • Parameters: Total 125B parameters (activated parameters 6B, N-gram Embedding 51B, MTP 4B)
  • Architecture: Mixture of Experts (MoE)
  • Number of experts: 512 (activated experts: routing 10 + shared 1)
  • Context length: 262,144 tokens natively (expandable up to 1,000,000 tokens)
  • Number of layers: 48
  • Hidden layer layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))

Performance

Language task benchmark comparisons published in the model card are as follows:

→ Scroll horizontally to see all columns

Qwen3.8-Flash-Next Qwen3.8-27B Qwen3.7-Plus DeepSeek-V4-Flash-0731 Claude-Opus-4.6 (Max)
# Params 125B 27B 397B 284B —
# Activated params 6B 27B 17B 13B —
# N-gram embedding params 51B — — — —
Coding
Agentic coding DeepSWE 1.1 58.7 42.2 16.5 54.4 —
Agentic coding SWE-bench Pro 62.5 61.7 55.8 56.0 53.4
Multilingual software engineering SWE-bench Multilingual 81.0 73.8 75.8 — 77.5
Repo-level code generation NL2Repo-Bench 48.1 42.3 41.1 54.2 47.6
Agent
Long-horizon office work CoWorkBench 73.9 70.7 65.1 45.1 68.2
Professional job tasks JobBench 55.7 33.4 27.6 41.3 36.6
Frontier agentic tasks Agents’ Last Exam Pass@1 24.3 Score 51.2 Pass@1 20.4 Score 42.9 Pass@1 13.2 Score 33.6 Pass@1 25.2 Score — —
Real-world tool use Toolathlon Verified (Pass@1) 73.5 67.1 50.6 70.3 —
General
Instruction following IFBench 81.3 79.5 79.1 79.2 62.5
Scientific reasoning GPQA Diamond 91.7 89.2 90.3 90.8 91.3
Multidisciplinary reasoning HLE 35.9 30.8 34.7 33.8 40.0
Competitive coding LiveCodeBench v6 91.9 90.3 89.6 90.6 88.8

According to the measurement results published by the creators, Qwen3.8-Flash-Next scores 62.5 on SWE-bench Pro and 81.0 on SWE-bench Multilingual, which measure the ability to edit and fix real GitHub issues, outperforming the other models in the comparison table. It also shows high scores of 91.9 on LiveCodeBench v6, which handles new competitive programming problems, and 91.7 on GPQA Diamond, which deals with PhD-level physics, chemistry, and biology problems. On the other hand, it stops at 35.9 on HLE, a very difficult collection of problems created by experts in various fields, falling short of Claude-Opus-4.6 (Max) which recorded 40.0. Additionally, it stands at 48.1 on NL2Repo-Bench, trailing DeepSeek-V4-Flash-0731 which showed 54.2.

Next, the measurement results for multimodal (Vision Language) performance are as follows:

→ Scroll horizontally to see all columns

Qwen3.8-Flash-Next Qwen3.8-27B Qwen3.7-Plus Claude-Opus-4.6 (Max)
Agentic Multimodal Intelligence
Multimodal tool use ClawEval-MM Pass@3 64.4 Average 60.4 Pass@3 57.4 Average 56.9 Pass@3 57.4 Average 60.1 Pass@3 52.5 Average 54.7
Application recreation RecreationBench 49.9 47.1 30.2 —
Mobile use AndroidWorld 84.5 81.9 81.0 62.0
Computer use OSWorld 2.0 Binary 19.4 Partial 52.3 Binary 19.4 Partial 48.0 Binary 2.8 Partial 21.5 —
Visual web development Vision2Web 64.0 62.9 42.1 —
General Multimodal Intelligence
Embodied intelligence ERQA 72.3 65.5 69.8 40.8
Long video understanding LVBench 76.6 72.4 76.2 63.0
Real-world perception RealWorldQA 88.5 85.9 86.9 73.9
Visual math problem solving MathVision Without CI 90.6 With CI 95.7 Without CI 90.0 With CI 94.6 Without CI 90.3 With CI 88.7 Without CI 65.5
Scientific chart analysis CharXiv (RQ) Without CI 84.6 With CI 90.6 Without CI 83.7 With CI 90.2 Without CI 85.8 With CI 85.9 Without CI 66.0

In multimodal domain measurement results as well, it records Binary 19.4 and Partial 52.3 on OSWorld 2.0, which evaluates whether tasks can be completed by operating actual OS screens, significantly increasing its numbers from the past model Qwen3.7-Plus (Binary 2.8, Partial 21.5). It also achieves 84.5 on AndroidWorld, greatly exceeding Claude-Opus-4.6 (Max)’s 62.0. However, looking at the Average score for ClawEval-MM, it is 60.4, with the difference from Qwen3.7-Plus’s 60.1 remaining very minor at just 0.3.

Strengths and Use Cases

Qwen3.8-Flash-Next is designed with a focus on agents performing long-duration autonomous tasks, software engineering, and multimodal tasks involving screen operations.

Specific strengths include the following areas:

  • Software Development and Repository Editing (Agentic Coding): It records high scores on SWE-bench Pro and SWE-bench Multilingual, which solve real repository issues, showing high aptitude for agent tasks such as debugging and adding features that look across the entire project, rather than stopping at single-function generation.
  • Long-Horizon Office Work and Practical Tasks (Long-Horizon Agents): It yields excellent results on CoWorkBench and JobBench, which evaluate office work and specialized tasks requiring lengthy steps, making it suitable for workflows that repeatedly plan and execute while autonomously calling tools.
  • GUI Operations and Multimodal Support: It integrates a vision encoder and supports image and video input. As demonstrated by benchmark results such as OSWorld 2.0, AndroidWorld, and Vision2Web, it exhibits strengths in agents that visually recognize and operate desktop and mobile device screens, as well as image-based web development.
  • Ultra-Long Context Processing and Preservation of Reasoning Processes: It supports a standard context length of 262,144 tokens and up to 1,000,000 tokens through expansion. Additionally, the thinking mode is enabled by default, equipped with a “Preserved Thinking" feature that retains the trajectory of reasoning across the entire past conversation history. This makes it possible to maintain consistency in judgment across multi-turn interactions and eliminate wasteful re-reasoning.

In addition, in the community project “Strata", variations such as “Coder", a derivative version tailored for code generation with a narrowed-down number of experts, and “Swift 1.5", a fine-tuned version that shortens thinking time to return quick responses, are also provided, envisioning tailored use cases such as programming-dedicated assistants and response-speed-prioritized chats.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 180.0B parameters

Your VRAM Quantization File size Est. memory needed
More than 402GB of VRAM (multi-GPU or CPU offload required) BF16 335.3GB 402.3GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-05): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): not registered. “Not registered" means the name is absent from that registry today, not that the model cannot run.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

The publisher distributes this model as safetensors.

License — qwen-community-1.0: A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.

Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Recent Models in the Same Size Class

Models with over 40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.

→ Scroll horizontally to see all columns

Model Parameters Smallest VRAM tier License Our article
ggml-org/GLM-5.3-Flash-GGUF 321.3B — other GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory (2026-10-03)
bartowski/Intern-S2-397B-GGUF 403.4B — apache-2.0 Intern-S2-397B-GGUF Vision-Language Model: ~102GB Memory (2026-09-15)
deepseek-ai/DeepSeek-V4.1-Flash 763.2B — mit DeepSeek-V4.1-Flash 552B Multimodal MoE Model: ~570GB Memory (2026-09-10)

How to Get It

The model weights for Qwen3.8-Flash-Next are published on Hugging Face in safetensors format. The “qwen-community-1.0" license applies, and it is distributed as an open-weight model that does not require specific advance applications.

Multiple methods are provided to use and run this model.

For server environments and high-performance execution environments, the official model card provides instructions on how to use the following inference frameworks:

  • SGLang
  • vLLM
  • TokenSpeed

Additionally, as an inference project designed to run on typical consumer PC environments, the open-source software “Strata", developed under the MIT license, has been released. It supports Windows and Linux environments, and after downloading the repository, running the following scripts allows setup and execution to proceed interactively:

  • For Windows: Run START-HERE.bat.
  • For Linux: Run ./setup.sh.

The installer detects hardware, selects an appropriate quantized model, and automatically handles everything from downloading to startup. If using AI coding assistants such as Cursor or Claude Code, it also features functionality to have the assistant delegate tasks from configuration checking to startup by passing the instruction prompts prepared in the repository.

After startup, you can access the Web UI (http://127.0.0.1:8080) from your browser to chat or check operation status, and you can also use it from various clients and development tools via the following compatible endpoints:

  • OpenAI-compatible API: Specify http://127.0.0.1:8080/v1 for the base URL.
  • Anthropic-compatible API: Specify http://127.0.0.1:8080/v1/messages.

Note that during the initial startup of the model, loading into main memory and allocation to the GPU take place, so depending on the environment, responses may appear to stop for a few minutes, but this is explained as normal background processing.

Related Articles

What to Read Next

Sources