Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory

At a Glance
| Item | Value |
|---|---|
| Repository | Qwen/Qwen3.8-Flash-Next |
| Publisher guide | Alibaba (Qwen): models and licenses |
| Published | 2026-08-24 |
| License | qwen-community-1.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
The Qwen team has released “Qwen3.8-Flash-Next", an open-weight model adopting a next-generation architecture. This model is an experimental preview release that provides an early look at the design forming the basis for the future Qwen4, and is a multimodal causal language model with an integrated vision encoder.
It introduces hybrid attention combining Qwen Sparse Attention (QSA), which performs sparse processing in micro-block units, and Gated DeltaNet, as well as N-gram Embedding, which scales parameters while suppressing computational load. Through these features, it has been developed to dramatically improve inference efficiency and reduce latency in environments where models are becoming increasingly large and long-context.
Specifications
- Parameters: Total 125B parameters (activated parameters 6B, N-gram Embedding 51B, MTP 4B)
- Architecture: Mixture of Experts (MoE)
- Number of experts: 512 (activated experts: routing 10 + shared 1)
- Context length: 262,144 tokens natively (expandable up to 1,000,000 tokens)
- Number of layers: 48
- Hidden layer layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))
Performance
Language task benchmark comparisons published in the model card are as follows:
→ Scroll horizontally to see all columns
| Qwen3.8-Flash-Next | Qwen3.8-27B | Qwen3.7-Plus | DeepSeek-V4-Flash-0731 | Claude-Opus-4.6 (Max) | |
|---|---|---|---|---|---|
| # Params | 125B | 27B | 397B | 284B | — |
| # Activated params | 6B | 27B | 17B | 13B | — |
| # N-gram embedding params | 51B | — | — | — | — |
| Coding | |||||
| Agentic coding DeepSWE 1.1 | 58.7 | 42.2 | 16.5 | 54.4 | — |
| Agentic coding SWE-bench Pro | 62.5 | 61.7 | 55.8 | 56.0 | 53.4 |
| Multilingual software engineering SWE-bench Multilingual | 81.0 | 73.8 | 75.8 | — | 77.5 |
| Repo-level code generation NL2Repo-Bench | 48.1 | 42.3 | 41.1 | 54.2 | 47.6 |
| Agent | |||||
| Long-horizon office work CoWorkBench | 73.9 | 70.7 | 65.1 | 45.1 | 68.2 |
| Professional job tasks JobBench | 55.7 | 33.4 | 27.6 | 41.3 | 36.6 |
| Frontier agentic tasks Agents’ Last Exam | Pass@1 24.3 Score 51.2 | Pass@1 20.4 Score 42.9 | Pass@1 13.2 Score 33.6 | Pass@1 25.2 Score — | — |
| Real-world tool use Toolathlon Verified (Pass@1) | 73.5 | 67.1 | 50.6 | 70.3 | — |
| General | |||||
| Instruction following IFBench | 81.3 | 79.5 | 79.1 | 79.2 | 62.5 |
| Scientific reasoning GPQA Diamond | 91.7 | 89.2 | 90.3 | 90.8 | 91.3 |
| Multidisciplinary reasoning HLE | 35.9 | 30.8 | 34.7 | 33.8 | 40.0 |
| Competitive coding LiveCodeBench v6 | 91.9 | 90.3 | 89.6 | 90.6 | 88.8 |
According to the measurement results published by the creators, Qwen3.8-Flash-Next scores 62.5 on SWE-bench Pro and 81.0 on SWE-bench Multilingual, which measure the ability to edit and fix real GitHub issues, outperforming the other models in the comparison table. It also shows high scores of 91.9 on LiveCodeBench v6, which handles new competitive programming problems, and 91.7 on GPQA Diamond, which deals with PhD-level physics, chemistry, and biology problems. On the other hand, it stops at 35.9 on HLE, a very difficult collection of problems created by experts in various fields, falling short of Claude-Opus-4.6 (Max) which recorded 40.0. Additionally, it stands at 48.1 on NL2Repo-Bench, trailing DeepSeek-V4-Flash-0731 which showed 54.2.
Next, the measurement results for multimodal (Vision Language) performance are as follows:
→ Scroll horizontally to see all columns
| Qwen3.8-Flash-Next | Qwen3.8-27B | Qwen3.7-Plus | Claude-Opus-4.6 (Max) | |
|---|---|---|---|---|
| Agentic Multimodal Intelligence | ||||
| Multimodal tool use ClawEval-MM | Pass@3 64.4 Average 60.4 | Pass@3 57.4 Average 56.9 | Pass@3 57.4 Average 60.1 | Pass@3 52.5 Average 54.7 |
| Application recreation RecreationBench | 49.9 | 47.1 | 30.2 | — |
| Mobile use AndroidWorld | 84.5 | 81.9 | 81.0 | 62.0 |
| Computer use OSWorld 2.0 | Binary 19.4 Partial 52.3 | Binary 19.4 Partial 48.0 | Binary 2.8 Partial 21.5 | — |
| Visual web development Vision2Web | 64.0 | 62.9 | 42.1 | — |
| General Multimodal Intelligence | ||||
| Embodied intelligence ERQA | 72.3 | 65.5 | 69.8 | 40.8 |
| Long video understanding LVBench | 76.6 | 72.4 | 76.2 | 63.0 |
| Real-world perception RealWorldQA | 88.5 | 85.9 | 86.9 | 73.9 |
| Visual math problem solving MathVision | Without CI 90.6 With CI 95.7 | Without CI 90.0 With CI 94.6 | Without CI 90.3 With CI 88.7 | Without CI 65.5 |
| Scientific chart analysis CharXiv (RQ) | Without CI 84.6 With CI 90.6 | Without CI 83.7 With CI 90.2 | Without CI 85.8 With CI 85.9 | Without CI 66.0 |
In multimodal domain measurement results as well, it records Binary 19.4 and Partial 52.3 on OSWorld 2.0, which evaluates whether tasks can be completed by operating actual OS screens, significantly increasing its numbers from the past model Qwen3.7-Plus (Binary 2.8, Partial 21.5). It also achieves 84.5 on AndroidWorld, greatly exceeding Claude-Opus-4.6 (Max)’s 62.0. However, looking at the Average score for ClawEval-MM, it is 60.4, with the difference from Qwen3.7-Plus’s 60.1 remaining very minor at just 0.3.
Strengths and Use Cases
Qwen3.8-Flash-Next is designed with a focus on agents performing long-duration autonomous tasks, software engineering, and multimodal tasks involving screen operations.
Specific strengths include the following areas:
- Software Development and Repository Editing (Agentic Coding): It records high scores on SWE-bench Pro and SWE-bench Multilingual, which solve real repository issues, showing high aptitude for agent tasks such as debugging and adding features that look across the entire project, rather than stopping at single-function generation.
- Long-Horizon Office Work and Practical Tasks (Long-Horizon Agents): It yields excellent results on CoWorkBench and JobBench, which evaluate office work and specialized tasks requiring lengthy steps, making it suitable for workflows that repeatedly plan and execute while autonomously calling tools.
- GUI Operations and Multimodal Support: It integrates a vision encoder and supports image and video input. As demonstrated by benchmark results such as OSWorld 2.0, AndroidWorld, and Vision2Web, it exhibits strengths in agents that visually recognize and operate desktop and mobile device screens, as well as image-based web development.
- Ultra-Long Context Processing and Preservation of Reasoning Processes: It supports a standard context length of 262,144 tokens and up to 1,000,000 tokens through expansion. Additionally, the thinking mode is enabled by default, equipped with a “Preserved Thinking" feature that retains the trajectory of reasoning across the entire past conversation history. This makes it possible to maintain consistency in judgment across multi-turn interactions and eliminate wasteful re-reasoning.
In addition, in the community project “Strata", variations such as “Coder", a derivative version tailored for code generation with a narrowed-down number of experts, and “Swift 1.5", a fine-tuned version that shortens thinking time to return quick responses, are also provided, envisioning tailored use cases such as programming-dedicated assistants and response-speed-prioritized chats.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 180.0B parameters
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| More than 402GB of VRAM (multi-GPU or CPU offload required) | BF16 | 335.3GB | 402.3GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-05): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): not registered. “Not registered" means the name is absent from that registry today, not that the model cannot run.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
The publisher distributes this model as safetensors.
License — qwen-community-1.0: A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.
Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Recent Models in the Same Size Class
Models with over 40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest VRAM tier | License | Our article |
|---|---|---|---|---|
| ggml-org/GLM-5.3-Flash-GGUF | 321.3B | — | other | GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory (2026-10-03) |
| bartowski/Intern-S2-397B-GGUF | 403.4B | — | apache-2.0 | Intern-S2-397B-GGUF Vision-Language Model: ~102GB Memory (2026-09-15) |
| deepseek-ai/DeepSeek-V4.1-Flash | 763.2B | — | mit | DeepSeek-V4.1-Flash 552B Multimodal MoE Model: ~570GB Memory (2026-09-10) |
How to Get It
The model weights for Qwen3.8-Flash-Next are published on Hugging Face in safetensors format. The “qwen-community-1.0" license applies, and it is distributed as an open-weight model that does not require specific advance applications.
Multiple methods are provided to use and run this model.
For server environments and high-performance execution environments, the official model card provides instructions on how to use the following inference frameworks:
- SGLang
- vLLM
- TokenSpeed
Additionally, as an inference project designed to run on typical consumer PC environments, the open-source software “Strata", developed under the MIT license, has been released. It supports Windows and Linux environments, and after downloading the repository, running the following scripts allows setup and execution to proceed interactively:
- For Windows: Run
START-HERE.bat. - For Linux: Run
./setup.sh.
The installer detects hardware, selects an appropriate quantized model, and automatically handles everything from downloading to startup. If using AI coding assistants such as Cursor or Claude Code, it also features functionality to have the assistant delegate tasks from configuration checking to startup by passing the instruction prompts prepared in the repository.
After startup, you can access the Web UI (http://127.0.0.1:8080) from your browser to chat or check operation status, and you can also use it from various clients and development tools via the following compatible endpoints:
- OpenAI-compatible API: Specify
http://127.0.0.1:8080/v1for the base URL. - Anthropic-compatible API: Specify
http://127.0.0.1:8080/v1/messages.
Note that during the initial startup of the model, loading into main memory and allocation to the GPU take place, so depending on the environment, responses may appear to stop for a few minutes, but this is explained as normal background processing.
Related Articles
- Qwen3.8-27B Vision-Language Model: Our Test Answers, 8GB+ VRAM
- Qwen-Image-2.1 Released: Open-Weight Image Gen & Editing
- DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds
- K2-Horizon-32B-NVFP4: Our Test Answers, 32GB+ VRAM
What to Read Next
- Find models by VRAM (This model does not fit even in 80GB; its smallest build needs about 402GB) → VRAM quick reference, including models over 80GB
- Engines that run this model → vLLM
- Formats this model is available in → Safetensors format guide and models
- Learn about the publisher → Alibaba (Qwen): models, licenses and articles
- Other models for the same task → Other vision-language models

