MiMo-V2.6 Guide: VRAM Requirements, GGUF Builds
Everything Local Model Watch has published about the MiMo-V2.6 family: 2 article(s) covering the base model and its fine-tunes, plus converted builds we tracked after publication. Memory requirements below are computed by this site from file sizes, not quoted from model cards. Part of our model family index.
At a Glance
| Item | Value |
|---|---|
| Base model(s) | XiaomiMiMo/MiMo-V2.6-Pro-RL, XiaomiMiMo/MiMo-V2.6-Flash-RL |
| Publisher | Xiaomi (MiMo) |
| License (model card) | mit |
| Articles | 2 |
Hardware Requirements
Estimated requirements (calculated by Local Model Watch)
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| More than 141GB of VRAM (multi-GPU or CPU offload required) | Q2_K | 117.5GB | 141.1GB |
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Figures are for ggml-org/MiMo-V2.6-Flash-RL-GGUF, the most recent model in this family with a hardware estimate.
Can You Run It Locally?
Runs in Ollama, LM Studio and llama.cpp as-is.
It is distributed in GGUF, so no conversion is needed.
License — mit (Commercial use allowed): Permits commercial use, modification and redistribution, provided the copyright notice and license text are retained.
Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.
This assessment is for ggml-org/MiMo-V2.6-Flash-RL-GGUF.
Quantized and Converted Variants
→ Scroll horizontally to see all columns
| Added | Publisher | Format | Repository | Smallest VRAM tier (build, est. memory) |
|---|---|---|---|---|
| 2026-09-27 | mlx-community | MXFP4 | mlx-community/MiMo-V2.6-Pro-RL-mxfp4-q8 | MLX 4bit 619.0GB (does not fit a single consumer GPU) |
File sizes of each build:
- Available builds in mlx-community/MiMo-V2.6-Pro-RL-mxfp4-q8: MLX 4bit 515.8GB
This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files.
Articles (the family’s own models first, then newest)
| Published | Model | Type | Article |
|---|---|---|---|
| 2026-09-27 | XiaomiMiMo/MiMo-V2.6-Pro-RL | New Models | Xiaomi Releases MiMo-V2.6-Pro-RL: 1.02T MoE Flagship |
| 2026-09-22 | ggml-org/MiMo-V2.6-Flash-RL-GGUF | New Models | MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory |
Repositories
Last updated 2026-09-27 (JST). Assembled by code from our article log; no text on this page is written by an AI model.