Xing4.0-29B-A4B 29B MoE Model Strong in Coding Agents: 24GB+ VRAM

- 1. At a Glance
- 2. Overview
- 3. Specifications
- 4. Performance
- 5. Strengths and Use Cases
- 6. How It Differs from Similar Models
- 7. Hardware Requirements
- 8. Can You Run It Locally?
- 9. Our Own Measurements
- 10. Recent Models in the Same Size Class
- 11. How to Get It
- 12. Quantized and Converted Variants
- 13. Other Models for the Same Task
- 14. What to Read Next
- 15. Sources
- 16. Update History
At a Glance
| Item | Value |
|---|---|
| Repository | XingChen-AGI/Xing4.0-29B-A4B |
| Family guide | Xing4.0 guide (1 articles) |
| Publisher guide | China Telecom (Xing, TeleChat): models and licenses |
| Published | 2026-09-16 |
| License | apache-2.0 |
| Formats | safetensors |
| Paper | arXiv:2512.24157, arXiv:2507.18013 |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
China Telecom’s AI subsidiary, China Telecom Artificial Intelligence Technology, has released a new large language model, Xing4.0-29B-A4B, on Hugging Face. It is a Mixture of Experts (MoE) model with a total of 29B parameters, of which only 4B are active per token. It handles a context length of 256K tokens by default, which can be extended to 512K tokens. Licensed under Apache-2.0, it is freely available including for commercial use.
Xing (Star) is the successor to models previously published by the company under the “TeleChat" name. The most significant feature of this release lies in its training infrastructure. According to the model card, Xing4.0-29B-A4B is the first model of this scale trained entirely on Huawei Ascend NPUs (Ascend 910C cluster) and the MindSpore framework. This means it achieved results matching leading open models of the same class on coding agent benchmarks without using NVIDIA GPUs.
The publisher has provided FP8 and GGUF versions directly alongside the weights (XingChen-AGI/Xing4.0-29B-A4B-FP8 and XingChen-AGI/Xing4.0-29B-A4B-GGUF). Because official GGUF files are available, it is easy to test using llama.cpp-based tools.
Specifications
Based on the model card configuration table:
- Parameter Count: Total 29B / Active 4B (MoE)
- Number of Layers: 40
- Hidden Dimension: 3,584
- FFN Intermediate Dimension: 9,216 for dense FFN, 1,024 per expert
- Experts: Uses 4 out of 64 routed experts, plus 1 always-on shared expert
- Attention: MLA (Multi-head Latent Attention)
- Context Length: 256K tokens (extendable up to 512K)
- Architectural Elements: The model card lists “mHC + MLA + MTP". MLA compresses the KV cache to reduce memory consumption in long contexts, while MTP (Multi-Token Prediction) simultaneously predicts multiple subsequent tokens and is also used to accelerate speculative decoding. The model card provides no formal name or explanation for mHC.
- Thinking Mode: The
enable_thinkingchat template parameter allows toggling whether to output the reasoning process.
Figures are also provided for training efficiency. By combining fine-grained MoE communication optimizations, selective recomputation, automatic operation fusion, and dedicated operators for mHC (Ascend C), training throughput on Ascend was reportedly increased by approximately 96% compared to unoptimized setups.
Performance
Transcribed directly from the model card’s evaluation table. The comparison targets are Gemma4-26B-A4B and Qwen3.6-35B-A3B, which have roughly the same scale of active parameters. All figures were measured by the publishers themselves.
| Benchmark | Xing4.0-29B-A4B | Gemma4-26B-A4B | Qwen3.6-35B-A3B |
|---|---|---|---|
| IFBench | 69.67 | 72.67 | 65.50 |
| AIME2026 | 90.00 | 88.30 | 92.70 |
| AA.LCR | 61.00 | 66.00 | 62.00 |
| Tau3-Bench | 64.63 | 58.90 | 67.20 |
| Claw-Eval | 76.55 | 71.49 | 74.54 |
| SWE-bench Verified | 75.00 | 53.00 | 76.00 |
| Terminal-Bench 2.1 | 57.50 | 30.00 | 51.50 |
| SWE-bench Multilingual | 66.00 | 51.00 | 67.20 |
| DeepresearchBII | 60.80 | 39.30 | 59.70 |
Its strengths lie in agentic tasks that utilize tools and follow structured steps. Terminal-Bench 2.1, which tests the completion of tasks in a terminal, scores 57.50—outperforming Qwen3.6-35B-A3B (51.50) by 6 points and Gemma4-26B-A4B (30.00) by 27.5 points, making it the highest in the table. The overall agent evaluation Claw-Eval (76.55) and DeepresearchBII (60.80), which involves conducting extensive research to produce reports, are also the highest among the three models. SWE-bench Verified, which fixes bugs in real repositories, scores 75.00, essentially tying with Qwen3.6-35B-A3B at 76.00 and beating Gemma4-26B-A4B (53.00) by a 22-point margin.
On the other hand, it scores lowest in some categories. AA.LCR, which evaluates reading comprehension and reasoning over long documents, scores 61.00, falling short of both Gemma4-26B-A4B (66.00) and Qwen3.6-35B-A3B (62.00). Being able to handle a long context of 256K to 512K is distinct from accurately comprehending long contexts, and based on this table, the latter is weaker than the other two models in its class. Mathematics on AIME2026 (90.00) trails Qwen3.6-35B-A3B (92.70), and instruction-following on IFBench (69.67) falls behind Gemma4-26B-A4B (72.67). Tau3-Bench (64.63), which measures tool-using dialogues like customer service, is also lower than Qwen3.6-35B-A3B (67.20).
In summary, as a coding agent, it is equal to or better than Qwen3.6-35B-A3B and clearly superior to Gemma4-26B-A4B, while performing roughly on par with or slightly below the other two models in mathematics, long-text reading, and instruction following. Active parameters for all three models fall in the same 3B–4B bracket, providing a similar speed feel when run locally.
Evaluation conditions require caution. According to the model card footnotes, SWE-bench used the SWE-agent framework with a 210K token context, while Terminal-Bench 2.1 averaged 3 runs using the terminus-2 framework. It is not specified whether the two comparison models were measured under identical conditions.
Strengths and Use Cases
The model card envisions its use as the core of coding agents and general-purpose agents. It is designed to maintain multi-step planning, tool invocation, and long chains of reasoning even across extended contexts. The model card explicitly states that it has been aligned with agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes, alongside output format integration.
Another intended use case is continual pre-training on proprietary data. It claims to enable low-cost adaptation for domain-specific applications such as intent classification, table understanding, contract review, and knowledge base Q&A. For continued training, it supports LLaMA-Factory and MindFormers, and deployment across multiple chips (including non-GPU hardware) is supported via BAAI’s FlagOS.
Recommended sampling settings are also provided by use case:
| Use Case | temperature | top_p | repetition_penalty |
|---|---|---|---|
| Complex Reasoning / General Tasks | 1.0 | 0.95 | 1.05 |
| Coding / Agents | 0.8 | 0.95 | 1.05 |
How It Differs from Similar Models
- Differences from Gemma4-26B-A4B: While active parameters are identically 4B, Xing outperforms Gemma by 15–27 points on coding agent benchmarks (SWE-bench Verified, Terminal-Bench 2.1, SWE-bench Multilingual). Conversely, Gemma scores higher on instruction following (IFBench) and long-text reading (AA.LCR).
- Differences from Qwen3.6-35B-A3B: Xing’s total parameter count is smaller by 6B. SWE-bench scores are nearly tied; Xing leads in Terminal-Bench 2.1, Claw-Eval, and DeepresearchBII, while Qwen leads in AIME2026, Tau3-Bench, and SWE-bench Multilingual. Strengths are divided, and neither model is universally superior.
- Training Infrastructure Differences: While the comparison models were trained on GPUs, Xing4.0 was trained exclusively on Ascend NPUs. From a user’s perspective, inference procedures remain unchanged, but it is significant as a demonstration that models of this class can be built on non-GPU training infrastructure.
Released under Apache-2.0 with official FP8 and GGUF versions provided concurrently, and with explicit support for agent frameworks, this release reflects an emphasis on making the model immediately usable by external developers rather than just publishing weights.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 31.2B parameters
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 24GB (RTX 4090 / 3090, etc.) | IQ4_NL | 18.7GB | 22.5GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-09-28): llama.cpp: not registered, vLLM: not registered, MLX (mlx-lm): not registered. “Not registered" means the name is absent from that registry today, not that the model cannot run.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. File sizes are measured from the converted build XingChen-AGI/Xing4.0-29B-A4B-GGUF. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
Runs in Ollama, LM Studio and llama.cpp via a converted build.
The publisher ships safetensors, but XingChen-AGI/Xing4.0-29B-A4B-GGUF provides a GGUF build you can use.
License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.
Compression: the IQ4_NL build measures 5.15 bits per weight — about 32% the size of the original 16-bit weights, calculated by this site from the actual file sizes.
Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Our Own Measurements
Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.
Inside the GGUF File
File: xing4_0-29b-IQ4_NL.gguf (18.72GB, IQ4_NL). We read only the header (metadata) of the file, not the weights.
| Item | Value |
|---|---|
| Architecture (as named in the GGUF) | xing4_0 |
| Maximum trained context length | 262,144 tokens |
| Layers | 41 |
| Experts | 4 active out of 64 |
| Vocabulary size | 131,072 |
| Chat template | Included (mentions tool calls, has a thinking switch) |
| imatrix | Used (calibration.txt, 8 chunks) |
| Weight types (share of parameters) | IQ4_NL 93.1% / BF16 5.3% / Q6_K 1.5% / F32 0.0% |
| Average bits per weight | 5.15 bits |
| Embedding / output layer type | BF16 / Q6_K |
The quant name in the file name describes the file as a whole; in practice layers mix several types. The average bits per weight is the measured file data divided by the number of weights.
Recent Models in the Same Size Class
Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest VRAM tier | License | Our article |
|---|---|---|---|---|
| Altworld/Hemmingway-1 | 26.9B | 12GB | cc-by-nc-4.0 | Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds (2026-09-28) |
| orcarouter/OrcaSAQ-2-27B | 27.8B | 16GB | apache-2.0 | OrcaSAQ-2-27B Text Generation Model: 16GB+ VRAM (2026-09-28) |
| prism-ml/Ternary-Bonsai-2-27B-gguf | 27.8B | 8GB | apache-2.0 | Ternary-Bonsai-2-27B-gguf Text Generation Model: 8GB+ VRAM (2026-09-18) |
| Edge0/Edge0-35B-A3B-preview | 36.0B | 24GB | apache-2.0 | Edge0-35B-A3B-preview 35B MoE Model for Phone-Class Memory: 24GB+ VRAM (2026-09-11) |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | 26.5B | 12GB | apache-2.0 | Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF: 12GB+ VRAM (2026-09-11) |
How to Get It
- Distribution Format: Transformers-format (
safetensors) weights are available atXingChen-AGI/Xing4.0-29B-A4Bon Hugging Face. Official FP8 and GGUF versions are hosted in separate repositories. It is also distributed on ModelScope and Modelers. - Example Download Command:
huggingface-cli download XingChen-AGI/Xing4.0-29B-A4B - Supported Engines: The model card lists vLLM, SGLang, and KTransformers, with startup procedures provided in the GitHub repository. Note that inference examples using GitHub’s Transformers require
trust_remote_code=True, executing the code bundled within the repository. - Running Locally: Using the official GGUF version allows testing with llama.cpp-based tools. Computational workload per token corresponds to 4 active parameters, making generation faster than dense models of the same total parameter count. However, sufficient memory must be available to load all 29B parameters. Refer to the table below for estimated memory requirements.
Quantized and Converted Variants
→ Scroll horizontally to see all columns
| Added | Publisher | Format | Repository | Smallest VRAM tier (build, est. memory) |
|---|---|---|---|---|
| 2026-09-28 | XingChen-AGI | GGUF (imatrix) | XingChen-AGI/Xing4.0-29B-A4B-GGUF | IQ4_NL 22.5GB (fits in 24GB VRAM) |
| 2026-09-28 | XingChen-AGI | FP8 | XingChen-AGI/Xing4.0-29B-A4B-FP8 | FP8 37.1GB (fits in 48GB VRAM) |
| 2026-09-28 | mlx-community | MLX | mlx-community/Xing4.0-29B-A4B-OptiQ-4bit | MLX 4bit 23.0GB (fits in 24GB VRAM) |
File sizes of each build:
- Available builds in XingChen-AGI/Xing4.0-29B-A4B-GGUF: IQ4_NL 18.7GB
- Available builds in XingChen-AGI/Xing4.0-29B-A4B-FP8: FP8 30.9GB
- Available builds in mlx-community/Xing4.0-29B-A4B-OptiQ-4bit: MLX 4bit 19.2GB
In addition, 10 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model’s publisher or established quantization maintainers.
This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: glossary.
Other Models for the Same Task
Recent text generation models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site’s estimate.
- Xiaomi Releases MiMo-V2.6-Flash-MOPD and Pro-MOPD Models
- Xiaomi Releases MiMo-V2.6-Pro-RL: 1.02T MoE Flagship
- DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds (1650.5B, Over 80GB)
- Tev1-0.8B-experimental Text Generation Model: 4GB+ VRAM (873M, 4GB)
See all text generation models →
What to Read Next
- Find models by VRAM (This model runs from the 24GB tier) → Other models that run on a 24GB GPU
- Explore the same model family → Xing4.0 family overview (1 articles, 3 converted builds)
- Engines that run this model → llama.cpp / Ollama
- What IQ4_NL mean and where to get this model → GGUF format guide and models / MLX format guide and models / FP8 format guide and models / Quantization and model-format glossary
- Learn about the publisher → China Telecom (Xing, TeleChat): models, licenses and articles
Sources
Update History
- 2026-09-28: Added converted builds to “Quantized and Converted Variants”: XingChen-AGI/Xing4.0-29B-A4B-GGUF, XingChen-AGI/Xing4.0-29B-A4B-FP8, mlx-community/Xing4.0-29B-A4B-OptiQ-4bit
- 2026-09-28: Updated the hardware requirements table with the actual file sizes of XingChen-AGI/Xing4.0-29B-A4B-GGUF.
- 2026-09-28: Changed the title to show what the article covers (VRAM requirements, file list, etc.).
- 2026-09-28: Added our own measurements: what is inside the GGUF file.

