Xing4.0-29B-A4B 29B MoE Model Strong in Coding Agents: 24GB+ VRAM

Xing4.0-29B-A4B 29B MoE Model Strong in Coding Agents: 24GB+ VRAM

At a Glance

Item Value
Repository XingChen-AGI/Xing4.0-29B-A4B
Family guide Xing4.0 guide (1 articles)
Publisher guide China Telecom (Xing, TeleChat): models and licenses
Published 2026-09-16
License apache-2.0
Formats safetensors
Paper arXiv:2512.24157, arXiv:2507.18013
Source type Primary source (the publisher itself)

Values determined by this site’s code at collection time. Dates are JST.

Overview

China Telecom’s AI subsidiary, China Telecom Artificial Intelligence Technology, has released a new large language model, Xing4.0-29B-A4B, on Hugging Face. It is a Mixture of Experts (MoE) model with a total of 29B parameters, of which only 4B are active per token. It handles a context length of 256K tokens by default, which can be extended to 512K tokens. Licensed under Apache-2.0, it is freely available including for commercial use.

Xing (Star) is the successor to models previously published by the company under the “TeleChat" name. The most significant feature of this release lies in its training infrastructure. According to the model card, Xing4.0-29B-A4B is the first model of this scale trained entirely on Huawei Ascend NPUs (Ascend 910C cluster) and the MindSpore framework. This means it achieved results matching leading open models of the same class on coding agent benchmarks without using NVIDIA GPUs.

The publisher has provided FP8 and GGUF versions directly alongside the weights (XingChen-AGI/Xing4.0-29B-A4B-FP8 and XingChen-AGI/Xing4.0-29B-A4B-GGUF). Because official GGUF files are available, it is easy to test using llama.cpp-based tools.

Specifications

Based on the model card configuration table:

  • Parameter Count: Total 29B / Active 4B (MoE)
  • Number of Layers: 40
  • Hidden Dimension: 3,584
  • FFN Intermediate Dimension: 9,216 for dense FFN, 1,024 per expert
  • Experts: Uses 4 out of 64 routed experts, plus 1 always-on shared expert
  • Attention: MLA (Multi-head Latent Attention)
  • Context Length: 256K tokens (extendable up to 512K)
  • Architectural Elements: The model card lists “mHC + MLA + MTP". MLA compresses the KV cache to reduce memory consumption in long contexts, while MTP (Multi-Token Prediction) simultaneously predicts multiple subsequent tokens and is also used to accelerate speculative decoding. The model card provides no formal name or explanation for mHC.
  • Thinking Mode: The enable_thinking chat template parameter allows toggling whether to output the reasoning process.

Figures are also provided for training efficiency. By combining fine-grained MoE communication optimizations, selective recomputation, automatic operation fusion, and dedicated operators for mHC (Ascend C), training throughput on Ascend was reportedly increased by approximately 96% compared to unoptimized setups.

Performance

Transcribed directly from the model card’s evaluation table. The comparison targets are Gemma4-26B-A4B and Qwen3.6-35B-A3B, which have roughly the same scale of active parameters. All figures were measured by the publishers themselves.

Benchmark Xing4.0-29B-A4B Gemma4-26B-A4B Qwen3.6-35B-A3B
IFBench 69.67 72.67 65.50
AIME2026 90.00 88.30 92.70
AA.LCR 61.00 66.00 62.00
Tau3-Bench 64.63 58.90 67.20
Claw-Eval 76.55 71.49 74.54
SWE-bench Verified 75.00 53.00 76.00
Terminal-Bench 2.1 57.50 30.00 51.50
SWE-bench Multilingual 66.00 51.00 67.20
DeepresearchBII 60.80 39.30 59.70

Its strengths lie in agentic tasks that utilize tools and follow structured steps. Terminal-Bench 2.1, which tests the completion of tasks in a terminal, scores 57.50—outperforming Qwen3.6-35B-A3B (51.50) by 6 points and Gemma4-26B-A4B (30.00) by 27.5 points, making it the highest in the table. The overall agent evaluation Claw-Eval (76.55) and DeepresearchBII (60.80), which involves conducting extensive research to produce reports, are also the highest among the three models. SWE-bench Verified, which fixes bugs in real repositories, scores 75.00, essentially tying with Qwen3.6-35B-A3B at 76.00 and beating Gemma4-26B-A4B (53.00) by a 22-point margin.

On the other hand, it scores lowest in some categories. AA.LCR, which evaluates reading comprehension and reasoning over long documents, scores 61.00, falling short of both Gemma4-26B-A4B (66.00) and Qwen3.6-35B-A3B (62.00). Being able to handle a long context of 256K to 512K is distinct from accurately comprehending long contexts, and based on this table, the latter is weaker than the other two models in its class. Mathematics on AIME2026 (90.00) trails Qwen3.6-35B-A3B (92.70), and instruction-following on IFBench (69.67) falls behind Gemma4-26B-A4B (72.67). Tau3-Bench (64.63), which measures tool-using dialogues like customer service, is also lower than Qwen3.6-35B-A3B (67.20).

In summary, as a coding agent, it is equal to or better than Qwen3.6-35B-A3B and clearly superior to Gemma4-26B-A4B, while performing roughly on par with or slightly below the other two models in mathematics, long-text reading, and instruction following. Active parameters for all three models fall in the same 3B–4B bracket, providing a similar speed feel when run locally.

Evaluation conditions require caution. According to the model card footnotes, SWE-bench used the SWE-agent framework with a 210K token context, while Terminal-Bench 2.1 averaged 3 runs using the terminus-2 framework. It is not specified whether the two comparison models were measured under identical conditions.

Strengths and Use Cases

The model card envisions its use as the core of coding agents and general-purpose agents. It is designed to maintain multi-step planning, tool invocation, and long chains of reasoning even across extended contexts. The model card explicitly states that it has been aligned with agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes, alongside output format integration.

Another intended use case is continual pre-training on proprietary data. It claims to enable low-cost adaptation for domain-specific applications such as intent classification, table understanding, contract review, and knowledge base Q&A. For continued training, it supports LLaMA-Factory and MindFormers, and deployment across multiple chips (including non-GPU hardware) is supported via BAAI’s FlagOS.

Recommended sampling settings are also provided by use case:

Use Case temperature top_p repetition_penalty
Complex Reasoning / General Tasks 1.0 0.95 1.05
Coding / Agents 0.8 0.95 1.05

How It Differs from Similar Models

  • Differences from Gemma4-26B-A4B: While active parameters are identically 4B, Xing outperforms Gemma by 15–27 points on coding agent benchmarks (SWE-bench Verified, Terminal-Bench 2.1, SWE-bench Multilingual). Conversely, Gemma scores higher on instruction following (IFBench) and long-text reading (AA.LCR).
  • Differences from Qwen3.6-35B-A3B: Xing’s total parameter count is smaller by 6B. SWE-bench scores are nearly tied; Xing leads in Terminal-Bench 2.1, Claw-Eval, and DeepresearchBII, while Qwen leads in AIME2026, Tau3-Bench, and SWE-bench Multilingual. Strengths are divided, and neither model is universally superior.
  • Training Infrastructure Differences: While the comparison models were trained on GPUs, Xing4.0 was trained exclusively on Ascend NPUs. From a user’s perspective, inference procedures remain unchanged, but it is significant as a demonstration that models of this class can be built on non-GPU training infrastructure.

Released under Apache-2.0 with official FP8 and GGUF versions provided concurrently, and with explicit support for agent frameworks, this release reflects an emphasis on making the model immediately usable by external developers rather than just publishing weights.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 31.2B parameters

Your VRAM Quantization File size Est. memory needed
24GB (RTX 4090 / 3090, etc.) IQ4_NL 18.7GB 22.5GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-09-28): llama.cpp: not registered, vLLM: not registered, MLX (mlx-lm): not registered. “Not registered" means the name is absent from that registry today, not that the model cannot run.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. File sizes are measured from the converted build XingChen-AGI/Xing4.0-29B-A4B-GGUF. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

Runs in Ollama, LM Studio and llama.cpp via a converted build.

The publisher ships safetensors, but XingChen-AGI/Xing4.0-29B-A4B-GGUF provides a GGUF build you can use.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compression: the IQ4_NL build measures 5.15 bits per weight — about 32% the size of the original 16-bit weights, calculated by this site from the actual file sizes.

Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Our Own Measurements

Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.

Inside the GGUF File

File: xing4_0-29b-IQ4_NL.gguf (18.72GB, IQ4_NL). We read only the header (metadata) of the file, not the weights.

Item Value
Architecture (as named in the GGUF) xing4_0
Maximum trained context length 262,144 tokens
Layers 41
Experts 4 active out of 64
Vocabulary size 131,072
Chat template Included (mentions tool calls, has a thinking switch)
imatrix Used (calibration.txt, 8 chunks)
Weight types (share of parameters) IQ4_NL 93.1% / BF16 5.3% / Q6_K 1.5% / F32 0.0%
Average bits per weight 5.15 bits
Embedding / output layer type BF16 / Q6_K

The quant name in the file name describes the file as a whole; in practice layers mix several types. The average bits per weight is the measured file data divided by the number of weights.

Recent Models in the Same Size Class

Models with 15–40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.

→ Scroll horizontally to see all columns

Model Parameters Smallest VRAM tier License Our article
Altworld/Hemmingway-1 26.9B 12GB cc-by-nc-4.0 Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds (2026-09-28)
orcarouter/OrcaSAQ-2-27B 27.8B 16GB apache-2.0 OrcaSAQ-2-27B Text Generation Model: 16GB+ VRAM (2026-09-28)
prism-ml/Ternary-Bonsai-2-27B-gguf 27.8B 8GB apache-2.0 Ternary-Bonsai-2-27B-gguf Text Generation Model: 8GB+ VRAM (2026-09-18)
Edge0/Edge0-35B-A3B-preview 36.0B 24GB apache-2.0 Edge0-35B-A3B-preview 35B MoE Model for Phone-Class Memory: 24GB+ VRAM (2026-09-11)
bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF 26.5B 12GB apache-2.0 Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF: 12GB+ VRAM (2026-09-11)

How to Get It

  • Distribution Format: Transformers-format (safetensors) weights are available at XingChen-AGI/Xing4.0-29B-A4B on Hugging Face. Official FP8 and GGUF versions are hosted in separate repositories. It is also distributed on ModelScope and Modelers.
  • Example Download Command: huggingface-cli download XingChen-AGI/Xing4.0-29B-A4B
  • Supported Engines: The model card lists vLLM, SGLang, and KTransformers, with startup procedures provided in the GitHub repository. Note that inference examples using GitHub’s Transformers require trust_remote_code=True, executing the code bundled within the repository.
  • Running Locally: Using the official GGUF version allows testing with llama.cpp-based tools. Computational workload per token corresponds to 4 active parameters, making generation faster than dense models of the same total parameter count. However, sufficient memory must be available to load all 29B parameters. Refer to the table below for estimated memory requirements.

Quantized and Converted Variants

→ Scroll horizontally to see all columns

Added Publisher Format Repository Smallest VRAM tier (build, est. memory)
2026-09-28 XingChen-AGI GGUF (imatrix) XingChen-AGI/Xing4.0-29B-A4B-GGUF IQ4_NL 22.5GB (fits in 24GB VRAM)
2026-09-28 XingChen-AGI FP8 XingChen-AGI/Xing4.0-29B-A4B-FP8 FP8 37.1GB (fits in 48GB VRAM)
2026-09-28 mlx-community MLX mlx-community/Xing4.0-29B-A4B-OptiQ-4bit MLX 4bit 23.0GB (fits in 24GB VRAM)

File sizes of each build:

  • Available builds in XingChen-AGI/Xing4.0-29B-A4B-GGUF: IQ4_NL 18.7GB
  • Available builds in XingChen-AGI/Xing4.0-29B-A4B-FP8: FP8 30.9GB
  • Available builds in mlx-community/Xing4.0-29B-A4B-OptiQ-4bit: MLX 4bit 19.2GB

In addition, 10 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model’s publisher or established quantization maintainers.

This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: glossary.

Other Models for the Same Task

Recent text generation models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site’s estimate.

See all text generation models →

What to Read Next

Sources

Update History

  • 2026-09-28: Added converted builds to “Quantized and Converted Variants”: XingChen-AGI/Xing4.0-29B-A4B-GGUF, XingChen-AGI/Xing4.0-29B-A4B-FP8, mlx-community/Xing4.0-29B-A4B-OptiQ-4bit
  • 2026-09-28: Updated the hardware requirements table with the actual file sizes of XingChen-AGI/Xing4.0-29B-A4B-GGUF.
  • 2026-09-28: Changed the title to show what the article covers (VRAM requirements, file list, etc.).
  • 2026-09-28: Added our own measurements: what is inside the GGUF file.