GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory

October 5, 2026

GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory

At a Glance

Item Value
Repository ggml-org/GLM-5.3-Flash-GGUF
Publisher guide Z.ai (Zhipu AI): models and licenses
Published 2026-10-04
License other
Formats GGUF
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

What we checked ourselves

  • It needs about 7% more tokens than the Qwen3 tokenizer for the same Japanese text, so only about 0.94× as much Japanese fits in the same context length.

Details and conditions are in “Our Own Measurements” below.

Overview

ggml-org has released “ggml-org/GLM-5.3-Flash-GGUF", which is a GGUF quantized version of the open-weight multimodal model “GLM-5.3-Flash-BF16" published by zai-org.

The original model, GLM-5.3-Flash, is the first native multimodal model in the GLM-5 series. It adopts a Mixture of Experts (MoE) configuration with a total parameter count of 320B and 18B active parameters. It is designed to balance efficient long-context processing and high performance by incorporating a hybrid architecture combining sparse and linear attention, along with Manifold-Constrained Hyper-Connections (mHC).

This repository provides the model converted into the GGUF format for use in inference environments such as llama.cpp, including the mmproj for the vision encoder and the MTP sidecar for speculative sampling.

Specifications

  • Total parameters: 320B
  • Active parameters: 18B
  • Architecture: MoE (288 routed experts per layer, 1 shared expert), mHC + DSA, 1 MTP layer

Performance

Since this model is a converted version in GGUF format, performance evaluations are based on the public information of the original model “GLM-5.3-Flash-BF16". Note that slight variations from the original precision may occur due to quantization.

According to explanations from the original model’s publishers, it delivers performance surpassing the previous-generation GLM-5.2 across various benchmarks and real-world workloads. Furthermore, it is reported to achieve performance levels approaching Claude Opus 4.8 on coding-related and agent-related benchmarks.

The evaluation procedures for the original model cite the following benchmarks:

  • HLE w/ tools (full set)
  • NL2Repo
  • DeepSWE
  • Terminal-Bench 2.1
  • Agent’s Last Exam
  • Toolathlon Verified
  • AutomationBench
  • GDPval-AA v2
  • BabyVision

The original model features a parameter (reasoning_effort) to control the amount of reasoning during inference, allowing users to adjust inference capabilities by specifying one of three stages: low, high, or max (default).

Strengths and Use Cases

The base model “GLM-5.3-Flash-BF16" is the first in the GLM-5 series to support native multimodal processing. It supports image-text input and conversation processing (image-text-to-text), making it useful for general tasks involving visual information.

The main strengths and expected use cases derived from the original model’s design and evaluation details are as follows:

Coding and Autonomous Agent Support

The original model is optimized to demonstrate high capabilities in coding and agent-related benchmarks. Evaluations adopt “NL2Repo" for repository generation, “DeepSWE" for autonomously solving software engineering tasks, “Terminal-Bench 2.1" for evaluating terminal operations, “Toolathlon Verified" for measuring tool utilization capabilities, and “AutomationBench" for verifying automated workflows. This makes it suitable for complex agent use cases such as not only creating and editing code, but also executing development tools and automating autonomous workflows.

Efficient Long-Context Processing

Architecturally, a hybrid structure combining sparse and linear attention is introduced. This aims to significantly reduce inference overhead when handling long contexts while maintaining context-grasping capabilities. It is suited for scenarios where context lengths tend to grow large, such as analyzing massive codebases or handling agent tasks that include multi-step tool-calling histories.

Thinking Budget Control and Dialogue

The original model features a parameter (reasoning_effort) that controls the amount of reasoning during inference, allowing users to adjust inference cost and depth according to their use case. Additionally, the chat template provides a clear_thinking option to control the display of the thought process, and explicitly setting clear_thinking=true is recommended for general conversational use. Supported language tags include English (en) and Chinese (zh).

High-Speed Inference in Local Environments

In this GGUF version, in addition to the mmproj file for the vision encoder, an MTP (Multi-Token Prediction) sidecar that enables speculative decoding is included. This is expected to achieve faster text generation than standard inference when using a compatible inference engine.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 321.3B parameters (taken from the base model zai-org/GLM-5.3-Flash-BF16)

Your VRAM Quantization File size Est. memory needed
More than 150GB of VRAM (multi-GPU or CPU offload required) Q2_K 124.6GB 149.5GB

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

Runs in Ollama, LM Studio and llama.cpp as-is.

It is distributed in GGUF, so no conversion is needed.

License — other: A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.

Compression: the Q2_K build measures 3.33 bits per weight — about 21% the size of the original 16-bit weights, calculated by this site from the actual file sizes.

Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Our Own Measurements

Values we checked ourselves on our server (no GPU) by actually reading this model’s files, without running the model — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.

Measurements We Did Not Take

We have not confirmed that this site meets the commercial-use terms of this model’s license (other). Because this site carries advertising, we did not run the model (no CPU run, answers, quantization comparison, conversion or generation). Only values we checked without running the model, such as token counts and the GGUF header, are shown.

Japanese Token Efficiency

Tokenizer Tokens per 1,000 Japanese characters Ratio to the same text in English
This model 733 1.34×
Qwen3 688 1.26×
Llama 3.2 744 1.36×
Gemma 3 564 1.03×
gpt-oss 795 1.45×
LLM-jp-3 497 0.85×

It needs about 7% more tokens than the Qwen3 tokenizer for the same Japanese text, so only about 0.94× as much Japanese fits in the same context length.

Counted with the tokenizer.json of zai-org/GLM-5.3-Flash-BF16 on a fixed text we wrote ourselves (876 Japanese characters across news, conversation, technical docs, a formal email, travel writing and a recipe) and its English translation. Fewer tokens mean more Japanese fits in the context window.

Inside the GGUF File

File: GLM-5.3-Flash-Q2_K-00001-of-00002.gguf (124.62GB, Q2_K). We read only the header (metadata) of the file, not the weights.

Item Value
Architecture (as named in the GGUF) glm5-next
Maximum trained context length 1,048,576 tokens
Layers 45
Experts 8 active out of 288
Vocabulary size 154,880
Chat template Included (mentions tool calls, has a thinking switch)
imatrix Not recorded in the file
Weight types (share of parameters) Q2_K 64.8% / Q4_K 32.4% / Q8_0 2.7% / BF16 0.1% / other 0.0%
Average bits per weight 3.42 bits
Embedding / output layer type Q8_0 / Q8_0

The quant name in the file name describes the file as a whole; in practice layers mix several types. The average bits per weight is the measured file data divided by the number of weights.

Recent Models in the Same Size Class

Models with over 40B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.

→ Scroll horizontally to see all columns

Model Parameters Smallest VRAM tier License Our article
Qwen/Qwen3.8-Flash-Next 180.0B — qwen-community-1.0 Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory (2026-10-04)
bartowski/Intern-S2-397B-GGUF 403.4B — apache-2.0 Intern-S2-397B-GGUF Vision-Language Model: ~102GB Memory (2026-09-15)
deepseek-ai/DeepSeek-V4.1-Flash 763.2B — mit DeepSeek-V4.1-Flash 552B Multimodal MoE Model: ~570GB Memory (2026-09-10)

How to Get It

This model is distributed in GGUF format on the Hugging Face repository “ggml-org/GLM-5.3-Flash-GGUF". Since it is not a gated model, you can download and use it without prior application or consent procedures.

When using CLI tools from the official llama.cpp family, you can start a server by directly specifying the Hugging Face repository with the following command:

llama serve -hf ggml-org/GLM-5.3-Flash-GGUF

The internal specifications of the quantized versions provided in this repository have the following features:

  • Q4_K: All routed experts are quantized with Q4_K.
  • Q2_K: Composed of Q4_K for routed down experts and Q2_K for gate/up experts.
  • Smaller non-expert tensors, such as embedding layers, attention layers, shared experts, and dense FFNs, are kept as Q8_0 to maintain quality.
  • Bundled with mmproj (Q8_0) for the vision encoder and MTP sidecars (Q8_0 / Q4_0) for speculative sampling.
  • Note that low-bit quantization is noted as not having undergone calibration using imatrix (importance matrix).

Regarding the license classification, the base model “GLM-5.3-Flash-BF16" is set under the MIT license, but this GGUF repository carries an other tag. Please check the repository’s license terms in advance before use.

Related Articles

What to Read Next

Sources

Update History

  • 2026-10-05: Added our own measurements: Japanese token efficiency, what is inside the GGUF file.