clef-flash Structured Output Optimized Model: 4GB+ VRAM, GGUF Builds

- 1. At a Glance
- 2. Specifications
- 3. Performance
- 4. Strengths and Use Cases
- 5. How It Differs from Similar Models
- 6. Hardware Requirements
- 7. Can You Run It Locally?
- 8. Our Own Measurements
- 9. How to Get It
- 10. Quantized and Converted Variants
- 11. Related Articles
- 12. What to Read Next
- 13. Sources
- 14. Update History
At a Glance
| Item | Value |
|---|---|
| Repository | Cloudflare/clef-flash |
| Publisher guide | Alibaba (Qwen): models and licenses |
| Published | 2026-10-01 |
| License | apache-2.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
What we checked ourselves
- Gave the 20 decision questions we give every decision model to this model, running the publisher’s own code on a cloud GPU (NVIDIA L4) with its network cut off: 18 correct in Japanese and 18 in English.
Details and conditions are in “Our Own Measurements” below.
Specifications
- Parameters: 9B
- Architecture: Qwen3.5-9B (Gated DeltaNet + Mixture of Experts: MoE) with a vision encoder and a custom Joint schema head
- Context Length: Native 262,144 tokens (can be extended up to 1,010,000 tokens using techniques such as YaRN)
Performance
Below are internal measurement results conducted by the publisher using the “Decision Index 0.2.1" suite. For comparison, figures for the higher-tier model Clef, as well as Jev and Kev 9B, are excerpted and listed.
Decision Index Measurement Results
→ Scroll horizontally to see all columns
| Benchmark | Clef | Clef-flash | Jev | Kev 9B |
|---|---|---|---|---|
| BFCL (case exact accuracy) | 98.5 | 98.8 | 95.8 | 94.5 |
| ToolRet (nDCG@10) | 69.2 | 66.4 | 65.3 | 64.3 |
| API-Bank (accuracy) | 91.9 | 93.1 | 88.2 | 56.3 |
| BANKING77 (macro-F1) | 94.2 | 90.9 | 79.7 | 84.8 |
| CLINC150+OOS (macro-F1) | 97.4 | 66.8 | 89.3 | 79.0 |
| RouterBench (selected quality) | 79.7 | 79.9 | 79.9 | 80.0 |
| Home appliance simulator (case exact accuracy) | 83.0 | 97.7 | 52.3 | 25.0 |
| SGD/SGD-X (macro-F1) | 43.8 | 34.2 | 43.0 | 64.0 |
| ContractNLI (macro-F1) | 81.4 | 84.3 | 71.7 | 57.8 |
| ANLI (macro-F1) | 69.8 | 59.1 | 74.8 | 56.3 |
| BPoMP (accuracy) | 96.9 | 95.4 | 90.6 | 67.0 |
| Humicroedit (accuracy) | 66.7 | 75.1 | 61.9 | 55.8 |
| POP909-CL (accuracy) | 15.8 | 1.6 | 18.1 | 10.8 |
| cfcolor (accuracy) | 66.0 | 65.8 | 64.7 | 56.3 |
| MMLU (accuracy) | 90.3 | 91.8 | 91.7 | 75.3 |
| GPQA Diamond (accuracy) | 48.0 | 51.0 | 78.3 | 38.8 |
| ARC-Easy (accuracy) | 99.0 | 99.5 | 99.3 | 97.7 |
| ARC-Challenge (accuracy) | 97.7 | 98.3 | 97.8 | 93.7 |
| WinoGrande (accuracy) | 93.5 | 97.5 | 92.0 | 73.2 |
| HellaSwag (accuracy) | 98.2 | 98.6 | 94.5 | 81.9 |
| GSM8K (accuracy) | 80.8 | 67.3 | 79.9 | 48.7 |
| ChessBench (accuracy) | 24.7 | 23.0 | 17.2 | 11.2 |
| MuSR (accuracy) | 83.5 | 86.0 | 66.1 | 57.9 |
| SATA-Bench (case exact accuracy) | 33.8 | 36.7 | 26.4 | 26.7 |
| BRIGHT (nDCG@10) | 45.9 | 39.3 | 47.5 | 38.5 |
| Amazon ESCI (macro-F1) | 57.5 | 57.4 | 55.2 | 49.2 |
| ACOS (per-review F1) | 33.3 | 25.9 | 29.5 | 18.3 |
| FinEntity (macro-F1) | 96.2 | 97.1 | 87.0 | 88.4 |
| VAST (macro-F1) | 59.5 | 49.6 | 64.6 | 55.4 |
| NLI4CT (macro-F1) | 82.9 | 78.6 | 84.1 | 74.9 |
| CRUXEval (accuracy) | 86.7 | 86.1 | 73.0 | 51.2 |
| CLadder (accuracy) | 94.0 | 97.7 | 72.6 | 62.0 |
| ForecastBench (Brier, lower is better) | 13.9 | 10.6 | 17.4 | 17.6 |
| Habermas Machine (accuracy) | 68.7 | 71.8 | 45.9 | 39.4 |
| PhishNChips (accuracy) | 79.6 | 75.0 | 62.5 | 50.7 |
| MMLU-Pro (accuracy) | 65.9 | 65.3 | 82.7 | 51.1 |
| BBH (accuracy) | 73.7 | 68.9 | 92.9 | 65.2 |
| RAGTruth (hallucination F1) | 79.4 | 35.6 | 76.5 | 46.2 |
| HoVer (accuracy) | 65.2 | 61.2 | 72.9 | 58.8 |
| When2Call MCQ (accuracy) | 72.4 | 65.6 | 81.0 | 49.6 |
| New Yorker (accuracy) | 69.5 | 66.1 | 70.1 | 58.1 |
| Median latency (ms) | 209.3 | 38.8 | 524.1 | 51.4 |
| p95 latency (ms) | 238.6 | 122.4 | 536.0 | 187.9 |
This table shows that Clef-Flash achieves an extremely high accuracy of 98.8% on BFCL, which measures tool-use capability as an agent, slightly outperforming the higher-tier Clef model. It also maintains a high standard of 91.8% on MMLU for general knowledge. Notably, it has low latency, with a p95 latency of 122.4ms, making it significantly faster than Clef (238.6ms) and Jev (536.0ms).
On the other hand, its RAGTruth score (measuring hallucinations) remains at 35.6%, which is significantly lower compared to Clef (79.4%) and Jev (76.5%). Additionally, in metrics such as intent classification (CLINC150+OOS), BBH, and MMLU-Pro which require reasoning, there are areas where it falls short of other models like Jev. While it excels in decision-making speed and specific tool-use accuracy, there are trade-offs regarding complex reasoning and hallucination suppression.
Workflow Evaluation
Accuracy measurement results from “Typesafe Evals", simulating actual business workflows, are as follows:
→ Scroll horizontally to see all columns
| Workflow | Metric | Clef | Clef-flash | Jev |
|---|---|---|---|---|
| Invoice processing | Exact actions | 64.7 | 57.1 | 61.8 |
| Invoice processing | Primary action | 86.2 | 73.3 | 83.1 |
| Customer service | Exact actions | 76.3 | 77.0 | 76.0 |
| Security incidents | Exact actions | 62.9 | 61.7 | 61.7 |
| Agent trace observability | Primary action | 68.5 | 69.8 | 71.6 |
In practical tasks as well, it records 77.0% in exact action selection for Customer service, showing performance that slightly surpasses the higher-tier model. Although it trails higher-tier models in complex tasks like Invoice processing, it maintains practical accuracy.
Furthermore, according to measurements by the publisher of the base model Qwen3.5-9B, it possesses high vision-language processing capabilities, such as 78.4% on MMMU for image understanding and 85.7% on MathVista for solving math problems involving figures and graphs. Clef-Flash inherits these powerful vision and language capabilities while being optimized for structured decision-making outputs.
Strengths and Use Cases
Clef-Flash excels not at standard chat generating free-form text, but at structured decision-making where passing a “state" and a “multiple-choice question schema" returns the probabilities for all choices of each question in a single inference. As indicated by tags such as structured-output, classification, and image-text-to-typed-output, its greatest feature is that output parsing is unnecessary, making it suitable for business processes like classification, routing, and intent determination. Since inputs accept images and video frames in addition to strings and JSON, it also supports use cases such as reading and evaluating invoice images.
According to the model card, actual measurements on the Decision Index show high values such as 98.8 in BFCL (tool-calling accuracy) and 97.7 in the Home appliance simulator (case exact accuracy). In business workflow evaluations, it also outperforms the higher-tier Clef and Jev models with 77.0 in customer service action selection. In addition, its p95 latency is significantly lower than other models at 122.4ms, making it suitable for real-time routing and classification processing where speed is required. Since the base Qwen/Qwen3.5-9B is a model equipped with multimodal capabilities, support for 201 languages, and long-context processing up to 1,010,000 tokens, it is believed to inherit those foundations even when handling multilingual documents or long state descriptions.
How It Differs from Similar Models
As related models built upon the same Qwen3.5-9B, our site has introduced “LensVLM-9B", a vision-language model that compresses long text into images for processing and “MiMo-V2.6-Distill-Qwen-9B-GGUF". LensVLM-9B is a model tuned by Apple in the direction of compressing long documents into images to reduce token count, while MiMo-V2.6-Distill-Qwen-9B is distilled and SFT-trained by Xiaomi MiMo for coding and general-purpose agent tasks, with both assuming free-form text generation. In contrast, Clef-Flash significantly differs by eliminating free-form text generation itself and specializing in schema-aligned probability output, making it a derivative with a completely different directional use case despite sharing the same base model.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 9.4B parameters
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 4GB (laptop iGPU / phone class) | IQ2_M | 3.3GB | 4.0GB |
| 8GB (RTX 4060 / 3060 Ti, etc.) | Q5_K_M | 6.4GB | 7.7GB |
| 12GB (RTX 4070 / 3060 12GB, etc.) | Q8_0 | 8.9GB | 10.7GB |
| 24GB (RTX 4090 / 3090, etc.) | BF16 | 16.7GB | 20.0GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. File sizes are measured from the converted build bartowski/Cloudflare_clef-flash-GGUF. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
Runs in Ollama, LM Studio and llama.cpp via a converted build.
The publisher ships safetensors, but bartowski/Cloudflare_clef-flash-GGUF provides a GGUF build you can use.
License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.
Compression: the IQ2_M build measures 3.01 bits per weight — about 19% the size of the original 16-bit weights, calculated by this site from the actual file sizes.
Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Our Own Measurements
Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.
Decision-Model Evaluation on Our Own Tasks
We gave this model the same 11 situations with 20 questions that we give every decision model (yes/no, named options and ordered levels; written by us, in Japanese and in English). llama.cpp cannot run this model’s decision head, so we ran the publisher’s own code (joint_schema_model.py and its SystemOne-compatible function) on a rented cloud GPU, inside a container with its network cut off. Each question states the rule to apply so that it has a single correct answer; our code decides right or wrong (yes when the probability of yes is 0.5 or more; for options and levels, the most probable one). These are not benchmark questions, and the result is not a general score of the model.
Conditions: Modal, NVIDIA L4 (peak VRAM 17.9GB), BF16, transformers 5.18.0, torch 2.14.1+cu130, publisher revision 17f0b0ad64. The time per decision is measured on that GPU from passing one situation (1 to 3 questions) to the publisher’s function until it returns; it is not a guide to the speed of your own GPU and cannot be compared with times measured on our CPU.
→ Scroll horizontally to see all columns
| Model | Method | Quant | Correct (Japanese) | Correct (English) | Mean probability on the correct answer (ja / en) | Time per decision, median (ja / en) | Peak memory |
|---|---|---|---|---|---|---|---|
| clef-flash | joint_schema_model.py |
BF16 |
18/20 | 18/20 | 0.88 / 0.88 | 170ms / 167ms (GPU: NVIDIA L4) | VRAM 17.9GB |
| Kev-4B (reference) | kev |
Q4_K_M |
18/20 | 19/20 | 0.87 / 0.87 | 7,616ms / 7,409ms | 6.3GB |
| clef | joint_schema_model.py |
BF16 |
20/20 | 19/20 | 0.98 / 0.94 | 190ms / 189ms (GPU: NVIDIA A100 80GB PCIe) | VRAM 51.4GB |
| lev (reference) | lev |
Q4_K_M |
16/20 | 16/20 | 0.76 / 0.80 | 14,061ms / 13,186ms | 6.0GB |
| Laya (reference) | laya |
Q8_0 |
10/20 | 15/20 | 0.47 / 0.68 | 968ms / 678ms | 0.9GB |
| Julia-1 (reference) | laya |
Q8_0 |
13/20 | 10/20 | 0.60 / 0.45 | 136ms / 107ms | 0.5GB |
It got the same number right in Japanese and English (18 of 20).
Other rows are other decision models measured the same way (“reference” rows are models we measure for comparison). Methods differ by model, and each model is measured with its own quantization.
Times marked (GPU) were measured on a cloud GPU and the others on our CPU, so the times cannot be compared across those rows; the numbers of correct answers can.
Answer to Each Question
→ Scroll horizontally to see all columns
| Situation | Question | Correct answer | Japanese | English |
|---|---|---|---|---|
| Routing a support message | Which team should handle this message? | billing |
✓ billing (p=0.98) |
✓ billing (p=0.96) |
| Routing a support message | Is the customer asking for money back? | yes | ✓ yes (p=0.95) | ✓ yes (p=0.80) |
| Routing a support message | Does the message report that a service is down? | no | ✓ no (p=1.00) | ✓ no (p=0.99) |
| Return eligibility (within the window) | Is today within 30 days of the delivery date? | yes | ✓ yes (p=0.96) | ✓ yes (p=0.95) |
| Return eligibility (within the window) | Can this item be returned under the policy? | yes | ✗ no (p=0.14) | ✗ no (p=0.17) |
| Return eligibility (past the window) | Is today within 30 days of the delivery date? | no | ✓ no (p=0.98) | ✓ no (p=0.99) |
| Return eligibility (past the window) | Can this item be returned under the policy? | no | ✓ no (p=0.96) | ✓ no (p=0.95) |
| Invoice handling (vendor) | What should happen to this invoice? | reject |
✓ reject (p=0.97) |
✓ reject (p=0.98) |
| Invoice handling (vendor) | Is the invoice total above 1,000 USD? | no | ✓ no (p=1.00) | ✓ no (p=1.00) |
| Invoice handling (amount) | What should happen to this invoice? | manager |
✓ manager (p=0.94) |
✓ manager (p=0.96) |
| Invoice handling (amount) | Is the invoice total above 1,000 USD? | yes | ✓ yes (p=0.98) | ✓ yes (p=0.98) |
| Incident severity | How widespread is the impact of this incident? | 3: All users affected | ✓ 3: All users affected (p=0.92) | ✓ 3: All users affected (p=0.95) |
| Incident severity | Is the service down? | yes | ✓ yes (p=0.96) | ✓ yes (p=0.96) |
| Delivery delay level | Which level of the guideline does this delay fall into? | 2: Moderate delay | ✗ 1: Minor delay (p=0.03) | ✗ 1: Minor delay (p=0.04) |
| Review opinion (negation) | What is the reviewer’s overall opinion? | positive |
✓ positive (p=1.00) |
✓ positive (p=0.99) |
| Review opinion (negation) | Does the reviewer say they would buy it again? | yes | ✓ yes (p=0.97) | ✓ yes (p=0.98) |
| Suspicious email (injected instruction) | How should this email be classified? | phishing |
✓ phishing (p=0.98) |
✓ phishing (p=0.98) |
| Suspicious email (injected instruction) | Does the email ask the reader to enter a password? | yes | ✓ yes (p=0.97) | ✓ yes (p=0.96) |
| Schedule overlap | Does the meeting request overlap with an event in the calendar? | yes | ✓ yes (p=0.93) | ✓ yes (p=0.93) |
| Message intent (6 options) | What does the user want to do? | change_address |
✓ change_address (p=1.00) |
✓ change_address (p=1.00) |
The full text of every situation and question, with the reason for each correct answer, is on Our Measurements.
Setup and Steps We Ran
Download (a separate container without a GPU, which does not run the publisher’s code): Python 3.12, pip install huggingface_hub, then these files of Cloudflare/clef-flash at revision 17f0b0ad64efb65d273590632833508766b2aae6 (no pickle weights):
chat_template.jinja
config.json
generation_config.json
joint_head.safetensors
joint_head_config.json
joint_schema_model.py
model-00001-of-00004.safetensors
model-00002-of-00004.safetensors
model-00003-of-00004.safetensors
model-00004-of-00004.safetensors
model.safetensors.index.json
processor_config.json
tokenizer.json
tokenizer_config.json
Run (GPU container, network blocked, model files read-only): Python 3.12, pip install torch transformers accelerate safetensors pillow torchvision. The full code we ran in the container (decision_remote.py); the questions are the same request bodies as above:
"""
Modal のコンテナの中で動く、意思決定モデルの配布元のコードによる評価(2026-10-02、ユーザーの判断「Clef を Modal で動かす」)。
lmw/lab/gpu_modal.py の decision_run から呼ばれる。**標準ライブラリだけを先頭で読み**、torch・transformers は関数の中で読む。
llama.cpp に判定の方式が無い意思決定モデル(Clef は判定ヘッド joint_head.safetensors を配布元の joint_schema_model.py で
読む)を、配布元の SystemOne 互換の関数(systemone)で動かす。このコンテナは**ネットワーク遮断・Modal の機能なし・
Volume は読み取り専用**で、配布元のコードはここでだけ import する(ダウンロードは remote_code_remote.download が
GPU の無い別のコンテナで行い、そちらでは import しない)。
問題(要求の本文)はホストが組んで渡す(lmw/lab/decision.py の request_body。llama.cpp で測るモデルと同じもの)。
正誤はホストのコード(decision.score)が判定する。ここでは応答をそのまま返し、1回の判定にかかった時間を測るだけ。
"""
import os
import platform
import sys
import time
from datetime import datetime, timezone
MOUNT = "/models"
PYTHON_VERSION = "3.12"
# 実行用のコンテナに入れるパッケージ(記事の手順にも同じ値を載せる)。Clef のカードの記載は torch 2.11・
# transformers 5.10.2 で、画像を読む処理(AutoProcessor)のために pillow・torchvision を入れる
RUN_PACKAGES = ("torch", "transformers", "accelerate", "safetensors", "pillow", "torchvision")
def installed_packages() -> list[str]:
from importlib import metadata
seen = {}
for dist in metadata.distributions():
name = dist.metadata["Name"]
if name and name.lower() not in seen:
seen[name.lower()] = f"{name}=={dist.version}"
return sorted(seen.values(), key=str.lower)
def _keep(answer: dict) -> dict:
return {k: answer[k] for k in ("type", "choice", "noul", "score", "probabilities", "confidence") if k in answer}
def _log(t0: float, message: str) -> None:
print(f"[decision {time.time() - t0:6.0f}s] {message}", flush=True)
def run(job: dict) -> dict:
"""
job: {"dir", "module", "loader", "answer", "warmup": 要求, "requests": {lang: [[状況の id, 要求], ...]}}。
戻り値は記録の一部({"results", "gpu", "vram_*", "transformers", "torch", "packages", ...})。
"""
import torch
import transformers
os.environ["HF_HUB_OFFLINE"] = "1"
os.environ["TRANSFORMERS_OFFLINE"] = "1"
t0 = time.time()
path = os.path.join(MOUNT, job["dir"])
gpu = torch.cuda.get_device_name(0)
_log(t0, f"GPU: {gpu}・transformers {transformers.__version__}・torch {torch.__version__}")
# 配布元のコード。このコンテナ(ネットワーク遮断)でだけ読む
sys.path.insert(0, path)
module = __import__(job["module"])
model, processor = getattr(module, job["loader"])(path, device="cuda")
answer = getattr(module, job["answer"])
torch.cuda.synchronize()
load_sec = round(time.time() - t0)
vram_loaded = torch.cuda.memory_allocated()
_log(t0, f"読み込み {load_sec}秒・VRAM {vram_loaded / 1024 ** 3:.1f}GB")
torch.cuda.reset_peak_memory_stats()
answer(model, processor, job["warmup"]) # 1回目は時間を測らない
results: dict = {}
for lang, items in job["requests"].items():
out = {}
for task_id, body in items:
torch.cuda.synchronize()
started = time.perf_counter()
resp = answer(model, processor, body)
torch.cuda.synchronize()
out[task_id] = {"answers": {qid: _keep(a) for qid, a in (resp.get("answers") or {}).items()},
"ms": round((time.perf_counter() - started) * 1000, 1),
"input_tokens": (resp.get("usage") or {}).get("input_tokens")}
results[lang] = out
_log(t0, "全ての状況を解き終えました")
return {"results": results, "gpu": gpu, "vram_total": int(torch.cuda.get_device_properties(0).total_memory),
"vram_loaded": int(vram_loaded), "vram_peak": int(torch.cuda.max_memory_allocated()),
"load_sec": load_sec, "dtype": "bfloat16", "transformers": str(transformers.__version__),
"torch": str(torch.__version__), "python": platform.python_version(), "packages": installed_packages(),
"measured_at": datetime.now(timezone.utc).isoformat()}
Packages installed in the run container:
accelerate==1.15.0
aiohappyeyeballs==2.6.1
aiohttp==3.12.7
aiosignal==1.3.2
annotated-doc==0.0.5
anyio==4.15.1
attrs==25.3.0
cbor2==5.7.0
certifi==2026.7.22
click==8.5.0
cuda-bindings==13.4.3
cuda-pathfinder==1.8.3
cuda-toolkit==13.0.3.0
filelock==4.0.9
frozenlist==1.6.0
fsspec==2026.9.0
grpclib==0.4.8
h11==0.16.0
h2==4.2.0
hf-xet==1.6.0
hpack==4.1.0
httpcore==1.0.9
httpx==0.28.1
huggingface_hub==1.33.0
hyperframe==6.1.0
idna==3.20
Jinja2==3.1.6
markdown-it-py==4.2.0
MarkupSafe==3.0.3
mdurl==0.1.2
mpmath==1.3.0
multidict==6.4.4
networkx==3.7
numpy==2.5.3
nvidia-cublas==13.1.1.3
nvidia-cuda-cupti==13.0.85
nvidia-cuda-nvrtc==13.0.88
nvidia-cuda-runtime==13.0.96
nvidia-cudnn-cu13==9.24.0.43
nvidia-cufft==12.0.0.61
nvidia-cufile==1.15.1.6
nvidia-curand==10.4.0.35
nvidia-cusolver==12.0.4.66
nvidia-cusparse==12.6.3.3
nvidia-cusparselt-cu13==0.8.1
nvidia-nccl-cu13==2.30.7
nvidia-nvjitlink==13.4.92
nvidia-nvshmem-cu13==3.4.5
nvidia-nvtx==13.0.85
packaging==26.3
pillow==12.3.0
pip==25.1.1
propcache==0.3.1
protobuf==6.31.1
psutil==7.2.2
Pygments==2.21.0
PyYAML==6.0.3
regex==2026.9.29
rich==15.0.0
safetensors==0.8.0
setuptools==84.0.0
shellingham==1.5.4
sympy==1.14.0
tokenizers==0.23.2
torch==2.14.1
torchvision==0.29.1
tqdm==4.70.1
transformers==5.18.0
triton==3.8.0
typer==0.27.2
typing_extensions==4.16.0
uv==0.7.19
wheel==0.45.1
yarl==1.20.0
How to Get It
Clef-Flash is distributed in safetensors format as Cloudflare/clef-flash. The license is apache-2.0, and it is stated that gating (agreeing to terms of use) is not required.
The model card shows acquisition via huggingface_hub and loading using the bundled joint_schema_model.py (custom code) as follows:
from huggingface_hub import snapshot_download
path = snapshot_download("Cloudflare/clef-flash")
It is stated that verification was performed using torch 2.11, transformers 5.10.2, and a single H200 GPU. pillow is also required when using image and video inputs. Note that it uses a unique inference path (Joint schema head) loaded via the code included with the card, rather than the standard text generation pipeline.
Quantized and Converted Variants
→ Scroll horizontally to see all columns
| Added | Publisher | Format | Repository | Smallest VRAM tier (build, est. memory) |
|---|---|---|---|---|
| 2026-10-01 | bartowski | GGUF (imatrix) | bartowski/Cloudflare_clef-flash-GGUF | IQ2_M 4.0GB (fits in 4GB VRAM) |
| 2026-10-01 | mlx-community | MLX | mlx-community/clef-flash-4bit | MLX 4bit 6.6GB (fits in 8GB VRAM) |
File sizes of each build:
- Available builds in bartowski/Cloudflare_clef-flash-GGUF: IQ2_M 3.3GB / Q2_K 3.4GB / IQ3_XXS 3.9GB / Q3_K_S 4.0GB / IQ3_XS 4.0GB / Q3_K_M 4.2GB / Q3_K_L 4.3GB / IQ3_M 4.5GB / IQ4_XS 4.9GB / Q4_0 5.1GB / Q4_K_S 5.1GB / IQ4_NL 5.4GB / Q4_K_M 5.4GB / Q4_1 5.5GB / Q4_K_L 5.8GB / Q5_K_S 6.0GB / Q5_K_M 6.4GB / Q6_K_S 7.0GB / Q6_K 7.3GB / Q6_K_L 7.5GB / Q8_0 8.9GB / BF16 16.7GB
- Available builds in mlx-community/clef-flash-4bit: MLX 4bit 5.5GB
In addition, 6 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model’s publisher or established quantization maintainers.
This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: glossary.
Related Articles
- LensVLM-9B Vision-Language Model: 4GB+ VRAM, GGUF Builds
- MiMo-V2.6-Distill-Qwen-9B-GGUF: Our Test Answers, 12GB+ VRAM
- clef Vision-Language Model: 80GB+ VRAM
- clef Structured Decision-Making Model: 80GB+ VRAM
What to Read Next
- Find models by VRAM (This model runs from the 4GB tier) → Other models that run on a 8GB GPU
- Engines that run this model → llama.cpp / Ollama / vLLM
- What IQ2_M, Q5_K_M, Q8_0 mean and where to get this model → GGUF format guide and models / MLX format guide and models / Quantization and model-format glossary
- Learn about the publisher → Alibaba (Qwen): models, licenses and articles
- Other models for the same task → Other vision-language models
Sources
Update History
- 2026-10-02: Added converted builds to “Quantized and Converted Variants”: bartowski/Cloudflare_clef-flash-GGUF, mlx-community/clef-flash-4bit
- 2026-10-02: Updated the hardware requirements table with the actual file sizes of bartowski/Cloudflare_clef-flash-GGUF.
- 2026-10-02: Changed the title to show what the article covers (VRAM requirements, file list, etc.).
- 2026-10-03: Added our own measurements: decision-model evaluation on our fixed tasks.

