clef Vision-Language Model: 80GB+ VRAM

At a Glance
| Item | Value |
|---|---|
| Repository | Cloudflare/clef |
| Family guide | Qwen3.8 guide (9 articles) |
| Publisher guide | Alibaba (Qwen): models and licenses |
| Published | 2026-10-01 |
| License | apache-2.0 |
| Formats | safetensors |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
What we checked ourselves
- Gave the 20 decision questions we give every decision model to this model, running the publisher’s own code on a cloud GPU (NVIDIA A100 80GB PCIe) with its network cut off: 20 correct in Japanese and 19 in English.
Details and conditions are in “Our Own Measurements” below.
Overview
Cloudflare has released Clef and its faster counterpart Clef-flash on Hugging Face, open-weight decision models that accept text, JSON, images, and video as inputs and return structured, probability-assigned outputs in a single batch. Built on top of Qwen3.8-27B through post-training, they are optimized specifically for decision-making tasks. Released under the Apache 2.0 license, they can be run and tested in local environments.
Specifications
- Parameter Count: 27B
- Architecture: Dense model based on Qwen3.8-27B (equipped with a vision encoder and a joint schema head powered by a small transformer head)
- Context Length: 64,000 tokens
Performance
Comparisons with major models based on the Decision Index (version 0.2.1) and Workflow evals evaluation tables from the model card are shown below (comparing Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev 9B, and Laya).
→ Scroll horizontally to see all columns
| Benchmark | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B |
|---|---|---|---|---|---|
| BFCL (case exact accuracy) | 98.5 | 98.8 | 95.8 | 96.5 | 94.5 |
| ToolRet (nDCG@10) | 69.2 | 66.4 | 65.3 | 61.2 | 64.3 |
| API-Bank (accuracy) | 91.9 | 93.1 | 88.2 | 83.7 | 56.3 |
| BANKING77 (macro-F1) | 94.2 | 90.9 | 79.7 | 74.3 | 84.8 |
| MMLU (accuracy) | 90.3 | 91.8 | 91.7 | 79.3 | 75.3 |
| GSM8K (accuracy) | 80.8 | 67.3 | 79.9 | 50.3 | 48.7 |
| Median latency (ms) | 209.3 | 38.8 | 524.1 | 84.4 | 51.4 |
Additionally, the Workflow evals results assessing four business workflows are as follows:
→ Scroll horizontally to see all columns
| Workflow | Metric | Clef | Clef-flash | Jev |
|---|---|---|---|---|
| Invoice processing | Exact actions | 64.7 | 57.1 | 61.8 |
| Invoice processing | Primary action | 86.2 | 73.3 | 83.1 |
| Customer service | Exact actions | 76.3 | 77.0 | 76.0 |
| Security incidents | Exact actions | 62.9 | 61.7 | 61.7 |
| Agent trace observability | Primary action | 68.5 | 69.8 | 71.6 |
According to measurements by the publishers, the Clef model outperforms other models such as Jev on metrics measuring tool use and function selection accuracy, including BFCL, ToolRet, and API-Bank, while holding a distinct advantage in latency (such as median latency). On the other hand, Jev scores higher on certain reasoning and commonsense reasoning metrics such as ANLI, BRIGHT, MMLU-Pro, and BBH, indicating that performance varies by task rather than being universally superior. Meanwhile, Clef-flash successfully reduces latency significantly while maintaining accuracy.
Strengths and Use Cases
Clef is a decision model that accepts text, JSON, images, and video as inputs and returns structured answers with probabilities for specified questions or schemas in a single forward pass, without requiring free-form text generation or output parsing. It is well-suited for use cases such as web domain classification, invoice processing, customer support routing, security incident triage, and agent decision processing. Furthermore, equipped with a vision encoder, it can also classify visual content contained in images and videos. The base model, Qwen3.8-27B, inherently boasts strengths in coding, professional tasks, research, and general long-context agent tasks.
How It Differs from Similar Models
Models based on Qwen3.8-27B include OrcaSAQ-2-27B, which focuses on lightweight deployment via quantization, and Hemmingway-1, which specializes in everyday text writing. While these are primarily aimed at standard text generation and composition, Clef specializes in classification and decision-making, significantly differing in that it returns fast, consistent typed outputs with probabilities without performing extraneous text generation.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 27.4B parameters
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 80GB class (A100 / H100) | BF16 | 51.0GB | 61.1GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
The publisher distributes this model as safetensors.
License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.
Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Our Own Measurements
Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.
Decision-Model Evaluation on Our Own Tasks
We gave this model the same 11 situations with 20 questions that we give every decision model (yes/no, named options and ordered levels; written by us, in Japanese and in English). llama.cpp cannot run this model’s decision head, so we ran the publisher’s own code (joint_schema_model.py and its SystemOne-compatible function) on a rented cloud GPU, inside a container with its network cut off. Each question states the rule to apply so that it has a single correct answer; our code decides right or wrong (yes when the probability of yes is 0.5 or more; for options and levels, the most probable one). These are not benchmark questions, and the result is not a general score of the model.
Conditions: Modal, NVIDIA A100 80GB PCIe (peak VRAM 51.4GB), BF16, transformers 5.18.0, torch 2.14.1+cu130, publisher revision 2f3de3dd85. The time per decision is measured on that GPU from passing one situation (1 to 3 questions) to the publisher’s function until it returns; it is not a guide to the speed of your own GPU and cannot be compared with times measured on our CPU.
→ Scroll horizontally to see all columns
| Model | Method | Quant | Correct (Japanese) | Correct (English) | Mean probability on the correct answer (ja / en) | Time per decision, median (ja / en) | Peak memory |
|---|---|---|---|---|---|---|---|
| clef | joint_schema_model.py |
BF16 |
20/20 | 19/20 | 0.98 / 0.94 | 190ms / 189ms (GPU: NVIDIA A100 80GB PCIe) | VRAM 51.4GB |
| Kev-4B (reference) | kev |
Q4_K_M |
18/20 | 19/20 | 0.87 / 0.87 | 7,616ms / 7,409ms | 6.3GB |
| clef-flash | joint_schema_model.py |
BF16 |
18/20 | 18/20 | 0.88 / 0.88 | 170ms / 167ms (GPU: NVIDIA L4) | VRAM 17.9GB |
| lev (reference) | lev |
Q4_K_M |
16/20 | 16/20 | 0.76 / 0.80 | 14,061ms / 13,186ms | 6.0GB |
| Laya (reference) | laya |
Q8_0 |
10/20 | 15/20 | 0.47 / 0.68 | 968ms / 678ms | 0.9GB |
| Julia-1 (reference) | laya |
Q8_0 |
13/20 | 10/20 | 0.60 / 0.45 | 136ms / 107ms | 0.5GB |
In Japanese it got 1 more right than in English (20 vs 19 of 20).
Other rows are other decision models measured the same way (“reference” rows are models we measure for comparison). Methods differ by model, and each model is measured with its own quantization.
Times marked (GPU) were measured on a cloud GPU and the others on our CPU, so the times cannot be compared across those rows; the numbers of correct answers can.
Answer to Each Question
→ Scroll horizontally to see all columns
| Situation | Question | Correct answer | Japanese | English |
|---|---|---|---|---|
| Routing a support message | Which team should handle this message? | billing |
✓ billing (p=0.99) |
✓ billing (p=0.99) |
| Routing a support message | Is the customer asking for money back? | yes | ✓ yes (p=0.99) | ✓ yes (p=0.99) |
| Routing a support message | Does the message report that a service is down? | no | ✓ no (p=0.99) | ✓ no (p=0.99) |
| Return eligibility (within the window) | Is today within 30 days of the delivery date? | yes | ✓ yes (p=0.99) | ✓ yes (p=0.98) |
| Return eligibility (within the window) | Can this item be returned under the policy? | yes | ✓ yes (p=0.99) | ✓ yes (p=0.94) |
| Return eligibility (past the window) | Is today within 30 days of the delivery date? | no | ✓ no (p=0.99) | ✓ no (p=1.00) |
| Return eligibility (past the window) | Can this item be returned under the policy? | no | ✓ no (p=0.99) | ✓ no (p=0.99) |
| Invoice handling (vendor) | What should happen to this invoice? | reject |
✓ reject (p=0.99) |
✓ reject (p=0.99) |
| Invoice handling (vendor) | Is the invoice total above 1,000 USD? | no | ✓ no (p=0.99) | ✓ no (p=0.99) |
| Invoice handling (amount) | What should happen to this invoice? | manager |
✓ manager (p=0.98) |
✓ manager (p=0.99) |
| Invoice handling (amount) | Is the invoice total above 1,000 USD? | yes | ✓ yes (p=0.99) | ✓ yes (p=0.99) |
| Incident severity | How widespread is the impact of this incident? | 3: All users affected | ✓ 3: All users affected (p=0.94) | ✓ 3: All users affected (p=0.95) |
| Incident severity | Is the service down? | yes | ✓ yes (p=0.98) | ✓ yes (p=0.99) |
| Delivery delay level | Which level of the guideline does this delay fall into? | 2: Moderate delay | ✓ 2: Moderate delay (p=0.90) | ✓ 2: Moderate delay (p=0.93) |
| Review opinion (negation) | What is the reviewer’s overall opinion? | positive |
✓ positive (p=0.98) |
✓ positive (p=0.99) |
| Review opinion (negation) | Does the reviewer say they would buy it again? | yes | ✓ yes (p=0.99) | ✓ yes (p=0.99) |
| Suspicious email (injected instruction) | How should this email be classified? | phishing |
✓ phishing (p=0.98) |
✓ phishing (p=0.99) |
| Suspicious email (injected instruction) | Does the email ask the reader to enter a password? | yes | ✓ yes (p=0.99) | ✓ yes (p=0.99) |
| Schedule overlap | Does the meeting request overlap with an event in the calendar? | yes | ✓ yes (p=0.96) | ✗ no (p=0.21) |
| Message intent (6 options) | What does the user want to do? | change_address |
✓ change_address (p=0.99) |
✓ change_address (p=0.99) |
The full text of every situation and question, with the reason for each correct answer, is on Our Measurements.
Setup and Steps We Ran
Download (a separate container without a GPU, which does not run the publisher’s code): Python 3.12, pip install huggingface_hub, then these files of Cloudflare/clef at revision 2f3de3dd85f379784083b0814d997ab627200f0c (no pickle weights):
chat_template.jinja
config.json
generation_config.json
joint_head.safetensors
joint_head_config.json
joint_schema_model.py
model-00001-of-00012.safetensors
model-00002-of-00012.safetensors
model-00003-of-00012.safetensors
model-00004-of-00012.safetensors
model-00005-of-00012.safetensors
model-00006-of-00012.safetensors
model-00007-of-00012.safetensors
model-00008-of-00012.safetensors
model-00009-of-00012.safetensors
model-00010-of-00012.safetensors
model-00011-of-00012.safetensors
model-00012-of-00012.safetensors
model.safetensors.index.json
processor_config.json
tokenizer.json
tokenizer_config.json
Run (GPU container, network blocked, model files read-only): Python 3.12, pip install torch transformers accelerate safetensors pillow torchvision. The full code we ran in the container (decision_remote.py); the questions are the same request bodies as above:
"""
Modal のコンテナの中で動く、意思決定モデルの配布元のコードによる評価(2026-10-02、ユーザーの判断「Clef を Modal で動かす」)。
lmw/lab/gpu_modal.py の decision_run から呼ばれる。**標準ライブラリだけを先頭で読み**、torch・transformers は関数の中で読む。
llama.cpp に判定の方式が無い意思決定モデル(Clef は判定ヘッド joint_head.safetensors を配布元の joint_schema_model.py で
読む)を、配布元の SystemOne 互換の関数(systemone)で動かす。このコンテナは**ネットワーク遮断・Modal の機能なし・
Volume は読み取り専用**で、配布元のコードはここでだけ import する(ダウンロードは remote_code_remote.download が
GPU の無い別のコンテナで行い、そちらでは import しない)。
問題(要求の本文)はホストが組んで渡す(lmw/lab/decision.py の request_body。llama.cpp で測るモデルと同じもの)。
正誤はホストのコード(decision.score)が判定する。ここでは応答をそのまま返し、1回の判定にかかった時間を測るだけ。
"""
import os
import platform
import sys
import time
from datetime import datetime, timezone
MOUNT = "/models"
PYTHON_VERSION = "3.12"
# 実行用のコンテナに入れるパッケージ(記事の手順にも同じ値を載せる)。Clef のカードの記載は torch 2.11・
# transformers 5.10.2 で、画像を読む処理(AutoProcessor)のために pillow・torchvision を入れる
RUN_PACKAGES = ("torch", "transformers", "accelerate", "safetensors", "pillow", "torchvision")
def installed_packages() -> list[str]:
from importlib import metadata
seen = {}
for dist in metadata.distributions():
name = dist.metadata["Name"]
if name and name.lower() not in seen:
seen[name.lower()] = f"{name}=={dist.version}"
return sorted(seen.values(), key=str.lower)
def _keep(answer: dict) -> dict:
return {k: answer[k] for k in ("type", "choice", "noul", "score", "probabilities", "confidence") if k in answer}
def _log(t0: float, message: str) -> None:
print(f"[decision {time.time() - t0:6.0f}s] {message}", flush=True)
def run(job: dict) -> dict:
"""
job: {"dir", "module", "loader", "answer", "warmup": 要求, "requests": {lang: [[状況の id, 要求], ...]}}。
戻り値は記録の一部({"results", "gpu", "vram_*", "transformers", "torch", "packages", ...})。
"""
import torch
import transformers
os.environ["HF_HUB_OFFLINE"] = "1"
os.environ["TRANSFORMERS_OFFLINE"] = "1"
t0 = time.time()
path = os.path.join(MOUNT, job["dir"])
gpu = torch.cuda.get_device_name(0)
_log(t0, f"GPU: {gpu}・transformers {transformers.__version__}・torch {torch.__version__}")
# 配布元のコード。このコンテナ(ネットワーク遮断)でだけ読む
sys.path.insert(0, path)
module = __import__(job["module"])
model, processor = getattr(module, job["loader"])(path, device="cuda")
answer = getattr(module, job["answer"])
torch.cuda.synchronize()
load_sec = round(time.time() - t0)
vram_loaded = torch.cuda.memory_allocated()
_log(t0, f"読み込み {load_sec}秒・VRAM {vram_loaded / 1024 ** 3:.1f}GB")
torch.cuda.reset_peak_memory_stats()
answer(model, processor, job["warmup"]) # 1回目は時間を測らない
results: dict = {}
for lang, items in job["requests"].items():
out = {}
for task_id, body in items:
torch.cuda.synchronize()
started = time.perf_counter()
resp = answer(model, processor, body)
torch.cuda.synchronize()
out[task_id] = {"answers": {qid: _keep(a) for qid, a in (resp.get("answers") or {}).items()},
"ms": round((time.perf_counter() - started) * 1000, 1),
"input_tokens": (resp.get("usage") or {}).get("input_tokens")}
results[lang] = out
_log(t0, "全ての状況を解き終えました")
return {"results": results, "gpu": gpu, "vram_total": int(torch.cuda.get_device_properties(0).total_memory),
"vram_loaded": int(vram_loaded), "vram_peak": int(torch.cuda.max_memory_allocated()),
"load_sec": load_sec, "dtype": "bfloat16", "transformers": str(transformers.__version__),
"torch": str(torch.__version__), "python": platform.python_version(), "packages": installed_packages(),
"measured_at": datetime.now(timezone.utc).isoformat()}
Packages installed in the run container:
accelerate==1.15.0
aiohappyeyeballs==2.6.1
aiohttp==3.12.7
aiosignal==1.3.2
annotated-doc==0.0.5
anyio==4.15.1
attrs==25.3.0
cbor2==5.7.0
certifi==2026.7.22
click==8.5.0
cuda-bindings==13.4.3
cuda-pathfinder==1.8.3
cuda-toolkit==13.0.3.0
filelock==4.0.9
frozenlist==1.6.0
fsspec==2026.9.0
grpclib==0.4.8
h11==0.16.0
h2==4.2.0
hf-xet==1.6.0
hpack==4.1.0
httpcore==1.0.9
httpx==0.28.1
huggingface_hub==1.33.0
hyperframe==6.1.0
idna==3.20
Jinja2==3.1.6
markdown-it-py==4.2.0
MarkupSafe==3.0.3
mdurl==0.1.2
mpmath==1.3.0
multidict==6.4.4
networkx==3.7
numpy==2.5.3
nvidia-cublas==13.1.1.3
nvidia-cuda-cupti==13.0.85
nvidia-cuda-nvrtc==13.0.88
nvidia-cuda-runtime==13.0.96
nvidia-cudnn-cu13==9.24.0.43
nvidia-cufft==12.0.0.61
nvidia-cufile==1.15.1.6
nvidia-curand==10.4.0.35
nvidia-cusolver==12.0.4.66
nvidia-cusparse==12.6.3.3
nvidia-cusparselt-cu13==0.8.1
nvidia-nccl-cu13==2.30.7
nvidia-nvjitlink==13.4.92
nvidia-nvshmem-cu13==3.4.5
nvidia-nvtx==13.0.85
packaging==26.3
pillow==12.3.0
pip==25.1.1
propcache==0.3.1
protobuf==6.31.1
psutil==7.2.2
Pygments==2.21.0
PyYAML==6.0.3
regex==2026.9.29
rich==15.0.0
safetensors==0.8.0
setuptools==84.0.0
shellingham==1.5.4
sympy==1.14.0
tokenizers==0.23.2
torch==2.14.1
torchvision==0.29.1
tqdm==4.70.1
transformers==5.18.0
triton==3.8.0
typer==0.27.2
typing_extensions==4.16.0
uv==0.7.19
wheel==0.45.1
yarl==1.20.0
How to Get It
Model weights are available from the Hugging Face repository. Provided under the Apache 2.0 license, commercial use and local execution are permitted.
huggingface-cli download Cloudflare/clef
It can be loaded using joint_schema_model.py included in the repository, and is compatible with transformers and torch. Hosted versions are also available via Cloudflare Workers AI.
Related Articles
- clef Structured Decision-Making Model: 80GB+ VRAM
- OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM
- Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds
- Qwen3.8-27B Vision-Language Model: Our Test Answers, 8GB+ VRAM
What to Read Next
- Find models by VRAM (This model needs at least 80GB) → VRAM quick reference
- Explore the same model family → Qwen3.8 family overview (9 articles, 11 converted builds)
- Engines that run this model → vLLM
- Formats this model is available in → Safetensors format guide and models
- Learn about the publisher → Alibaba (Qwen): models, licenses and articles
Sources
- https://blog.cloudflare.com/clef-decision-models/
- https://huggingface.co/Cloudflare/clef
- https://huggingface.co/Qwen/Qwen3.8-27B
Update History
- 2026-10-03: Added our own measurements: decision-model evaluation on our fixed tasks.

