clef Vision-Language Model: 80GB+ VRAM

October 3, 2026

clef Vision-Language Model: 80GB+ VRAM

At a Glance

Item Value
Repository Cloudflare/clef
Family guide Qwen3.8 guide (9 articles)
Publisher guide Alibaba (Qwen): models and licenses
Published 2026-10-01
License apache-2.0
Formats safetensors
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

What we checked ourselves

  • Gave the 20 decision questions we give every decision model to this model, running the publisher’s own code on a cloud GPU (NVIDIA A100 80GB PCIe) with its network cut off: 20 correct in Japanese and 19 in English.

Details and conditions are in “Our Own Measurements” below.

Overview

Cloudflare has released Clef and its faster counterpart Clef-flash on Hugging Face, open-weight decision models that accept text, JSON, images, and video as inputs and return structured, probability-assigned outputs in a single batch. Built on top of Qwen3.8-27B through post-training, they are optimized specifically for decision-making tasks. Released under the Apache 2.0 license, they can be run and tested in local environments.

Specifications

  • Parameter Count: 27B
  • Architecture: Dense model based on Qwen3.8-27B (equipped with a vision encoder and a joint schema head powered by a small transformer head)
  • Context Length: 64,000 tokens

Performance

Comparisons with major models based on the Decision Index (version 0.2.1) and Workflow evals evaluation tables from the model card are shown below (comparing Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev 9B, and Laya).

→ Scroll horizontally to see all columns

Benchmark Clef Clef-flash Jev DiffusionGemma Jev Kev 9B
BFCL (case exact accuracy) 98.5 98.8 95.8 96.5 94.5
ToolRet (nDCG@10) 69.2 66.4 65.3 61.2 64.3
API-Bank (accuracy) 91.9 93.1 88.2 83.7 56.3
BANKING77 (macro-F1) 94.2 90.9 79.7 74.3 84.8
MMLU (accuracy) 90.3 91.8 91.7 79.3 75.3
GSM8K (accuracy) 80.8 67.3 79.9 50.3 48.7
Median latency (ms) 209.3 38.8 524.1 84.4 51.4

Additionally, the Workflow evals results assessing four business workflows are as follows:

→ Scroll horizontally to see all columns

Workflow Metric Clef Clef-flash Jev
Invoice processing Exact actions 64.7 57.1 61.8
Invoice processing Primary action 86.2 73.3 83.1
Customer service Exact actions 76.3 77.0 76.0
Security incidents Exact actions 62.9 61.7 61.7
Agent trace observability Primary action 68.5 69.8 71.6

According to measurements by the publishers, the Clef model outperforms other models such as Jev on metrics measuring tool use and function selection accuracy, including BFCL, ToolRet, and API-Bank, while holding a distinct advantage in latency (such as median latency). On the other hand, Jev scores higher on certain reasoning and commonsense reasoning metrics such as ANLI, BRIGHT, MMLU-Pro, and BBH, indicating that performance varies by task rather than being universally superior. Meanwhile, Clef-flash successfully reduces latency significantly while maintaining accuracy.

Strengths and Use Cases

Clef is a decision model that accepts text, JSON, images, and video as inputs and returns structured answers with probabilities for specified questions or schemas in a single forward pass, without requiring free-form text generation or output parsing. It is well-suited for use cases such as web domain classification, invoice processing, customer support routing, security incident triage, and agent decision processing. Furthermore, equipped with a vision encoder, it can also classify visual content contained in images and videos. The base model, Qwen3.8-27B, inherently boasts strengths in coding, professional tasks, research, and general long-context agent tasks.

How It Differs from Similar Models

Models based on Qwen3.8-27B include OrcaSAQ-2-27B, which focuses on lightweight deployment via quantization, and Hemmingway-1, which specializes in everyday text writing. While these are primarily aimed at standard text generation and composition, Clef specializes in classification and decision-making, significantly differing in that it returns fast, consistent typed outputs with probabilities without performing extraneous text generation.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 27.4B parameters

Your VRAM Quantization File size Est. memory needed
80GB class (A100 / H100) BF16 51.0GB 61.1GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

The publisher distributes this model as safetensors.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Our Own Measurements

Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.

Decision-Model Evaluation on Our Own Tasks

We gave this model the same 11 situations with 20 questions that we give every decision model (yes/no, named options and ordered levels; written by us, in Japanese and in English). llama.cpp cannot run this model’s decision head, so we ran the publisher’s own code (joint_schema_model.py and its SystemOne-compatible function) on a rented cloud GPU, inside a container with its network cut off. Each question states the rule to apply so that it has a single correct answer; our code decides right or wrong (yes when the probability of yes is 0.5 or more; for options and levels, the most probable one). These are not benchmark questions, and the result is not a general score of the model.

Conditions: Modal, NVIDIA A100 80GB PCIe (peak VRAM 51.4GB), BF16, transformers 5.18.0, torch 2.14.1+cu130, publisher revision 2f3de3dd85. The time per decision is measured on that GPU from passing one situation (1 to 3 questions) to the publisher’s function until it returns; it is not a guide to the speed of your own GPU and cannot be compared with times measured on our CPU.

→ Scroll horizontally to see all columns

Model Method Quant Correct (Japanese) Correct (English) Mean probability on the correct answer (ja / en) Time per decision, median (ja / en) Peak memory
clef joint_schema_model.py BF16 20/20 19/20 0.98 / 0.94 190ms / 189ms (GPU: NVIDIA A100 80GB PCIe) VRAM 51.4GB
Kev-4B (reference) kev Q4_K_M 18/20 19/20 0.87 / 0.87 7,616ms / 7,409ms 6.3GB
clef-flash joint_schema_model.py BF16 18/20 18/20 0.88 / 0.88 170ms / 167ms (GPU: NVIDIA L4) VRAM 17.9GB
lev (reference) lev Q4_K_M 16/20 16/20 0.76 / 0.80 14,061ms / 13,186ms 6.0GB
Laya (reference) laya Q8_0 10/20 15/20 0.47 / 0.68 968ms / 678ms 0.9GB
Julia-1 (reference) laya Q8_0 13/20 10/20 0.60 / 0.45 136ms / 107ms 0.5GB

In Japanese it got 1 more right than in English (20 vs 19 of 20).

Other rows are other decision models measured the same way (“reference” rows are models we measure for comparison). Methods differ by model, and each model is measured with its own quantization.

Times marked (GPU) were measured on a cloud GPU and the others on our CPU, so the times cannot be compared across those rows; the numbers of correct answers can.

Answer to Each Question

→ Scroll horizontally to see all columns

Situation Question Correct answer Japanese English
Routing a support message Which team should handle this message? billing ✓ billing (p=0.99) ✓ billing (p=0.99)
Routing a support message Is the customer asking for money back? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Routing a support message Does the message report that a service is down? no ✓ no (p=0.99) ✓ no (p=0.99)
Return eligibility (within the window) Is today within 30 days of the delivery date? yes ✓ yes (p=0.99) ✓ yes (p=0.98)
Return eligibility (within the window) Can this item be returned under the policy? yes ✓ yes (p=0.99) ✓ yes (p=0.94)
Return eligibility (past the window) Is today within 30 days of the delivery date? no ✓ no (p=0.99) ✓ no (p=1.00)
Return eligibility (past the window) Can this item be returned under the policy? no ✓ no (p=0.99) ✓ no (p=0.99)
Invoice handling (vendor) What should happen to this invoice? reject ✓ reject (p=0.99) ✓ reject (p=0.99)
Invoice handling (vendor) Is the invoice total above 1,000 USD? no ✓ no (p=0.99) ✓ no (p=0.99)
Invoice handling (amount) What should happen to this invoice? manager ✓ manager (p=0.98) ✓ manager (p=0.99)
Invoice handling (amount) Is the invoice total above 1,000 USD? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Incident severity How widespread is the impact of this incident? 3: All users affected ✓ 3: All users affected (p=0.94) ✓ 3: All users affected (p=0.95)
Incident severity Is the service down? yes ✓ yes (p=0.98) ✓ yes (p=0.99)
Delivery delay level Which level of the guideline does this delay fall into? 2: Moderate delay ✓ 2: Moderate delay (p=0.90) ✓ 2: Moderate delay (p=0.93)
Review opinion (negation) What is the reviewer’s overall opinion? positive ✓ positive (p=0.98) ✓ positive (p=0.99)
Review opinion (negation) Does the reviewer say they would buy it again? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Suspicious email (injected instruction) How should this email be classified? phishing ✓ phishing (p=0.98) ✓ phishing (p=0.99)
Suspicious email (injected instruction) Does the email ask the reader to enter a password? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Schedule overlap Does the meeting request overlap with an event in the calendar? yes ✓ yes (p=0.96) ✗ no (p=0.21)
Message intent (6 options) What does the user want to do? change_address ✓ change_address (p=0.99) ✓ change_address (p=0.99)

The full text of every situation and question, with the reason for each correct answer, is on Our Measurements.

Setup and Steps We Ran

Download (a separate container without a GPU, which does not run the publisher’s code): Python 3.12, pip install huggingface_hub, then these files of Cloudflare/clef at revision 2f3de3dd85f379784083b0814d997ab627200f0c (no pickle weights):

chat_template.jinja
config.json
generation_config.json
joint_head.safetensors
joint_head_config.json
joint_schema_model.py
model-00001-of-00012.safetensors
model-00002-of-00012.safetensors
model-00003-of-00012.safetensors
model-00004-of-00012.safetensors
model-00005-of-00012.safetensors
model-00006-of-00012.safetensors
model-00007-of-00012.safetensors
model-00008-of-00012.safetensors
model-00009-of-00012.safetensors
model-00010-of-00012.safetensors
model-00011-of-00012.safetensors
model-00012-of-00012.safetensors
model.safetensors.index.json
processor_config.json
tokenizer.json
tokenizer_config.json

Run (GPU container, network blocked, model files read-only): Python 3.12, pip install torch transformers accelerate safetensors pillow torchvision. The full code we ran in the container (decision_remote.py); the questions are the same request bodies as above:

"""
Modal のコンテナの中で動く、意思決定モデルの配布元のコードによる評価(2026-10-02、ユーザーの判断「Clef を Modal で動かす」)。
lmw/lab/gpu_modal.py の decision_run から呼ばれる。**標準ライブラリだけを先頭で読み**、torch・transformers は関数の中で読む。

llama.cpp に判定の方式が無い意思決定モデル(Clef は判定ヘッド joint_head.safetensors を配布元の joint_schema_model.py で
読む)を、配布元の SystemOne 互換の関数(systemone)で動かす。このコンテナは**ネットワーク遮断・Modal の機能なし・
Volume は読み取り専用**で、配布元のコードはここでだけ import する(ダウンロードは remote_code_remote.download が
GPU の無い別のコンテナで行い、そちらでは import しない)。

問題(要求の本文)はホストが組んで渡す(lmw/lab/decision.py の request_body。llama.cpp で測るモデルと同じもの)。
正誤はホストのコード(decision.score)が判定する。ここでは応答をそのまま返し、1回の判定にかかった時間を測るだけ。
"""
import os
import platform
import sys
import time
from datetime import datetime, timezone

MOUNT = "/models"
PYTHON_VERSION = "3.12"
# 実行用のコンテナに入れるパッケージ(記事の手順にも同じ値を載せる)。Clef のカードの記載は torch 2.11・
# transformers 5.10.2 で、画像を読む処理(AutoProcessor)のために pillow・torchvision を入れる
RUN_PACKAGES = ("torch", "transformers", "accelerate", "safetensors", "pillow", "torchvision")

def installed_packages() -> list[str]:
    from importlib import metadata
    seen = {}
    for dist in metadata.distributions():
        name = dist.metadata["Name"]
        if name and name.lower() not in seen:
            seen[name.lower()] = f"{name}=={dist.version}"
    return sorted(seen.values(), key=str.lower)

def _keep(answer: dict) -> dict:
    return {k: answer[k] for k in ("type", "choice", "noul", "score", "probabilities", "confidence") if k in answer}

def _log(t0: float, message: str) -> None:
    print(f"[decision {time.time() - t0:6.0f}s] {message}", flush=True)

def run(job: dict) -> dict:
    """
    job: {"dir", "module", "loader", "answer", "warmup": 要求, "requests": {lang: [[状況の id, 要求], ...]}}。
    戻り値は記録の一部({"results", "gpu", "vram_*", "transformers", "torch", "packages", ...})。
    """
    import torch
    import transformers
    os.environ["HF_HUB_OFFLINE"] = "1"
    os.environ["TRANSFORMERS_OFFLINE"] = "1"
    t0 = time.time()
    path = os.path.join(MOUNT, job["dir"])
    gpu = torch.cuda.get_device_name(0)
    _log(t0, f"GPU: {gpu}・transformers {transformers.__version__}・torch {torch.__version__}")
    # 配布元のコード。このコンテナ(ネットワーク遮断)でだけ読む
    sys.path.insert(0, path)
    module = __import__(job["module"])
    model, processor = getattr(module, job["loader"])(path, device="cuda")
    answer = getattr(module, job["answer"])
    torch.cuda.synchronize()
    load_sec = round(time.time() - t0)
    vram_loaded = torch.cuda.memory_allocated()
    _log(t0, f"読み込み {load_sec}秒・VRAM {vram_loaded / 1024 ** 3:.1f}GB")
    torch.cuda.reset_peak_memory_stats()
    answer(model, processor, job["warmup"])  # 1回目は時間を測らない
    results: dict = {}
    for lang, items in job["requests"].items():
        out = {}
        for task_id, body in items:
            torch.cuda.synchronize()
            started = time.perf_counter()
            resp = answer(model, processor, body)
            torch.cuda.synchronize()
            out[task_id] = {"answers": {qid: _keep(a) for qid, a in (resp.get("answers") or {}).items()},
                            "ms": round((time.perf_counter() - started) * 1000, 1),
                            "input_tokens": (resp.get("usage") or {}).get("input_tokens")}
        results[lang] = out
    _log(t0, "全ての状況を解き終えました")
    return {"results": results, "gpu": gpu, "vram_total": int(torch.cuda.get_device_properties(0).total_memory),
            "vram_loaded": int(vram_loaded), "vram_peak": int(torch.cuda.max_memory_allocated()),
            "load_sec": load_sec, "dtype": "bfloat16", "transformers": str(transformers.__version__),
            "torch": str(torch.__version__), "python": platform.python_version(), "packages": installed_packages(),
            "measured_at": datetime.now(timezone.utc).isoformat()}

Packages installed in the run container:

accelerate==1.15.0
aiohappyeyeballs==2.6.1
aiohttp==3.12.7
aiosignal==1.3.2
annotated-doc==0.0.5
anyio==4.15.1
attrs==25.3.0
cbor2==5.7.0
certifi==2026.7.22
click==8.5.0
cuda-bindings==13.4.3
cuda-pathfinder==1.8.3
cuda-toolkit==13.0.3.0
filelock==4.0.9
frozenlist==1.6.0
fsspec==2026.9.0
grpclib==0.4.8
h11==0.16.0
h2==4.2.0
hf-xet==1.6.0
hpack==4.1.0
httpcore==1.0.9
httpx==0.28.1
huggingface_hub==1.33.0
hyperframe==6.1.0
idna==3.20
Jinja2==3.1.6
markdown-it-py==4.2.0
MarkupSafe==3.0.3
mdurl==0.1.2
mpmath==1.3.0
multidict==6.4.4
networkx==3.7
numpy==2.5.3
nvidia-cublas==13.1.1.3
nvidia-cuda-cupti==13.0.85
nvidia-cuda-nvrtc==13.0.88
nvidia-cuda-runtime==13.0.96
nvidia-cudnn-cu13==9.24.0.43
nvidia-cufft==12.0.0.61
nvidia-cufile==1.15.1.6
nvidia-curand==10.4.0.35
nvidia-cusolver==12.0.4.66
nvidia-cusparse==12.6.3.3
nvidia-cusparselt-cu13==0.8.1
nvidia-nccl-cu13==2.30.7
nvidia-nvjitlink==13.4.92
nvidia-nvshmem-cu13==3.4.5
nvidia-nvtx==13.0.85
packaging==26.3
pillow==12.3.0
pip==25.1.1
propcache==0.3.1
protobuf==6.31.1
psutil==7.2.2
Pygments==2.21.0
PyYAML==6.0.3
regex==2026.9.29
rich==15.0.0
safetensors==0.8.0
setuptools==84.0.0
shellingham==1.5.4
sympy==1.14.0
tokenizers==0.23.2
torch==2.14.1
torchvision==0.29.1
tqdm==4.70.1
transformers==5.18.0
triton==3.8.0
typer==0.27.2
typing_extensions==4.16.0
uv==0.7.19
wheel==0.45.1
yarl==1.20.0

How to Get It

Model weights are available from the Hugging Face repository. Provided under the Apache 2.0 license, commercial use and local execution are permitted.

huggingface-cli download Cloudflare/clef

It can be loaded using joint_schema_model.py included in the repository, and is compatible with transformers and torch. Hosted versions are also available via Cloudflare Workers AI.

Related Articles

What to Read Next

Sources

Update History

  • 2026-10-03: Added our own measurements: decision-model evaluation on our fixed tasks.