clef Structured Decision-Making Model: 80GB+ VRAM

October 3, 2026

clef Structured Decision-Making Model: 80GB+ VRAM

At a Glance

Item Value
Repository Cloudflare/clef
Family guide Qwen3.8 guide (9 articles)
Publisher guide Alibaba (Qwen): models and licenses
Published 2026-10-01
License apache-2.0
Formats safetensors
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

What we checked ourselves

  • Gave the 20 decision questions we give every decision model to this model, running the publisher’s own code on a cloud GPU (NVIDIA A100 80GB PCIe) with its network cut off: 20 correct in Japanese and 19 in English.

Details and conditions are in “Our Own Measurements” below.

Overview

Cloudflare has released “Clef", a 27B parameter multimodal decision-making model based on Qwen3.8-27B. This model is designed to accept inputs such as text, JSON, images, video as “state", along with typed “question schemas", and directly output the probabilities of each choice. It is capable of making structured decisions in a single forward pass without requiring free-form text generation or output parsing.

Clef has undergone post-training based on Qwen3.8-27B, and according to Cloudflare’s blog, it is positioned as a “decision model" that reads state and makes schema-based decisions. The API is fully compatible with Jev and SystemOne, making it easy to integrate into practical workflows.

Source: Clef decision models on the Cloudflare blog

Specifications

  • Parameter Count: 27B
  • Architecture: Based on Qwen3.5 architecture (Dense model), equipped with a vision encoder
  • Context Length: 262,144 tokens (native), expandable up to 1,000,000 tokens
  • Input Formats: Text, JSON, images, video

Performance

Here are the comparison results from the decision task benchmark suite “Decision Index" listed in the model card. The table focuses on four columns: Clef, the lightweight version Clef-flash, and existing models Jev and Kev 9B.

Decision Index

→ Scroll horizontally to see all columns

Benchmark Clef Clef-flash Jev Kev 9B
BFCL (case exact accuracy) 98.5 98.8 95.8 94.5
ToolRet (nDCG@10) 69.2 66.4 65.3 64.3
API-Bank (accuracy) 91.9 93.1 88.2 56.3
BANKING77 (macro-F1) 94.2 90.9 79.7 84.8
CLINC150+OOS (macro-F1) 97.4 66.8 89.3 79.0
RouterBench (selected quality) 79.7 79.9 79.9 80.0
ANLI (macro-F1) 69.8 59.1 74.8 56.3
MMLU (accuracy) 90.3 91.8 91.7 75.3
GPQA Diamond (accuracy) 48.0 51.0 78.3 38.8
ARC-Challenge (accuracy) 97.7 98.3 97.8 93.7
WinoGrande (accuracy) 93.5 97.5 92.0 73.2
HellaSwag (accuracy) 98.2 98.6 94.5 81.9
GSM8K (accuracy) 80.8 67.3 79.9 48.7
CRUXEval (accuracy) 86.7 86.1 73.0 51.2
BBH (accuracy) 73.7 68.9 92.9 65.2
RAGTruth (hallucination F1) 79.4 35.6 76.5 46.2

According to measurements by the publishers, Clef records an extremely high score of 98.5% in BFCL (accuracy of function selection and argument assembly), showing exceptional tool-use capabilities as an agent. It also marks the highest values among the comparison targets in classification tasks such as BANKING77 and CLINC150+OOS. On the other hand, it stays at 48.0% in GPQA Diamond (difficult scientific questions), lagging behind Jev’s 78.3% in reasoning tasks requiring specialized scientific knowledge. Jev also holds the advantage in BBH (reasoning tasks).

Next are the results of “Workflow evals", which simulate business operations.

Workflow evals

→ Scroll horizontally to see all columns

Workflow Metric Clef Clef-flash Jev
Invoice processing Exact actions 64.7 57.1 61.8
Invoice processing Primary action 86.2 73.3 83.1
Customer service Exact actions 76.3 77.0 76.0
Security incidents Exact actions 62.9 61.7 61.7

In practical workflows such as invoice processing and security incident response, Clef demonstrates higher accuracy than the existing Jev. It records a particularly high figure of 86.2% in identifying primary actions for invoice processing.

Additionally, the performance of the base model Qwen3.8-27B itself is extremely high; in publisher benchmarks, it records 61.7% on SWE-bench Pro (real repository fixes) and 90.3% on LiveCodeBench v6 (competitive programming), possessing performance that overwhelms models of the same scale in coding and agent tasks. Clef is a model built upon this powerful foundation with specialized training for structured outputs.

Strengths and Use Cases

Clef excels at decision-making tasks where it computes and returns probabilities for predefined choices based on a given input “state" and typed “question schemas". Because it can take not only text and JSON data but also images and video frames simultaneously as state, it enables classification and judgment based on multimodal information.

Supported formats in the question schema include the noul type for boolean judgments, the choice type for selecting from multiple options, and the score type for graded evaluations. Inside the model, a small transformer head (joint schema head) reads the final hidden states of the backbone Qwen3.8-27B (including the vision encoder), and calculates scores (log-probabilities) for all question choices in a single forward pass.

This mechanism eliminates the need for free-form text generation or complex parsing processes to convert output results into JSON or the like, which were traditionally required with LLMs. Specifically, the following practical use cases are anticipated:

  • Automatic Invoice and Receipt Judgment: Evaluates payment statuses and amount conditions from image or text-format invoice data.
  • Customer Support Routing: Determines department assignments (billing, technical support, etc.) and urgency (same-day response, within this week, etc.) based on the body of user inquiries.
  • Security Incident Classification: Automatically identifies the presence or absence of service outages from incident reports and log messages.

Furthermore, because it has full compatibility with the APIs of existing decision systems Jev and SystemOne (POST /v1/systemone), it can be integrated directly into existing workflows that support these formats.

How It Differs from Similar Models

Compared to Qwen3.8-27B derivative models previously featured on this site, Clef adopts a clearly distinct approach.

OrcaSAQ-2-27B Text Generation Model: Example Responses and Required VRAM 16GB+ is a text model compressed to an average of 3.21-bit using proprietary quantization technology (SAQ2) with the objective of running the original 54GB model on a single 16GB GPU. In contrast, Clef does not compress parameters for lightweighting; instead, it retains the vision encoder in the backbone, adds a structured decision head, and supports multimodal inputs.

Additionally, Hemmingway-1 Everyday Text Language Model: Required VRAM 12GB+ / GGUF Available is a model specialized in directly generating human-written everyday texts such as emails and messages. Clef, on the other hand, does not generate natural language text; it outputs only probability values (probability distributions) for predefined question choices. The major difference is that while Hemmingway-1 aims to create text, Clef specializes in automating classification and decision-making.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 27.4B parameters

Your VRAM Quantization File size Est. memory needed
80GB class (A100 / H100) BF16 51.0GB 61.1GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

Not usable in Ollama, LM Studio and llama.cpp yet — we have found no GGUF build.

The publisher ships safetensors only. However, llama.cpp’s registry does list this architecture, so conversion to GGUF is possible and the model will run once someone publishes a converted build. 4 converted build(s) from other uploaders exist. Today it can be run with transformers or vLLM, using the memory figures in the table above.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compiled by this site’s code from the published formats, converted builds we have found, and each engine’s own model registry. “Not found" means we have not seen such a build, not that none exists. License summaries are not legal advice — check the publisher’s original terms before relying on them.

Our Own Measurements

Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.

Decision-Model Evaluation on Our Own Tasks

We gave this model the same 11 situations with 20 questions that we give every decision model (yes/no, named options and ordered levels; written by us, in Japanese and in English). llama.cpp cannot run this model’s decision head, so we ran the publisher’s own code (joint_schema_model.py and its SystemOne-compatible function) on a rented cloud GPU, inside a container with its network cut off. Each question states the rule to apply so that it has a single correct answer; our code decides right or wrong (yes when the probability of yes is 0.5 or more; for options and levels, the most probable one). These are not benchmark questions, and the result is not a general score of the model.

Conditions: Modal, NVIDIA A100 80GB PCIe (peak VRAM 51.4GB), BF16, transformers 5.18.0, torch 2.14.1+cu130, publisher revision 2f3de3dd85. The time per decision is measured on that GPU from passing one situation (1 to 3 questions) to the publisher’s function until it returns; it is not a guide to the speed of your own GPU and cannot be compared with times measured on our CPU.

→ Scroll horizontally to see all columns

Model Method Quant Correct (Japanese) Correct (English) Mean probability on the correct answer (ja / en) Time per decision, median (ja / en) Peak memory
clef joint_schema_model.py BF16 20/20 19/20 0.98 / 0.94 190ms / 189ms (GPU: NVIDIA A100 80GB PCIe) VRAM 51.4GB
Kev-4B (reference) kev Q4_K_M 18/20 19/20 0.87 / 0.87 7,616ms / 7,409ms 6.3GB
clef-flash joint_schema_model.py BF16 18/20 18/20 0.88 / 0.88 170ms / 167ms (GPU: NVIDIA L4) VRAM 17.9GB
lev (reference) lev Q4_K_M 16/20 16/20 0.76 / 0.80 14,061ms / 13,186ms 6.0GB
Laya (reference) laya Q8_0 10/20 15/20 0.47 / 0.68 968ms / 678ms 0.9GB
Julia-1 (reference) laya Q8_0 13/20 10/20 0.60 / 0.45 136ms / 107ms 0.5GB

In Japanese it got 1 more right than in English (20 vs 19 of 20).

Other rows are other decision models measured the same way (“reference” rows are models we measure for comparison). Methods differ by model, and each model is measured with its own quantization.

Times marked (GPU) were measured on a cloud GPU and the others on our CPU, so the times cannot be compared across those rows; the numbers of correct answers can.

Answer to Each Question

→ Scroll horizontally to see all columns

Situation Question Correct answer Japanese English
Routing a support message Which team should handle this message? billing ✓ billing (p=0.99) ✓ billing (p=0.99)
Routing a support message Is the customer asking for money back? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Routing a support message Does the message report that a service is down? no ✓ no (p=0.99) ✓ no (p=0.99)
Return eligibility (within the window) Is today within 30 days of the delivery date? yes ✓ yes (p=0.99) ✓ yes (p=0.98)
Return eligibility (within the window) Can this item be returned under the policy? yes ✓ yes (p=0.99) ✓ yes (p=0.94)
Return eligibility (past the window) Is today within 30 days of the delivery date? no ✓ no (p=0.99) ✓ no (p=1.00)
Return eligibility (past the window) Can this item be returned under the policy? no ✓ no (p=0.99) ✓ no (p=0.99)
Invoice handling (vendor) What should happen to this invoice? reject ✓ reject (p=0.99) ✓ reject (p=0.99)
Invoice handling (vendor) Is the invoice total above 1,000 USD? no ✓ no (p=0.99) ✓ no (p=0.99)
Invoice handling (amount) What should happen to this invoice? manager ✓ manager (p=0.98) ✓ manager (p=0.99)
Invoice handling (amount) Is the invoice total above 1,000 USD? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Incident severity How widespread is the impact of this incident? 3: All users affected ✓ 3: All users affected (p=0.94) ✓ 3: All users affected (p=0.95)
Incident severity Is the service down? yes ✓ yes (p=0.98) ✓ yes (p=0.99)
Delivery delay level Which level of the guideline does this delay fall into? 2: Moderate delay ✓ 2: Moderate delay (p=0.90) ✓ 2: Moderate delay (p=0.93)
Review opinion (negation) What is the reviewer’s overall opinion? positive ✓ positive (p=0.98) ✓ positive (p=0.99)
Review opinion (negation) Does the reviewer say they would buy it again? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Suspicious email (injected instruction) How should this email be classified? phishing ✓ phishing (p=0.98) ✓ phishing (p=0.99)
Suspicious email (injected instruction) Does the email ask the reader to enter a password? yes ✓ yes (p=0.99) ✓ yes (p=0.99)
Schedule overlap Does the meeting request overlap with an event in the calendar? yes ✓ yes (p=0.96) ✗ no (p=0.21)
Message intent (6 options) What does the user want to do? change_address ✓ change_address (p=0.99) ✓ change_address (p=0.99)

The full text of every situation and question, with the reason for each correct answer, is on Our Measurements.

Setup and Steps We Ran

Download (a separate container without a GPU, which does not run the publisher’s code): Python 3.12, pip install huggingface_hub, then these files of Cloudflare/clef at revision 2f3de3dd85f379784083b0814d997ab627200f0c (no pickle weights):

chat_template.jinja
config.json
generation_config.json
joint_head.safetensors
joint_head_config.json
joint_schema_model.py
model-00001-of-00012.safetensors
model-00002-of-00012.safetensors
model-00003-of-00012.safetensors
model-00004-of-00012.safetensors
model-00005-of-00012.safetensors
model-00006-of-00012.safetensors
model-00007-of-00012.safetensors
model-00008-of-00012.safetensors
model-00009-of-00012.safetensors
model-00010-of-00012.safetensors
model-00011-of-00012.safetensors
model-00012-of-00012.safetensors
model.safetensors.index.json
processor_config.json
tokenizer.json
tokenizer_config.json

Run (GPU container, network blocked, model files read-only): Python 3.12, pip install torch transformers accelerate safetensors pillow torchvision. The full code we ran in the container (decision_remote.py); the questions are the same request bodies as above:

"""
Modal のコンテナの中で動く、意思決定モデルの配布元のコードによる評価(2026-10-02、ユーザーの判断「Clef を Modal で動かす」)。
lmw/lab/gpu_modal.py の decision_run から呼ばれる。**標準ライブラリだけを先頭で読み**、torch・transformers は関数の中で読む。

llama.cpp に判定の方式が無い意思決定モデル(Clef は判定ヘッド joint_head.safetensors を配布元の joint_schema_model.py で
読む)を、配布元の SystemOne 互換の関数(systemone)で動かす。このコンテナは**ネットワーク遮断・Modal の機能なし・
Volume は読み取り専用**で、配布元のコードはここでだけ import する(ダウンロードは remote_code_remote.download が
GPU の無い別のコンテナで行い、そちらでは import しない)。

問題(要求の本文)はホストが組んで渡す(lmw/lab/decision.py の request_body。llama.cpp で測るモデルと同じもの)。
正誤はホストのコード(decision.score)が判定する。ここでは応答をそのまま返し、1回の判定にかかった時間を測るだけ。
"""
import os
import platform
import sys
import time
from datetime import datetime, timezone

MOUNT = "/models"
PYTHON_VERSION = "3.12"
# 実行用のコンテナに入れるパッケージ(記事の手順にも同じ値を載せる)。Clef のカードの記載は torch 2.11・
# transformers 5.10.2 で、画像を読む処理(AutoProcessor)のために pillow・torchvision を入れる
RUN_PACKAGES = ("torch", "transformers", "accelerate", "safetensors", "pillow", "torchvision")

def installed_packages() -> list[str]:
    from importlib import metadata
    seen = {}
    for dist in metadata.distributions():
        name = dist.metadata["Name"]
        if name and name.lower() not in seen:
            seen[name.lower()] = f"{name}=={dist.version}"
    return sorted(seen.values(), key=str.lower)

def _keep(answer: dict) -> dict:
    return {k: answer[k] for k in ("type", "choice", "noul", "score", "probabilities", "confidence") if k in answer}

def _log(t0: float, message: str) -> None:
    print(f"[decision {time.time() - t0:6.0f}s] {message}", flush=True)

def run(job: dict) -> dict:
    """
    job: {"dir", "module", "loader", "answer", "warmup": 要求, "requests": {lang: [[状況の id, 要求], ...]}}。
    戻り値は記録の一部({"results", "gpu", "vram_*", "transformers", "torch", "packages", ...})。
    """
    import torch
    import transformers
    os.environ["HF_HUB_OFFLINE"] = "1"
    os.environ["TRANSFORMERS_OFFLINE"] = "1"
    t0 = time.time()
    path = os.path.join(MOUNT, job["dir"])
    gpu = torch.cuda.get_device_name(0)
    _log(t0, f"GPU: {gpu}・transformers {transformers.__version__}・torch {torch.__version__}")
    # 配布元のコード。このコンテナ(ネットワーク遮断)でだけ読む
    sys.path.insert(0, path)
    module = __import__(job["module"])
    model, processor = getattr(module, job["loader"])(path, device="cuda")
    answer = getattr(module, job["answer"])
    torch.cuda.synchronize()
    load_sec = round(time.time() - t0)
    vram_loaded = torch.cuda.memory_allocated()
    _log(t0, f"読み込み {load_sec}秒・VRAM {vram_loaded / 1024 ** 3:.1f}GB")
    torch.cuda.reset_peak_memory_stats()
    answer(model, processor, job["warmup"])  # 1回目は時間を測らない
    results: dict = {}
    for lang, items in job["requests"].items():
        out = {}
        for task_id, body in items:
            torch.cuda.synchronize()
            started = time.perf_counter()
            resp = answer(model, processor, body)
            torch.cuda.synchronize()
            out[task_id] = {"answers": {qid: _keep(a) for qid, a in (resp.get("answers") or {}).items()},
                            "ms": round((time.perf_counter() - started) * 1000, 1),
                            "input_tokens": (resp.get("usage") or {}).get("input_tokens")}
        results[lang] = out
    _log(t0, "全ての状況を解き終えました")
    return {"results": results, "gpu": gpu, "vram_total": int(torch.cuda.get_device_properties(0).total_memory),
            "vram_loaded": int(vram_loaded), "vram_peak": int(torch.cuda.max_memory_allocated()),
            "load_sec": load_sec, "dtype": "bfloat16", "transformers": str(transformers.__version__),
            "torch": str(torch.__version__), "python": platform.python_version(), "packages": installed_packages(),
            "measured_at": datetime.now(timezone.utc).isoformat()}

Packages installed in the run container:

accelerate==1.15.0
aiohappyeyeballs==2.6.1
aiohttp==3.12.7
aiosignal==1.3.2
annotated-doc==0.0.5
anyio==4.15.1
attrs==25.3.0
cbor2==5.7.0
certifi==2026.7.22
click==8.5.0
cuda-bindings==13.4.3
cuda-pathfinder==1.8.3
cuda-toolkit==13.0.3.0
filelock==4.0.9
frozenlist==1.6.0
fsspec==2026.9.0
grpclib==0.4.8
h11==0.16.0
h2==4.2.0
hf-xet==1.6.0
hpack==4.1.0
httpcore==1.0.9
httpx==0.28.1
huggingface_hub==1.33.0
hyperframe==6.1.0
idna==3.20
Jinja2==3.1.6
markdown-it-py==4.2.0
MarkupSafe==3.0.3
mdurl==0.1.2
mpmath==1.3.0
multidict==6.4.4
networkx==3.7
numpy==2.5.3
nvidia-cublas==13.1.1.3
nvidia-cuda-cupti==13.0.85
nvidia-cuda-nvrtc==13.0.88
nvidia-cuda-runtime==13.0.96
nvidia-cudnn-cu13==9.24.0.43
nvidia-cufft==12.0.0.61
nvidia-cufile==1.15.1.6
nvidia-curand==10.4.0.35
nvidia-cusolver==12.0.4.66
nvidia-cusparse==12.6.3.3
nvidia-cusparselt-cu13==0.8.1
nvidia-nccl-cu13==2.30.7
nvidia-nvjitlink==13.4.92
nvidia-nvshmem-cu13==3.4.5
nvidia-nvtx==13.0.85
packaging==26.3
pillow==12.3.0
pip==25.1.1
propcache==0.3.1
protobuf==6.31.1
psutil==7.2.2
Pygments==2.21.0
PyYAML==6.0.3
regex==2026.9.29
rich==15.0.0
safetensors==0.8.0
setuptools==84.0.0
shellingham==1.5.4
sympy==1.14.0
tokenizers==0.23.2
torch==2.14.1
torchvision==0.29.1
tqdm==4.70.1
transformers==5.18.0
triton==3.8.0
typer==0.27.2
typing_extensions==4.16.0
uv==0.7.19
wheel==0.45.1
yarl==1.20.0

How to Get It

Clef weights and related code are available from the Hugging Face repository under the Apache-2.0 license. Since it is not a gated model, it can be downloaded directly without any special consent procedures.

The distribution format is standard safetensors, and in addition to the core model weights (model-*.safetensors), decision head weights (joint_head.safetensors) and Python scripts (joint_schema_model.py) are included.

As prerequisites for operation, the Python libraries torch (tested with 2.11) and transformers (tested with 5.10.2) are required. When handling images or videos as inputs, the image processing library pillow must also be prepared.

Model acquisition and loading are performed using the snapshot_download function from the huggingface_hub library. Below is an example Python code snippet to download the model and execute inference:

import sys
import torch
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))

If you want to perform inference in a format compatible with the Jev / SystemOne API, you can use the included systemone function to obtain results in a data structure similar to a POST /v1/systemone request.

Related Articles

What to Read Next

Sources

Update History

  • 2026-10-03: Added our own measurements: decision-model evaluation on our fixed tasks.