clef Vision-Language Model: 80GB+ VRAM

October 3, 2026

clef Vision-Language Model: 80GB+ VRAM

At a Glance

Item Value
Repository Cloudflare/clef
Family guide Qwen3.8 guide (9 articles)
Publisher guide Alibaba (Qwen): models and licenses
Published 2026-10-01
License apache-2.0
Formats safetensors
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

Cloudflare has released Clef and its faster counterpart Clef-flash on Hugging Face, open-weight decision models that accept text, JSON, images, and video as inputs and return structured, probability-assigned outputs in a single batch. Built on top of Qwen3.8-27B through post-training, they are optimized specifically for decision-making tasks. Released under the Apache 2.0 license, they can be run and tested in local environments.

Specifications

  • Parameter Count: 27B
  • Architecture: Dense model based on Qwen3.8-27B (equipped with a vision encoder and a joint schema head powered by a small transformer head)
  • Context Length: 64,000 tokens

Performance

Comparisons with major models based on the Decision Index (version 0.2.1) and Workflow evals evaluation tables from the model card are shown below (comparing Clef, Clef-flash, Jev, DiffusionGemma Jev, Kev 9B, and Laya).

→ Scroll horizontally to see all columns

Benchmark Clef Clef-flash Jev DiffusionGemma Jev Kev 9B
BFCL (case exact accuracy) 98.5 98.8 95.8 96.5 94.5
ToolRet (nDCG@10) 69.2 66.4 65.3 61.2 64.3
API-Bank (accuracy) 91.9 93.1 88.2 83.7 56.3
BANKING77 (macro-F1) 94.2 90.9 79.7 74.3 84.8
MMLU (accuracy) 90.3 91.8 91.7 79.3 75.3
GSM8K (accuracy) 80.8 67.3 79.9 50.3 48.7
Median latency (ms) 209.3 38.8 524.1 84.4 51.4

Additionally, the Workflow evals results assessing four business workflows are as follows:

→ Scroll horizontally to see all columns

Workflow Metric Clef Clef-flash Jev
Invoice processing Exact actions 64.7 57.1 61.8
Invoice processing Primary action 86.2 73.3 83.1
Customer service Exact actions 76.3 77.0 76.0
Security incidents Exact actions 62.9 61.7 61.7
Agent trace observability Primary action 68.5 69.8 71.6

According to measurements by the publishers, the Clef model outperforms other models such as Jev on metrics measuring tool use and function selection accuracy, including BFCL, ToolRet, and API-Bank, while holding a distinct advantage in latency (such as median latency). On the other hand, Jev scores higher on certain reasoning and commonsense reasoning metrics such as ANLI, BRIGHT, MMLU-Pro, and BBH, indicating that performance varies by task rather than being universally superior. Meanwhile, Clef-flash successfully reduces latency significantly while maintaining accuracy.

Strengths and Use Cases

Clef is a decision model that accepts text, JSON, images, and video as inputs and returns structured answers with probabilities for specified questions or schemas in a single forward pass, without requiring free-form text generation or output parsing. It is well-suited for use cases such as web domain classification, invoice processing, customer support routing, security incident triage, and agent decision processing. Furthermore, equipped with a vision encoder, it can also classify visual content contained in images and videos. The base model, Qwen3.8-27B, inherently boasts strengths in coding, professional tasks, research, and general long-context agent tasks.

How It Differs from Similar Models

Models based on Qwen3.8-27B include OrcaSAQ-2-27B, which focuses on lightweight deployment via quantization, and Hemmingway-1, which specializes in everyday text writing. While these are primarily aimed at standard text generation and composition, Clef specializes in classification and decision-making, significantly differing in that it returns fast, consistent typed outputs with probabilities without performing extraneous text generation.

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 27.4B parameters

Your VRAM Quantization File size Est. memory needed
80GB class (A100 / H100) BF16 51.0GB 61.1GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Can You Run It Locally?

The publisher distributes this model as safetensors.

License — apache-2.0 (Commercial use allowed): Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.

Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.

How to Get It

Model weights are available from the Hugging Face repository. Provided under the Apache 2.0 license, commercial use and local execution are permitted.

huggingface-cli download Cloudflare/clef

It can be loaded using joint_schema_model.py included in the repository, and is compatible with transformers and torch. Hosted versions are also available via Cloudflare Workers AI.

Related Articles

What to Read Next

Sources