{"id":8824,"date":"2026-10-02T04:12:16","date_gmt":"2026-10-01T19:12:16","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/cloudflare-clef-27b-multimodal-decision-model\/"},"modified":"2026-10-03T00:41:19","modified_gmt":"2026-10-02T15:41:19","slug":"cloudflare-clef-27b-multimodal-decision-model","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-clef-27b-multimodal-decision-model\/","title":{"rendered":"clef Structured Decision-Making Model: 80GB+ VRAM"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/Cloudflare\/clef\">Cloudflare\/clef<\/a><\/td>\n<\/tr>\n<tr>\n<td>Family guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-qwen-qwen3-8-en\/\">Qwen3.8 guide (9 articles)<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-alibaba-en\/\">Alibaba (Qwen): models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-01<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<p><!-- lmw:lab-summary --><\/p>\n<p><strong>What we checked ourselves<\/strong><\/p>\n<ul>\n<li>Gave the 20 decision questions we give every decision model to this model, running the publisher&#8217;s own code on a cloud GPU (NVIDIA A100 80GB PCIe) with its network cut off: 20 correct in Japanese and 19 in English.<\/li>\n<\/ul>\n<p>Details and conditions are in \u201cOur Own Measurements\u201d below.<\/p>\n<p><!-- \/lmw:lab-summary --><\/p>\n<h2>Overview<\/h2>\n<p>Cloudflare has released &#8220;Clef&#8221;, a 27B parameter multimodal decision-making model based on Qwen3.8-27B. This model is designed to accept inputs such as text, JSON, images, video as &#8220;state&#8221;, along with typed &#8220;question schemas&#8221;, and directly output the probabilities of each choice. It is capable of making structured decisions in a single forward pass without requiring free-form text generation or output parsing.<\/p>\n<p>Clef has undergone post-training based on Qwen3.8-27B, and according to Cloudflare&#8217;s blog, it is positioned as a &#8220;decision model&#8221; that reads state and makes schema-based decisions. The API is fully compatible with Jev and SystemOne, making it easy to integrate into practical workflows.<\/p>\n<p>Source: <a href=\"https:\/\/blog.cloudflare.com\/clef-decision-models\">Clef decision models on the Cloudflare blog<\/a><\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li><strong>Parameter Count<\/strong>: 27B<\/li>\n<li><strong>Architecture<\/strong>: Based on Qwen3.5 architecture (Dense model), equipped with a vision encoder<\/li>\n<li><strong>Context Length<\/strong>: 262,144 tokens (native), expandable up to 1,000,000 tokens<\/li>\n<li><strong>Input Formats<\/strong>: Text, JSON, images, video<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Here are the comparison results from the decision task benchmark suite &#8220;Decision Index&#8221; listed in the model card. The table focuses on four columns: Clef, the lightweight version Clef-flash, and existing models Jev and Kev 9B.<\/p>\n<h3>Decision Index<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>Clef<\/th>\n<th>Clef-flash<\/th>\n<th>Jev<\/th>\n<th>Kev 9B<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>BFCL (case exact accuracy)<\/td>\n<td>98.5<\/td>\n<td>98.8<\/td>\n<td>95.8<\/td>\n<td>94.5<\/td>\n<\/tr>\n<tr>\n<td>ToolRet (nDCG@10)<\/td>\n<td>69.2<\/td>\n<td>66.4<\/td>\n<td>65.3<\/td>\n<td>64.3<\/td>\n<\/tr>\n<tr>\n<td>API-Bank (accuracy)<\/td>\n<td>91.9<\/td>\n<td>93.1<\/td>\n<td>88.2<\/td>\n<td>56.3<\/td>\n<\/tr>\n<tr>\n<td>BANKING77 (macro-F1)<\/td>\n<td>94.2<\/td>\n<td>90.9<\/td>\n<td>79.7<\/td>\n<td>84.8<\/td>\n<\/tr>\n<tr>\n<td>CLINC150+OOS (macro-F1)<\/td>\n<td>97.4<\/td>\n<td>66.8<\/td>\n<td>89.3<\/td>\n<td>79.0<\/td>\n<\/tr>\n<tr>\n<td>RouterBench (selected quality)<\/td>\n<td>79.7<\/td>\n<td>79.9<\/td>\n<td>79.9<\/td>\n<td>80.0<\/td>\n<\/tr>\n<tr>\n<td>ANLI (macro-F1)<\/td>\n<td>69.8<\/td>\n<td>59.1<\/td>\n<td>74.8<\/td>\n<td>56.3<\/td>\n<\/tr>\n<tr>\n<td>MMLU (accuracy)<\/td>\n<td>90.3<\/td>\n<td>91.8<\/td>\n<td>91.7<\/td>\n<td>75.3<\/td>\n<\/tr>\n<tr>\n<td>GPQA Diamond (accuracy)<\/td>\n<td>48.0<\/td>\n<td>51.0<\/td>\n<td>78.3<\/td>\n<td>38.8<\/td>\n<\/tr>\n<tr>\n<td>ARC-Challenge (accuracy)<\/td>\n<td>97.7<\/td>\n<td>98.3<\/td>\n<td>97.8<\/td>\n<td>93.7<\/td>\n<\/tr>\n<tr>\n<td>WinoGrande (accuracy)<\/td>\n<td>93.5<\/td>\n<td>97.5<\/td>\n<td>92.0<\/td>\n<td>73.2<\/td>\n<\/tr>\n<tr>\n<td>HellaSwag (accuracy)<\/td>\n<td>98.2<\/td>\n<td>98.6<\/td>\n<td>94.5<\/td>\n<td>81.9<\/td>\n<\/tr>\n<tr>\n<td>GSM8K (accuracy)<\/td>\n<td>80.8<\/td>\n<td>67.3<\/td>\n<td>79.9<\/td>\n<td>48.7<\/td>\n<\/tr>\n<tr>\n<td>CRUXEval (accuracy)<\/td>\n<td>86.7<\/td>\n<td>86.1<\/td>\n<td>73.0<\/td>\n<td>51.2<\/td>\n<\/tr>\n<tr>\n<td>BBH (accuracy)<\/td>\n<td>73.7<\/td>\n<td>68.9<\/td>\n<td>92.9<\/td>\n<td>65.2<\/td>\n<\/tr>\n<tr>\n<td>RAGTruth (hallucination F1)<\/td>\n<td>79.4<\/td>\n<td>35.6<\/td>\n<td>76.5<\/td>\n<td>46.2<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>According to measurements by the publishers, Clef records an extremely high score of 98.5% in BFCL (accuracy of function selection and argument assembly), showing exceptional tool-use capabilities as an agent. It also marks the highest values among the comparison targets in classification tasks such as BANKING77 and CLINC150+OOS. On the other hand, it stays at 48.0% in GPQA Diamond (difficult scientific questions), lagging behind Jev&#8217;s 78.3% in reasoning tasks requiring specialized scientific knowledge. Jev also holds the advantage in BBH (reasoning tasks).<\/p>\n<p>Next are the results of &#8220;Workflow evals&#8221;, which simulate business operations.<\/p>\n<h3>Workflow evals<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Workflow<\/th>\n<th>Metric<\/th>\n<th style=\"text-align: right;\">Clef<\/th>\n<th style=\"text-align: right;\">Clef-flash<\/th>\n<th style=\"text-align: right;\">Jev<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Invoice processing<\/td>\n<td>Exact actions<\/td>\n<td style=\"text-align: right;\">64.7<\/td>\n<td style=\"text-align: right;\">57.1<\/td>\n<td style=\"text-align: right;\">61.8<\/td>\n<\/tr>\n<tr>\n<td>Invoice processing<\/td>\n<td>Primary action<\/td>\n<td style=\"text-align: right;\">86.2<\/td>\n<td style=\"text-align: right;\">73.3<\/td>\n<td style=\"text-align: right;\">83.1<\/td>\n<\/tr>\n<tr>\n<td>Customer service<\/td>\n<td>Exact actions<\/td>\n<td style=\"text-align: right;\">76.3<\/td>\n<td style=\"text-align: right;\">77.0<\/td>\n<td style=\"text-align: right;\">76.0<\/td>\n<\/tr>\n<tr>\n<td>Security incidents<\/td>\n<td>Exact actions<\/td>\n<td style=\"text-align: right;\">62.9<\/td>\n<td style=\"text-align: right;\">61.7<\/td>\n<td style=\"text-align: right;\">61.7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>In practical workflows such as invoice processing and security incident response, Clef demonstrates higher accuracy than the existing Jev. It records a particularly high figure of 86.2% in identifying primary actions for invoice processing.<\/p>\n<p>Additionally, the performance of the base model Qwen3.8-27B itself is extremely high; in publisher benchmarks, it records 61.7% on SWE-bench Pro (real repository fixes) and 90.3% on LiveCodeBench v6 (competitive programming), possessing performance that overwhelms models of the same scale in coding and agent tasks. Clef is a model built upon this powerful foundation with specialized training for structured outputs.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Clef excels at decision-making tasks where it computes and returns probabilities for predefined choices based on a given input &#8220;state&#8221; and typed &#8220;question schemas&#8221;. Because it can take not only text and JSON data but also images and video frames simultaneously as state, it enables classification and judgment based on multimodal information.<\/p>\n<p>Supported formats in the question schema include the <code>noul<\/code> type for boolean judgments, the <code>choice<\/code> type for selecting from multiple options, and the <code>score<\/code> type for graded evaluations. Inside the model, a small transformer head (joint schema head) reads the final hidden states of the backbone Qwen3.8-27B (including the vision encoder), and calculates scores (log-probabilities) for all question choices in a single forward pass.<\/p>\n<p>This mechanism eliminates the need for free-form text generation or complex parsing processes to convert output results into JSON or the like, which were traditionally required with LLMs. Specifically, the following practical use cases are anticipated:<\/p>\n<ul>\n<li><strong>Automatic Invoice and Receipt Judgment<\/strong>: Evaluates payment statuses and amount conditions from image or text-format invoice data.<\/li>\n<li><strong>Customer Support Routing<\/strong>: Determines department assignments (billing, technical support, etc.) and urgency (same-day response, within this week, etc.) based on the body of user inquiries.<\/li>\n<li><strong>Security Incident Classification<\/strong>: Automatically identifies the presence or absence of service outages from incident reports and log messages.<\/li>\n<\/ul>\n<p>Furthermore, because it has full compatibility with the APIs of existing decision systems Jev and SystemOne (<code>POST \/v1\/systemone<\/code>), it can be integrated directly into existing workflows that support these formats.<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<p>Compared to Qwen3.8-27B derivative models previously featured on this site, Clef adopts a clearly distinct approach.<\/p>\n<p><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/orcarouter-orcasaq-2-27b\/\">OrcaSAQ-2-27B Text Generation Model: Example Responses and Required VRAM 16GB+<\/a> is a text model compressed to an average of 3.21-bit using proprietary quantization technology (SAQ2) with the objective of running the original 54GB model on a single 16GB GPU. In contrast, Clef does not compress parameters for lightweighting; instead, it retains the vision encoder in the backbone, adds a structured decision head, and supports multimodal inputs.<\/p>\n<p>Additionally, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/altworld-hemmingway-1\/\">Hemmingway-1 Everyday Text Language Model: Required VRAM 12GB+ \/ GGUF Available<\/a> is a model specialized in directly generating human-written everyday texts such as emails and messages. Clef, on the other hand, does not generate natural language text; it outputs only probability values (probability distributions) for predefined question choices. The major difference is that while Hemmingway-1 aims to create text, Clef specializes in automating classification and decision-making.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 27.4B parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>80GB class (A100 \/ H100)<\/td>\n<td>BF16<\/td>\n<td>51.0GB<\/td>\n<td>61.1GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Not usable in Ollama, LM Studio and llama.cpp yet \u2014 we have found no GGUF build.<\/strong><\/p>\n<p>The publisher ships safetensors only. However, llama.cpp&#8217;s registry does list this architecture, so <strong>conversion to GGUF is possible<\/strong> and the model will run once someone publishes a converted build. 4 converted build(s) from other uploaders exist. Today it can be run with transformers or vLLM, using the memory figures in the table above.<\/p>\n<p><strong>License \u2014 <code>apache-2.0<\/code> (Commercial use allowed):<\/strong> Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we measured ourselves on our server (no GPU) by actually reading and running this model&#8217;s files \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Decision-Model Evaluation on Our Own Tasks<\/h3>\n<p>We gave this model the same 11 situations with 20 questions that we give every decision model (yes\/no, named options and ordered levels; written by us, in Japanese and in English). llama.cpp cannot run this model&#8217;s decision head, so we ran the publisher&#8217;s own code (<code>joint_schema_model.py<\/code> and its SystemOne-compatible function) on a rented cloud GPU, inside a container with its network cut off. Each question states the rule to apply so that it has a single correct answer; our code decides right or wrong (yes when the probability of yes is 0.5 or more; for options and levels, the most probable one). These are not benchmark questions, and the result is not a general score of the model.<\/p>\n<p>Conditions: Modal, NVIDIA A100 80GB PCIe (peak VRAM 51.4GB), BF16, transformers 5.18.0, torch 2.14.1+cu130, publisher revision <code>2f3de3dd85<\/code>. The time per decision is measured on that GPU from passing one situation (1 to 3 questions) to the publisher&#8217;s function until it returns; it is not a guide to the speed of your own GPU and cannot be compared with times measured on our CPU.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Method<\/th>\n<th>Quant<\/th>\n<th>Correct (Japanese)<\/th>\n<th>Correct (English)<\/th>\n<th>Mean probability on the correct answer (ja \/ en)<\/th>\n<th>Time per decision, median (ja \/ en)<\/th>\n<th>Peak memory<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>clef<\/strong><\/td>\n<td><strong><code>joint_schema_model.py<\/code><\/strong><\/td>\n<td><strong><code>BF16<\/code><\/strong><\/td>\n<td><strong>20\/20<\/strong><\/td>\n<td><strong>19\/20<\/strong><\/td>\n<td><strong>0.98 \/ 0.94<\/strong><\/td>\n<td><strong>190ms \/ 189ms (GPU: NVIDIA A100 80GB PCIe)<\/strong><\/td>\n<td><strong>VRAM 51.4GB<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Kev-4B (reference)<\/td>\n<td><code>kev<\/code><\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>18\/20<\/td>\n<td>19\/20<\/td>\n<td>0.87 \/ 0.87<\/td>\n<td>7,616ms \/ 7,409ms<\/td>\n<td>6.3GB<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/clef-flash-specs-performance-and-use-cases\/\">clef-flash<\/a><\/td>\n<td><code>joint_schema_model.py<\/code><\/td>\n<td><code>BF16<\/code><\/td>\n<td>18\/20<\/td>\n<td>18\/20<\/td>\n<td>0.88 \/ 0.88<\/td>\n<td>170ms \/ 167ms (GPU: NVIDIA L4)<\/td>\n<td>VRAM 17.9GB<\/td>\n<\/tr>\n<tr>\n<td>lev (reference)<\/td>\n<td><code>lev<\/code><\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>16\/20<\/td>\n<td>16\/20<\/td>\n<td>0.76 \/ 0.80<\/td>\n<td>14,061ms \/ 13,186ms<\/td>\n<td>6.0GB<\/td>\n<\/tr>\n<tr>\n<td>Laya (reference)<\/td>\n<td><code>laya<\/code><\/td>\n<td><code>Q8_0<\/code><\/td>\n<td>10\/20<\/td>\n<td>15\/20<\/td>\n<td>0.47 \/ 0.68<\/td>\n<td>968ms \/ 678ms<\/td>\n<td>0.9GB<\/td>\n<\/tr>\n<tr>\n<td>Julia-1 (reference)<\/td>\n<td><code>laya<\/code><\/td>\n<td><code>Q8_0<\/code><\/td>\n<td>13\/20<\/td>\n<td>10\/20<\/td>\n<td>0.60 \/ 0.45<\/td>\n<td>136ms \/ 107ms<\/td>\n<td>0.5GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>In Japanese it got 1 more right than in English (20 vs 19 of 20).<\/p>\n<p>Other rows are other decision models measured the same way (\u201creference\u201d rows are models we measure for comparison). Methods differ by model, and each model is measured with its own quantization.<\/p>\n<p>Times marked (GPU) were measured on a cloud GPU and the others on our CPU, so the times cannot be compared across those rows; the numbers of correct answers can.<\/p>\n<h4>Answer to Each Question<\/h4>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Situation<\/th>\n<th>Question<\/th>\n<th>Correct answer<\/th>\n<th>Japanese<\/th>\n<th>English<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Routing a support message<\/td>\n<td>Which team should handle this message?<\/td>\n<td><code>billing<\/code><\/td>\n<td>\u2713 <code>billing<\/code> (p=0.99)<\/td>\n<td>\u2713 <code>billing<\/code> (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Routing a support message<\/td>\n<td>Is the customer asking for money back?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Routing a support message<\/td>\n<td>Does the message report that a service is down?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (within the window)<\/td>\n<td>Is today within 30 days of the delivery date?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<td>\u2713 yes (p=0.98)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (within the window)<\/td>\n<td>Can this item be returned under the policy?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<td>\u2713 yes (p=0.94)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (past the window)<\/td>\n<td>Is today within 30 days of the delivery date?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<td>\u2713 no (p=1.00)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (past the window)<\/td>\n<td>Can this item be returned under the policy?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (vendor)<\/td>\n<td>What should happen to this invoice?<\/td>\n<td><code>reject<\/code><\/td>\n<td>\u2713 <code>reject<\/code> (p=0.99)<\/td>\n<td>\u2713 <code>reject<\/code> (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (vendor)<\/td>\n<td>Is the invoice total above 1,000 USD?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (amount)<\/td>\n<td>What should happen to this invoice?<\/td>\n<td><code>manager<\/code><\/td>\n<td>\u2713 <code>manager<\/code> (p=0.98)<\/td>\n<td>\u2713 <code>manager<\/code> (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (amount)<\/td>\n<td>Is the invoice total above 1,000 USD?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Incident severity<\/td>\n<td>How widespread is the impact of this incident?<\/td>\n<td>3: All users affected<\/td>\n<td>\u2713 3: All users affected (p=0.94)<\/td>\n<td>\u2713 3: All users affected (p=0.95)<\/td>\n<\/tr>\n<tr>\n<td>Incident severity<\/td>\n<td>Is the service down?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.98)<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Delivery delay level<\/td>\n<td>Which level of the guideline does this delay fall into?<\/td>\n<td>2: Moderate delay<\/td>\n<td>\u2713 2: Moderate delay (p=0.90)<\/td>\n<td>\u2713 2: Moderate delay (p=0.93)<\/td>\n<\/tr>\n<tr>\n<td>Review opinion (negation)<\/td>\n<td>What is the reviewer&#8217;s overall opinion?<\/td>\n<td><code>positive<\/code><\/td>\n<td>\u2713 <code>positive<\/code> (p=0.98)<\/td>\n<td>\u2713 <code>positive<\/code> (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Review opinion (negation)<\/td>\n<td>Does the reviewer say they would buy it again?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Suspicious email (injected instruction)<\/td>\n<td>How should this email be classified?<\/td>\n<td><code>phishing<\/code><\/td>\n<td>\u2713 <code>phishing<\/code> (p=0.98)<\/td>\n<td>\u2713 <code>phishing<\/code> (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Suspicious email (injected instruction)<\/td>\n<td>Does the email ask the reader to enter a password?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<td>\u2713 yes (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Schedule overlap<\/td>\n<td>Does the meeting request overlap with an event in the calendar?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.96)<\/td>\n<td>\u2717 no (p=0.21)<\/td>\n<\/tr>\n<tr>\n<td>Message intent (6 options)<\/td>\n<td>What does the user want to do?<\/td>\n<td><code>change_address<\/code><\/td>\n<td>\u2713 <code>change_address<\/code> (p=0.99)<\/td>\n<td>\u2713 <code>change_address<\/code> (p=0.99)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The full text of every situation and question, with the reason for each correct answer, is on <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h4>Setup and Steps We Ran<\/h4>\n<p>Download (a separate container without a GPU, which does not run the publisher&#8217;s code): Python 3.12, <code>pip install huggingface_hub<\/code>, then these files of <code>Cloudflare\/clef<\/code> at revision <code>2f3de3dd85f379784083b0814d997ab627200f0c<\/code> (no pickle weights):<\/p>\n<pre><code class=\"language-text\">chat_template.jinja\nconfig.json\ngeneration_config.json\njoint_head.safetensors\njoint_head_config.json\njoint_schema_model.py\nmodel-00001-of-00012.safetensors\nmodel-00002-of-00012.safetensors\nmodel-00003-of-00012.safetensors\nmodel-00004-of-00012.safetensors\nmodel-00005-of-00012.safetensors\nmodel-00006-of-00012.safetensors\nmodel-00007-of-00012.safetensors\nmodel-00008-of-00012.safetensors\nmodel-00009-of-00012.safetensors\nmodel-00010-of-00012.safetensors\nmodel-00011-of-00012.safetensors\nmodel-00012-of-00012.safetensors\nmodel.safetensors.index.json\nprocessor_config.json\ntokenizer.json\ntokenizer_config.json\n<\/code><\/pre>\n<p>Run (GPU container, network blocked, model files read-only): Python 3.12, <code>pip install torch transformers accelerate safetensors pillow torchvision<\/code>. The full code we ran in the container (<code>decision_remote.py<\/code>); the questions are the same request bodies as above:<\/p>\n<pre><code class=\"language-python\">&quot;&quot;&quot;\nModal \u306e\u30b3\u30f3\u30c6\u30ca\u306e\u4e2d\u3067\u52d5\u304f\u3001\u610f\u601d\u6c7a\u5b9a\u30e2\u30c7\u30eb\u306e\u914d\u5e03\u5143\u306e\u30b3\u30fc\u30c9\u306b\u3088\u308b\u8a55\u4fa1(2026-10-02\u3001\u30e6\u30fc\u30b6\u30fc\u306e\u5224\u65ad\u300cClef \u3092 Modal \u3067\u52d5\u304b\u3059\u300d)\u3002\nlmw\/lab\/gpu_modal.py \u306e decision_run \u304b\u3089\u547c\u3070\u308c\u308b\u3002**\u6a19\u6e96\u30e9\u30a4\u30d6\u30e9\u30ea\u3060\u3051\u3092\u5148\u982d\u3067\u8aad\u307f**\u3001torch\u30fbtransformers \u306f\u95a2\u6570\u306e\u4e2d\u3067\u8aad\u3080\u3002\n\nllama.cpp \u306b\u5224\u5b9a\u306e\u65b9\u5f0f\u304c\u7121\u3044\u610f\u601d\u6c7a\u5b9a\u30e2\u30c7\u30eb(Clef \u306f\u5224\u5b9a\u30d8\u30c3\u30c9 joint_head.safetensors \u3092\u914d\u5e03\u5143\u306e joint_schema_model.py \u3067\n\u8aad\u3080)\u3092\u3001\u914d\u5e03\u5143\u306e SystemOne \u4e92\u63db\u306e\u95a2\u6570(systemone)\u3067\u52d5\u304b\u3059\u3002\u3053\u306e\u30b3\u30f3\u30c6\u30ca\u306f**\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u906e\u65ad\u30fbModal \u306e\u6a5f\u80fd\u306a\u3057\u30fb\nVolume \u306f\u8aad\u307f\u53d6\u308a\u5c02\u7528**\u3067\u3001\u914d\u5e03\u5143\u306e\u30b3\u30fc\u30c9\u306f\u3053\u3053\u3067\u3060\u3051 import \u3059\u308b(\u30c0\u30a6\u30f3\u30ed\u30fc\u30c9\u306f remote_code_remote.download \u304c\nGPU \u306e\u7121\u3044\u5225\u306e\u30b3\u30f3\u30c6\u30ca\u3067\u884c\u3044\u3001\u305d\u3061\u3089\u3067\u306f import \u3057\u306a\u3044)\u3002\n\n\u554f\u984c(\u8981\u6c42\u306e\u672c\u6587)\u306f\u30db\u30b9\u30c8\u304c\u7d44\u3093\u3067\u6e21\u3059(lmw\/lab\/decision.py \u306e request_body\u3002llama.cpp \u3067\u6e2c\u308b\u30e2\u30c7\u30eb\u3068\u540c\u3058\u3082\u306e)\u3002\n\u6b63\u8aa4\u306f\u30db\u30b9\u30c8\u306e\u30b3\u30fc\u30c9(decision.score)\u304c\u5224\u5b9a\u3059\u308b\u3002\u3053\u3053\u3067\u306f\u5fdc\u7b54\u3092\u305d\u306e\u307e\u307e\u8fd4\u3057\u30011\u56de\u306e\u5224\u5b9a\u306b\u304b\u304b\u3063\u305f\u6642\u9593\u3092\u6e2c\u308b\u3060\u3051\u3002\n&quot;&quot;&quot;\nimport os\nimport platform\nimport sys\nimport time\nfrom datetime import datetime, timezone\n\nMOUNT = &quot;\/models&quot;\nPYTHON_VERSION = &quot;3.12&quot;\n# \u5b9f\u884c\u7528\u306e\u30b3\u30f3\u30c6\u30ca\u306b\u5165\u308c\u308b\u30d1\u30c3\u30b1\u30fc\u30b8(\u8a18\u4e8b\u306e\u624b\u9806\u306b\u3082\u540c\u3058\u5024\u3092\u8f09\u305b\u308b)\u3002Clef \u306e\u30ab\u30fc\u30c9\u306e\u8a18\u8f09\u306f torch 2.11\u30fb\n# transformers 5.10.2 \u3067\u3001\u753b\u50cf\u3092\u8aad\u3080\u51e6\u7406(AutoProcessor)\u306e\u305f\u3081\u306b pillow\u30fbtorchvision \u3092\u5165\u308c\u308b\nRUN_PACKAGES = (&quot;torch&quot;, &quot;transformers&quot;, &quot;accelerate&quot;, &quot;safetensors&quot;, &quot;pillow&quot;, &quot;torchvision&quot;)\n\ndef installed_packages() -&gt; list[str]:\n    from importlib import metadata\n    seen = {}\n    for dist in metadata.distributions():\n        name = dist.metadata[&quot;Name&quot;]\n        if name and name.lower() not in seen:\n            seen[name.lower()] = f&quot;{name}=={dist.version}&quot;\n    return sorted(seen.values(), key=str.lower)\n\ndef _keep(answer: dict) -&gt; dict:\n    return {k: answer[k] for k in (&quot;type&quot;, &quot;choice&quot;, &quot;noul&quot;, &quot;score&quot;, &quot;probabilities&quot;, &quot;confidence&quot;) if k in answer}\n\ndef _log(t0: float, message: str) -&gt; None:\n    print(f&quot;[decision {time.time() - t0:6.0f}s] {message}&quot;, flush=True)\n\ndef run(job: dict) -&gt; dict:\n    &quot;&quot;&quot;\n    job: {&quot;dir&quot;, &quot;module&quot;, &quot;loader&quot;, &quot;answer&quot;, &quot;warmup&quot;: \u8981\u6c42, &quot;requests&quot;: {lang: [[\u72b6\u6cc1\u306e id, \u8981\u6c42], ...]}}\u3002\n    \u623b\u308a\u5024\u306f\u8a18\u9332\u306e\u4e00\u90e8({&quot;results&quot;, &quot;gpu&quot;, &quot;vram_*&quot;, &quot;transformers&quot;, &quot;torch&quot;, &quot;packages&quot;, ...})\u3002\n    &quot;&quot;&quot;\n    import torch\n    import transformers\n    os.environ[&quot;HF_HUB_OFFLINE&quot;] = &quot;1&quot;\n    os.environ[&quot;TRANSFORMERS_OFFLINE&quot;] = &quot;1&quot;\n    t0 = time.time()\n    path = os.path.join(MOUNT, job[&quot;dir&quot;])\n    gpu = torch.cuda.get_device_name(0)\n    _log(t0, f&quot;GPU: {gpu}\u30fbtransformers {transformers.__version__}\u30fbtorch {torch.__version__}&quot;)\n    # \u914d\u5e03\u5143\u306e\u30b3\u30fc\u30c9\u3002\u3053\u306e\u30b3\u30f3\u30c6\u30ca(\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u906e\u65ad)\u3067\u3060\u3051\u8aad\u3080\n    sys.path.insert(0, path)\n    module = __import__(job[&quot;module&quot;])\n    model, processor = getattr(module, job[&quot;loader&quot;])(path, device=&quot;cuda&quot;)\n    answer = getattr(module, job[&quot;answer&quot;])\n    torch.cuda.synchronize()\n    load_sec = round(time.time() - t0)\n    vram_loaded = torch.cuda.memory_allocated()\n    _log(t0, f&quot;\u8aad\u307f\u8fbc\u307f {load_sec}\u79d2\u30fbVRAM {vram_loaded \/ 1024 ** 3:.1f}GB&quot;)\n    torch.cuda.reset_peak_memory_stats()\n    answer(model, processor, job[&quot;warmup&quot;])  # 1\u56de\u76ee\u306f\u6642\u9593\u3092\u6e2c\u3089\u306a\u3044\n    results: dict = {}\n    for lang, items in job[&quot;requests&quot;].items():\n        out = {}\n        for task_id, body in items:\n            torch.cuda.synchronize()\n            started = time.perf_counter()\n            resp = answer(model, processor, body)\n            torch.cuda.synchronize()\n            out[task_id] = {&quot;answers&quot;: {qid: _keep(a) for qid, a in (resp.get(&quot;answers&quot;) or {}).items()},\n                            &quot;ms&quot;: round((time.perf_counter() - started) * 1000, 1),\n                            &quot;input_tokens&quot;: (resp.get(&quot;usage&quot;) or {}).get(&quot;input_tokens&quot;)}\n        results[lang] = out\n    _log(t0, &quot;\u5168\u3066\u306e\u72b6\u6cc1\u3092\u89e3\u304d\u7d42\u3048\u307e\u3057\u305f&quot;)\n    return {&quot;results&quot;: results, &quot;gpu&quot;: gpu, &quot;vram_total&quot;: int(torch.cuda.get_device_properties(0).total_memory),\n            &quot;vram_loaded&quot;: int(vram_loaded), &quot;vram_peak&quot;: int(torch.cuda.max_memory_allocated()),\n            &quot;load_sec&quot;: load_sec, &quot;dtype&quot;: &quot;bfloat16&quot;, &quot;transformers&quot;: str(transformers.__version__),\n            &quot;torch&quot;: str(torch.__version__), &quot;python&quot;: platform.python_version(), &quot;packages&quot;: installed_packages(),\n            &quot;measured_at&quot;: datetime.now(timezone.utc).isoformat()}\n<\/code><\/pre>\n<p>Packages installed in the run container:<\/p>\n<pre><code class=\"language-text\">accelerate==1.15.0\naiohappyeyeballs==2.6.1\naiohttp==3.12.7\naiosignal==1.3.2\nannotated-doc==0.0.5\nanyio==4.15.1\nattrs==25.3.0\ncbor2==5.7.0\ncertifi==2026.7.22\nclick==8.5.0\ncuda-bindings==13.4.3\ncuda-pathfinder==1.8.3\ncuda-toolkit==13.0.3.0\nfilelock==4.0.9\nfrozenlist==1.6.0\nfsspec==2026.9.0\ngrpclib==0.4.8\nh11==0.16.0\nh2==4.2.0\nhf-xet==1.6.0\nhpack==4.1.0\nhttpcore==1.0.9\nhttpx==0.28.1\nhuggingface_hub==1.33.0\nhyperframe==6.1.0\nidna==3.20\nJinja2==3.1.6\nmarkdown-it-py==4.2.0\nMarkupSafe==3.0.3\nmdurl==0.1.2\nmpmath==1.3.0\nmultidict==6.4.4\nnetworkx==3.7\nnumpy==2.5.3\nnvidia-cublas==13.1.1.3\nnvidia-cuda-cupti==13.0.85\nnvidia-cuda-nvrtc==13.0.88\nnvidia-cuda-runtime==13.0.96\nnvidia-cudnn-cu13==9.24.0.43\nnvidia-cufft==12.0.0.61\nnvidia-cufile==1.15.1.6\nnvidia-curand==10.4.0.35\nnvidia-cusolver==12.0.4.66\nnvidia-cusparse==12.6.3.3\nnvidia-cusparselt-cu13==0.8.1\nnvidia-nccl-cu13==2.30.7\nnvidia-nvjitlink==13.4.92\nnvidia-nvshmem-cu13==3.4.5\nnvidia-nvtx==13.0.85\npackaging==26.3\npillow==12.3.0\npip==25.1.1\npropcache==0.3.1\nprotobuf==6.31.1\npsutil==7.2.2\nPygments==2.21.0\nPyYAML==6.0.3\nregex==2026.9.29\nrich==15.0.0\nsafetensors==0.8.0\nsetuptools==84.0.0\nshellingham==1.5.4\nsympy==1.14.0\ntokenizers==0.23.2\ntorch==2.14.1\ntorchvision==0.29.1\ntqdm==4.70.1\ntransformers==5.18.0\ntriton==3.8.0\ntyper==0.27.2\ntyping_extensions==4.16.0\nuv==0.7.19\nwheel==0.45.1\nyarl==1.20.0\n<\/code><\/pre>\n<p><!-- \/lmw:lab --><\/p>\n<h2>How to Get It<\/h2>\n<p>Clef weights and related code are available from the Hugging Face repository under the Apache-2.0 license. Since it is not a gated model, it can be downloaded directly without any special consent procedures.<\/p>\n<p>The distribution format is standard <code>safetensors<\/code>, and in addition to the core model weights (<code>model-*.safetensors<\/code>), decision head weights (<code>joint_head.safetensors<\/code>) and Python scripts (<code>joint_schema_model.py<\/code>) are included.<\/p>\n<p>As prerequisites for operation, the Python libraries <code>torch<\/code> (tested with 2.11) and <code>transformers<\/code> (tested with 5.10.2) are required. When handling images or videos as inputs, the image processing library <code>pillow<\/code> must also be prepared.<\/p>\n<p>Model acquisition and loading are performed using the <code>snapshot_download<\/code> function from the <code>huggingface_hub<\/code> library. Below is an example Python code snippet to download the model and execute inference:<\/p>\n<pre><code class=\"language-python\">import sys\nimport torch\nfrom huggingface_hub import snapshot_download\n\npath = snapshot_download(&quot;Cloudflare\/clef&quot;)\nsys.path.insert(0, path)\nfrom joint_schema_model import collate_records, encode_record, load_release_model\n\nmodel, processor = load_release_model(path, device=&quot;cuda&quot;)\n\nrecord = {\n    &quot;state&quot;: {&quot;invoice&quot;: {&quot;vendor&quot;: &quot;Acme&quot;, &quot;total&quot;: 1250.0, &quot;currency&quot;: &quot;USD&quot;, &quot;status&quot;: &quot;overdue&quot;}},\n    &quot;questions&quot;: {\n        &quot;status&quot;: {\n            &quot;type&quot;: &quot;choice&quot;,\n            &quot;instructions&quot;: &quot;What is the invoice status?&quot;,\n            &quot;criteria&quot;: {&quot;paid&quot;: &quot;Invoice is paid.&quot;, &quot;overdue&quot;: &quot;Invoice is past due.&quot;, &quot;draft&quot;: &quot;Not sent.&quot;},\n        },\n        &quot;large&quot;: {&quot;type&quot;: &quot;noul&quot;, &quot;instructions&quot;: &quot;Is the total above 1000 USD?&quot;},\n    },\n}\n\nencoded = encode_record(processor.tokenizer, record, processor=processor)\nbatch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device(&quot;cuda&quot;))\nwith torch.inference_mode():\n    logits = model(batch)[0]\n\nfor question, question_logits in zip(encoded.questions, logits):\n    probabilities = question_logits.float().softmax(-1).tolist()\n    print(question.question_id, dict(zip(question.option_ids, probabilities)))\n<\/code><\/pre>\n<p>If you want to perform inference in a format compatible with the Jev \/ SystemOne API, you can use the included <code>systemone<\/code> function to obtain results in a data structure similar to a <code>POST \/v1\/systemone<\/code> request.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-releases-clef-and-clef-flash\/\">clef Vision-Language Model: 80GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B Vision-Language Model: Our Test Answers, 8GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf\/\">Ternary-Bonsai-2-27B-gguf Text Generation Model: 8GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/orcasaq2-27b-qwen3-8-27b-quantized\/\">OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model needs at least 80GB) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a><\/li>\n<li><strong>Explore the same model family<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-qwen-qwen3-8-en\/\">Qwen3.8 family overview (9 articles, 11 converted builds)<\/a><\/li>\n<li><strong>Engines that run this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<li><strong>Formats this model is available in<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">Safetensors format guide and models<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-alibaba-en\/\">Alibaba (Qwen): models, licenses and articles<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/Cloudflare\/clef\">https:\/\/huggingface.co\/Cloudflare\/clef<\/a><\/li>\n<li><a href=\"https:\/\/blog.cloudflare.com\/clef-decision-models\">https:\/\/blog.cloudflare.com\/clef-decision-models<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-10-03: Added our own measurements: decision-model evaluation on our fixed tasks.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Cloudflare has released Clef, a 27B multimodal decision-making model based on Qwen3.8-27B. Learn about its specs, performance, and how to get it.<\/p>\n","protected":false},"author":1,"featured_media":8823,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[2808,2810,520,522,1547,1952,896,2812],"class_list":["post-8824","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-clef-en","tag-cloudflare-en","tag-qwen-en","tag-qwen3-8-en","tag-verified","tag-vlm-en","tag--en"],"lang":"en","translations":{"en":8824,"ja":8822},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8824","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=8824"}],"version-history":[{"count":5,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8824\/revisions"}],"predecessor-version":[{"id":9318,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8824\/revisions\/9318"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/8823"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=8824"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=8824"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=8824"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}