{"id":8939,"date":"2026-10-02T08:17:05","date_gmt":"2026-10-01T23:17:05","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/clef-flash-specs-performance-and-use-cases\/"},"modified":"2026-10-03T00:41:21","modified_gmt":"2026-10-02T15:41:21","slug":"clef-flash-specs-performance-and-use-cases","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/clef-flash-specs-performance-and-use-cases\/","title":{"rendered":"clef-flash Structured Output Optimized Model: 4GB+ VRAM, GGUF Builds"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/Cloudflare\/clef-flash\">Cloudflare\/clef-flash<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-alibaba-en\/\">Alibaba (Qwen): models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-01<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<p><!-- lmw:lab-summary --><\/p>\n<p><strong>What we checked ourselves<\/strong><\/p>\n<ul>\n<li>Gave the 20 decision questions we give every decision model to this model, running the publisher&#8217;s own code on a cloud GPU (NVIDIA L4) with its network cut off: 18 correct in Japanese and 18 in English.<\/li>\n<\/ul>\n<p>Details and conditions are in \u201cOur Own Measurements\u201d below.<\/p>\n<p><!-- \/lmw:lab-summary --><\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameters: 9B<\/li>\n<li>Architecture: Qwen3.5-9B (Gated DeltaNet + Mixture of Experts: MoE) with a vision encoder and a custom Joint schema head<\/li>\n<li>Context Length: Native 262,144 tokens (can be extended up to 1,010,000 tokens using techniques such as YaRN)<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Below are internal measurement results conducted by the publisher using the &#8220;Decision Index 0.2.1&#8221; suite. For comparison, figures for the higher-tier model Clef, as well as Jev and Kev 9B, are excerpted and listed.<\/p>\n<h3>Decision Index Measurement Results<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>Clef<\/th>\n<th>Clef-flash<\/th>\n<th>Jev<\/th>\n<th>Kev 9B<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>BFCL (case exact accuracy)<\/td>\n<td>98.5<\/td>\n<td>98.8<\/td>\n<td>95.8<\/td>\n<td>94.5<\/td>\n<\/tr>\n<tr>\n<td>ToolRet (nDCG@10)<\/td>\n<td>69.2<\/td>\n<td>66.4<\/td>\n<td>65.3<\/td>\n<td>64.3<\/td>\n<\/tr>\n<tr>\n<td>API-Bank (accuracy)<\/td>\n<td>91.9<\/td>\n<td>93.1<\/td>\n<td>88.2<\/td>\n<td>56.3<\/td>\n<\/tr>\n<tr>\n<td>BANKING77 (macro-F1)<\/td>\n<td>94.2<\/td>\n<td>90.9<\/td>\n<td>79.7<\/td>\n<td>84.8<\/td>\n<\/tr>\n<tr>\n<td>CLINC150+OOS (macro-F1)<\/td>\n<td>97.4<\/td>\n<td>66.8<\/td>\n<td>89.3<\/td>\n<td>79.0<\/td>\n<\/tr>\n<tr>\n<td>RouterBench (selected quality)<\/td>\n<td>79.7<\/td>\n<td>79.9<\/td>\n<td>79.9<\/td>\n<td>80.0<\/td>\n<\/tr>\n<tr>\n<td>Home appliance simulator (case exact accuracy)<\/td>\n<td>83.0<\/td>\n<td>97.7<\/td>\n<td>52.3<\/td>\n<td>25.0<\/td>\n<\/tr>\n<tr>\n<td>SGD\/SGD-X (macro-F1)<\/td>\n<td>43.8<\/td>\n<td>34.2<\/td>\n<td>43.0<\/td>\n<td>64.0<\/td>\n<\/tr>\n<tr>\n<td>ContractNLI (macro-F1)<\/td>\n<td>81.4<\/td>\n<td>84.3<\/td>\n<td>71.7<\/td>\n<td>57.8<\/td>\n<\/tr>\n<tr>\n<td>ANLI (macro-F1)<\/td>\n<td>69.8<\/td>\n<td>59.1<\/td>\n<td>74.8<\/td>\n<td>56.3<\/td>\n<\/tr>\n<tr>\n<td>BPoMP (accuracy)<\/td>\n<td>96.9<\/td>\n<td>95.4<\/td>\n<td>90.6<\/td>\n<td>67.0<\/td>\n<\/tr>\n<tr>\n<td>Humicroedit (accuracy)<\/td>\n<td>66.7<\/td>\n<td>75.1<\/td>\n<td>61.9<\/td>\n<td>55.8<\/td>\n<\/tr>\n<tr>\n<td>POP909-CL (accuracy)<\/td>\n<td>15.8<\/td>\n<td>1.6<\/td>\n<td>18.1<\/td>\n<td>10.8<\/td>\n<\/tr>\n<tr>\n<td>cfcolor (accuracy)<\/td>\n<td>66.0<\/td>\n<td>65.8<\/td>\n<td>64.7<\/td>\n<td>56.3<\/td>\n<\/tr>\n<tr>\n<td>MMLU (accuracy)<\/td>\n<td>90.3<\/td>\n<td>91.8<\/td>\n<td>91.7<\/td>\n<td>75.3<\/td>\n<\/tr>\n<tr>\n<td>GPQA Diamond (accuracy)<\/td>\n<td>48.0<\/td>\n<td>51.0<\/td>\n<td>78.3<\/td>\n<td>38.8<\/td>\n<\/tr>\n<tr>\n<td>ARC-Easy (accuracy)<\/td>\n<td>99.0<\/td>\n<td>99.5<\/td>\n<td>99.3<\/td>\n<td>97.7<\/td>\n<\/tr>\n<tr>\n<td>ARC-Challenge (accuracy)<\/td>\n<td>97.7<\/td>\n<td>98.3<\/td>\n<td>97.8<\/td>\n<td>93.7<\/td>\n<\/tr>\n<tr>\n<td>WinoGrande (accuracy)<\/td>\n<td>93.5<\/td>\n<td>97.5<\/td>\n<td>92.0<\/td>\n<td>73.2<\/td>\n<\/tr>\n<tr>\n<td>HellaSwag (accuracy)<\/td>\n<td>98.2<\/td>\n<td>98.6<\/td>\n<td>94.5<\/td>\n<td>81.9<\/td>\n<\/tr>\n<tr>\n<td>GSM8K (accuracy)<\/td>\n<td>80.8<\/td>\n<td>67.3<\/td>\n<td>79.9<\/td>\n<td>48.7<\/td>\n<\/tr>\n<tr>\n<td>ChessBench (accuracy)<\/td>\n<td>24.7<\/td>\n<td>23.0<\/td>\n<td>17.2<\/td>\n<td>11.2<\/td>\n<\/tr>\n<tr>\n<td>MuSR (accuracy)<\/td>\n<td>83.5<\/td>\n<td>86.0<\/td>\n<td>66.1<\/td>\n<td>57.9<\/td>\n<\/tr>\n<tr>\n<td>SATA-Bench (case exact accuracy)<\/td>\n<td>33.8<\/td>\n<td>36.7<\/td>\n<td>26.4<\/td>\n<td>26.7<\/td>\n<\/tr>\n<tr>\n<td>BRIGHT (nDCG@10)<\/td>\n<td>45.9<\/td>\n<td>39.3<\/td>\n<td>47.5<\/td>\n<td>38.5<\/td>\n<\/tr>\n<tr>\n<td>Amazon ESCI (macro-F1)<\/td>\n<td>57.5<\/td>\n<td>57.4<\/td>\n<td>55.2<\/td>\n<td>49.2<\/td>\n<\/tr>\n<tr>\n<td>ACOS (per-review F1)<\/td>\n<td>33.3<\/td>\n<td>25.9<\/td>\n<td>29.5<\/td>\n<td>18.3<\/td>\n<\/tr>\n<tr>\n<td>FinEntity (macro-F1)<\/td>\n<td>96.2<\/td>\n<td>97.1<\/td>\n<td>87.0<\/td>\n<td>88.4<\/td>\n<\/tr>\n<tr>\n<td>VAST (macro-F1)<\/td>\n<td>59.5<\/td>\n<td>49.6<\/td>\n<td>64.6<\/td>\n<td>55.4<\/td>\n<\/tr>\n<tr>\n<td>NLI4CT (macro-F1)<\/td>\n<td>82.9<\/td>\n<td>78.6<\/td>\n<td>84.1<\/td>\n<td>74.9<\/td>\n<\/tr>\n<tr>\n<td>CRUXEval (accuracy)<\/td>\n<td>86.7<\/td>\n<td>86.1<\/td>\n<td>73.0<\/td>\n<td>51.2<\/td>\n<\/tr>\n<tr>\n<td>CLadder (accuracy)<\/td>\n<td>94.0<\/td>\n<td>97.7<\/td>\n<td>72.6<\/td>\n<td>62.0<\/td>\n<\/tr>\n<tr>\n<td>ForecastBench (Brier, lower is better)<\/td>\n<td>13.9<\/td>\n<td>10.6<\/td>\n<td>17.4<\/td>\n<td>17.6<\/td>\n<\/tr>\n<tr>\n<td>Habermas Machine (accuracy)<\/td>\n<td>68.7<\/td>\n<td>71.8<\/td>\n<td>45.9<\/td>\n<td>39.4<\/td>\n<\/tr>\n<tr>\n<td>PhishNChips (accuracy)<\/td>\n<td>79.6<\/td>\n<td>75.0<\/td>\n<td>62.5<\/td>\n<td>50.7<\/td>\n<\/tr>\n<tr>\n<td>MMLU-Pro (accuracy)<\/td>\n<td>65.9<\/td>\n<td>65.3<\/td>\n<td>82.7<\/td>\n<td>51.1<\/td>\n<\/tr>\n<tr>\n<td>BBH (accuracy)<\/td>\n<td>73.7<\/td>\n<td>68.9<\/td>\n<td>92.9<\/td>\n<td>65.2<\/td>\n<\/tr>\n<tr>\n<td>RAGTruth (hallucination F1)<\/td>\n<td>79.4<\/td>\n<td>35.6<\/td>\n<td>76.5<\/td>\n<td>46.2<\/td>\n<\/tr>\n<tr>\n<td>HoVer (accuracy)<\/td>\n<td>65.2<\/td>\n<td>61.2<\/td>\n<td>72.9<\/td>\n<td>58.8<\/td>\n<\/tr>\n<tr>\n<td>When2Call MCQ (accuracy)<\/td>\n<td>72.4<\/td>\n<td>65.6<\/td>\n<td>81.0<\/td>\n<td>49.6<\/td>\n<\/tr>\n<tr>\n<td>New Yorker (accuracy)<\/td>\n<td>69.5<\/td>\n<td>66.1<\/td>\n<td>70.1<\/td>\n<td>58.1<\/td>\n<\/tr>\n<tr>\n<td>Median latency (ms)<\/td>\n<td>209.3<\/td>\n<td>38.8<\/td>\n<td>524.1<\/td>\n<td>51.4<\/td>\n<\/tr>\n<tr>\n<td>p95 latency (ms)<\/td>\n<td>238.6<\/td>\n<td>122.4<\/td>\n<td>536.0<\/td>\n<td>187.9<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>This table shows that Clef-Flash achieves an extremely high accuracy of 98.8% on BFCL, which measures tool-use capability as an agent, slightly outperforming the higher-tier Clef model. It also maintains a high standard of 91.8% on MMLU for general knowledge. Notably, it has low latency, with a p95 latency of 122.4ms, making it significantly faster than Clef (238.6ms) and Jev (536.0ms).<\/p>\n<p>On the other hand, its RAGTruth score (measuring hallucinations) remains at 35.6%, which is significantly lower compared to Clef (79.4%) and Jev (76.5%). Additionally, in metrics such as intent classification (CLINC150+OOS), BBH, and MMLU-Pro which require reasoning, there are areas where it falls short of other models like Jev. While it excels in decision-making speed and specific tool-use accuracy, there are trade-offs regarding complex reasoning and hallucination suppression.<\/p>\n<h3>Workflow Evaluation<\/h3>\n<p>Accuracy measurement results from &#8220;Typesafe Evals&#8221;, simulating actual business workflows, are as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Workflow<\/th>\n<th>Metric<\/th>\n<th style=\"text-align: right;\">Clef<\/th>\n<th style=\"text-align: right;\">Clef-flash<\/th>\n<th style=\"text-align: right;\">Jev<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Invoice processing<\/td>\n<td>Exact actions<\/td>\n<td style=\"text-align: right;\">64.7<\/td>\n<td style=\"text-align: right;\">57.1<\/td>\n<td style=\"text-align: right;\">61.8<\/td>\n<\/tr>\n<tr>\n<td>Invoice processing<\/td>\n<td>Primary action<\/td>\n<td style=\"text-align: right;\">86.2<\/td>\n<td style=\"text-align: right;\">73.3<\/td>\n<td style=\"text-align: right;\">83.1<\/td>\n<\/tr>\n<tr>\n<td>Customer service<\/td>\n<td>Exact actions<\/td>\n<td style=\"text-align: right;\">76.3<\/td>\n<td style=\"text-align: right;\">77.0<\/td>\n<td style=\"text-align: right;\">76.0<\/td>\n<\/tr>\n<tr>\n<td>Security incidents<\/td>\n<td>Exact actions<\/td>\n<td style=\"text-align: right;\">62.9<\/td>\n<td style=\"text-align: right;\">61.7<\/td>\n<td style=\"text-align: right;\">61.7<\/td>\n<\/tr>\n<tr>\n<td>Agent trace observability<\/td>\n<td>Primary action<\/td>\n<td style=\"text-align: right;\">68.5<\/td>\n<td style=\"text-align: right;\">69.8<\/td>\n<td style=\"text-align: right;\">71.6<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>In practical tasks as well, it records 77.0% in exact action selection for Customer service, showing performance that slightly surpasses the higher-tier model. Although it trails higher-tier models in complex tasks like Invoice processing, it maintains practical accuracy.<\/p>\n<p>Furthermore, according to measurements by the publisher of the base model Qwen3.5-9B, it possesses high vision-language processing capabilities, such as 78.4% on MMMU for image understanding and 85.7% on MathVista for solving math problems involving figures and graphs. Clef-Flash inherits these powerful vision and language capabilities while being optimized for structured decision-making outputs.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Clef-Flash excels not at standard chat generating free-form text, but at <strong>structured decision-making<\/strong> where passing a &#8220;state&#8221; and a &#8220;multiple-choice question schema&#8221; returns the probabilities for all choices of each question in a single inference. As indicated by tags such as <code>structured-output<\/code>, <code>classification<\/code>, and <code>image-text-to-typed-output<\/code>, its greatest feature is that output parsing is unnecessary, making it suitable for business processes like classification, routing, and intent determination. Since inputs accept images and video frames in addition to strings and JSON, it also supports use cases such as reading and evaluating invoice images.<\/p>\n<p>According to the model card, actual measurements on the Decision Index show high values such as 98.8 in BFCL (tool-calling accuracy) and 97.7 in the Home appliance simulator (case exact accuracy). In business workflow evaluations, it also outperforms the higher-tier Clef and Jev models with 77.0 in customer service action selection. In addition, its p95 latency is significantly lower than other models at 122.4ms, making it suitable for real-time routing and classification processing where speed is required. Since the base Qwen\/Qwen3.5-9B is a model equipped with multimodal capabilities, support for 201 languages, and long-context processing up to 1,010,000 tokens, it is believed to inherit those foundations even when handling multilingual documents or long state descriptions.<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<p>As related models built upon the same Qwen3.5-9B, our site has introduced <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/apple-lensvlm-9b\/\">&#8220;LensVLM-9B&#8221;, a vision-language model that compresses long text into images for processing<\/a> and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/22\/mimo-v26-distill-qwen-9b-gguf\/\">&#8220;MiMo-V2.6-Distill-Qwen-9B-GGUF&#8221;<\/a>. LensVLM-9B is a model tuned by Apple in the direction of compressing long documents into images to reduce token count, while MiMo-V2.6-Distill-Qwen-9B is distilled and SFT-trained by Xiaomi MiMo for coding and general-purpose agent tasks, with both assuming free-form text generation. In contrast, Clef-Flash significantly differs by eliminating free-form text generation itself and specializing in schema-aligned probability output, making it a derivative with a completely different directional use case despite sharing the same base model.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 9.4B parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>4GB (laptop iGPU \/ phone class)<\/td>\n<td>IQ2_M<\/td>\n<td>3.3GB<\/td>\n<td>4.0GB<\/td>\n<\/tr>\n<tr>\n<td>8GB (RTX 4060 \/ 3060 Ti, etc.)<\/td>\n<td>Q5_K_M<\/td>\n<td>6.4GB<\/td>\n<td>7.7GB<\/td>\n<\/tr>\n<tr>\n<td>12GB (RTX 4070 \/ 3060 12GB, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>8.9GB<\/td>\n<td>10.7GB<\/td>\n<\/tr>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>BF16<\/td>\n<td>16.7GB<\/td>\n<td>20.0GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-10-03): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. File sizes are measured from the converted build <a href=\"https:\/\/huggingface.co\/bartowski\/Cloudflare_clef-flash-GGUF\">bartowski\/Cloudflare_clef-flash-GGUF<\/a>. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Runs in Ollama, LM Studio and llama.cpp via a converted build.<\/strong><\/p>\n<p>The publisher ships safetensors, but <a href=\"https:\/\/huggingface.co\/bartowski\/Cloudflare_clef-flash-GGUF\">bartowski\/Cloudflare_clef-flash-GGUF<\/a> provides a GGUF build you can use.<\/p>\n<p><strong>License \u2014 <code>apache-2.0<\/code> (Commercial use allowed):<\/strong> Permits commercial use, modification and redistribution. Redistribution requires including the license and stating changes; includes a patent grant.<\/p>\n<p><strong>Compression:<\/strong> the IQ2_M build measures 3.01 bits per weight \u2014 about 19% the size of the original 16-bit weights, calculated by this site from the actual file sizes.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we measured ourselves on our server (no GPU) by actually reading and running this model&#8217;s files \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Decision-Model Evaluation on Our Own Tasks<\/h3>\n<p>We gave this model the same 11 situations with 20 questions that we give every decision model (yes\/no, named options and ordered levels; written by us, in Japanese and in English). llama.cpp cannot run this model&#8217;s decision head, so we ran the publisher&#8217;s own code (<code>joint_schema_model.py<\/code> and its SystemOne-compatible function) on a rented cloud GPU, inside a container with its network cut off. Each question states the rule to apply so that it has a single correct answer; our code decides right or wrong (yes when the probability of yes is 0.5 or more; for options and levels, the most probable one). These are not benchmark questions, and the result is not a general score of the model.<\/p>\n<p>Conditions: Modal, NVIDIA L4 (peak VRAM 17.9GB), BF16, transformers 5.18.0, torch 2.14.1+cu130, publisher revision <code>17f0b0ad64<\/code>. The time per decision is measured on that GPU from passing one situation (1 to 3 questions) to the publisher&#8217;s function until it returns; it is not a guide to the speed of your own GPU and cannot be compared with times measured on our CPU.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Method<\/th>\n<th>Quant<\/th>\n<th>Correct (Japanese)<\/th>\n<th>Correct (English)<\/th>\n<th>Mean probability on the correct answer (ja \/ en)<\/th>\n<th>Time per decision, median (ja \/ en)<\/th>\n<th>Peak memory<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>clef-flash<\/strong><\/td>\n<td><strong><code>joint_schema_model.py<\/code><\/strong><\/td>\n<td><strong><code>BF16<\/code><\/strong><\/td>\n<td><strong>18\/20<\/strong><\/td>\n<td><strong>18\/20<\/strong><\/td>\n<td><strong>0.88 \/ 0.88<\/strong><\/td>\n<td><strong>170ms \/ 167ms (GPU: NVIDIA L4)<\/strong><\/td>\n<td><strong>VRAM 17.9GB<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Kev-4B (reference)<\/td>\n<td><code>kev<\/code><\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>18\/20<\/td>\n<td>19\/20<\/td>\n<td>0.87 \/ 0.87<\/td>\n<td>7,616ms \/ 7,409ms<\/td>\n<td>6.3GB<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-clef-27b-multimodal-decision-model\/\">clef<\/a><\/td>\n<td><code>joint_schema_model.py<\/code><\/td>\n<td><code>BF16<\/code><\/td>\n<td>20\/20<\/td>\n<td>19\/20<\/td>\n<td>0.98 \/ 0.94<\/td>\n<td>190ms \/ 189ms (GPU: NVIDIA A100 80GB PCIe)<\/td>\n<td>VRAM 51.4GB<\/td>\n<\/tr>\n<tr>\n<td>lev (reference)<\/td>\n<td><code>lev<\/code><\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>16\/20<\/td>\n<td>16\/20<\/td>\n<td>0.76 \/ 0.80<\/td>\n<td>14,061ms \/ 13,186ms<\/td>\n<td>6.0GB<\/td>\n<\/tr>\n<tr>\n<td>Laya (reference)<\/td>\n<td><code>laya<\/code><\/td>\n<td><code>Q8_0<\/code><\/td>\n<td>10\/20<\/td>\n<td>15\/20<\/td>\n<td>0.47 \/ 0.68<\/td>\n<td>968ms \/ 678ms<\/td>\n<td>0.9GB<\/td>\n<\/tr>\n<tr>\n<td>Julia-1 (reference)<\/td>\n<td><code>laya<\/code><\/td>\n<td><code>Q8_0<\/code><\/td>\n<td>13\/20<\/td>\n<td>10\/20<\/td>\n<td>0.60 \/ 0.45<\/td>\n<td>136ms \/ 107ms<\/td>\n<td>0.5GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>It got the same number right in Japanese and English (18 of 20).<\/p>\n<p>Other rows are other decision models measured the same way (\u201creference\u201d rows are models we measure for comparison). Methods differ by model, and each model is measured with its own quantization.<\/p>\n<p>Times marked (GPU) were measured on a cloud GPU and the others on our CPU, so the times cannot be compared across those rows; the numbers of correct answers can.<\/p>\n<h4>Answer to Each Question<\/h4>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Situation<\/th>\n<th>Question<\/th>\n<th>Correct answer<\/th>\n<th>Japanese<\/th>\n<th>English<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Routing a support message<\/td>\n<td>Which team should handle this message?<\/td>\n<td><code>billing<\/code><\/td>\n<td>\u2713 <code>billing<\/code> (p=0.98)<\/td>\n<td>\u2713 <code>billing<\/code> (p=0.96)<\/td>\n<\/tr>\n<tr>\n<td>Routing a support message<\/td>\n<td>Is the customer asking for money back?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.95)<\/td>\n<td>\u2713 yes (p=0.80)<\/td>\n<\/tr>\n<tr>\n<td>Routing a support message<\/td>\n<td>Does the message report that a service is down?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=1.00)<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (within the window)<\/td>\n<td>Is today within 30 days of the delivery date?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.96)<\/td>\n<td>\u2713 yes (p=0.95)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (within the window)<\/td>\n<td>Can this item be returned under the policy?<\/td>\n<td>yes<\/td>\n<td>\u2717 no (p=0.14)<\/td>\n<td>\u2717 no (p=0.17)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (past the window)<\/td>\n<td>Is today within 30 days of the delivery date?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=0.98)<\/td>\n<td>\u2713 no (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Return eligibility (past the window)<\/td>\n<td>Can this item be returned under the policy?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=0.96)<\/td>\n<td>\u2713 no (p=0.95)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (vendor)<\/td>\n<td>What should happen to this invoice?<\/td>\n<td><code>reject<\/code><\/td>\n<td>\u2713 <code>reject<\/code> (p=0.97)<\/td>\n<td>\u2713 <code>reject<\/code> (p=0.98)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (vendor)<\/td>\n<td>Is the invoice total above 1,000 USD?<\/td>\n<td>no<\/td>\n<td>\u2713 no (p=1.00)<\/td>\n<td>\u2713 no (p=1.00)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (amount)<\/td>\n<td>What should happen to this invoice?<\/td>\n<td><code>manager<\/code><\/td>\n<td>\u2713 <code>manager<\/code> (p=0.94)<\/td>\n<td>\u2713 <code>manager<\/code> (p=0.96)<\/td>\n<\/tr>\n<tr>\n<td>Invoice handling (amount)<\/td>\n<td>Is the invoice total above 1,000 USD?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.98)<\/td>\n<td>\u2713 yes (p=0.98)<\/td>\n<\/tr>\n<tr>\n<td>Incident severity<\/td>\n<td>How widespread is the impact of this incident?<\/td>\n<td>3: All users affected<\/td>\n<td>\u2713 3: All users affected (p=0.92)<\/td>\n<td>\u2713 3: All users affected (p=0.95)<\/td>\n<\/tr>\n<tr>\n<td>Incident severity<\/td>\n<td>Is the service down?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.96)<\/td>\n<td>\u2713 yes (p=0.96)<\/td>\n<\/tr>\n<tr>\n<td>Delivery delay level<\/td>\n<td>Which level of the guideline does this delay fall into?<\/td>\n<td>2: Moderate delay<\/td>\n<td>\u2717 1: Minor delay (p=0.03)<\/td>\n<td>\u2717 1: Minor delay (p=0.04)<\/td>\n<\/tr>\n<tr>\n<td>Review opinion (negation)<\/td>\n<td>What is the reviewer&#8217;s overall opinion?<\/td>\n<td><code>positive<\/code><\/td>\n<td>\u2713 <code>positive<\/code> (p=1.00)<\/td>\n<td>\u2713 <code>positive<\/code> (p=0.99)<\/td>\n<\/tr>\n<tr>\n<td>Review opinion (negation)<\/td>\n<td>Does the reviewer say they would buy it again?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.97)<\/td>\n<td>\u2713 yes (p=0.98)<\/td>\n<\/tr>\n<tr>\n<td>Suspicious email (injected instruction)<\/td>\n<td>How should this email be classified?<\/td>\n<td><code>phishing<\/code><\/td>\n<td>\u2713 <code>phishing<\/code> (p=0.98)<\/td>\n<td>\u2713 <code>phishing<\/code> (p=0.98)<\/td>\n<\/tr>\n<tr>\n<td>Suspicious email (injected instruction)<\/td>\n<td>Does the email ask the reader to enter a password?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.97)<\/td>\n<td>\u2713 yes (p=0.96)<\/td>\n<\/tr>\n<tr>\n<td>Schedule overlap<\/td>\n<td>Does the meeting request overlap with an event in the calendar?<\/td>\n<td>yes<\/td>\n<td>\u2713 yes (p=0.93)<\/td>\n<td>\u2713 yes (p=0.93)<\/td>\n<\/tr>\n<tr>\n<td>Message intent (6 options)<\/td>\n<td>What does the user want to do?<\/td>\n<td><code>change_address<\/code><\/td>\n<td>\u2713 <code>change_address<\/code> (p=1.00)<\/td>\n<td>\u2713 <code>change_address<\/code> (p=1.00)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The full text of every situation and question, with the reason for each correct answer, is on <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h4>Setup and Steps We Ran<\/h4>\n<p>Download (a separate container without a GPU, which does not run the publisher&#8217;s code): Python 3.12, <code>pip install huggingface_hub<\/code>, then these files of <code>Cloudflare\/clef-flash<\/code> at revision <code>17f0b0ad64efb65d273590632833508766b2aae6<\/code> (no pickle weights):<\/p>\n<pre><code class=\"language-text\">chat_template.jinja\nconfig.json\ngeneration_config.json\njoint_head.safetensors\njoint_head_config.json\njoint_schema_model.py\nmodel-00001-of-00004.safetensors\nmodel-00002-of-00004.safetensors\nmodel-00003-of-00004.safetensors\nmodel-00004-of-00004.safetensors\nmodel.safetensors.index.json\nprocessor_config.json\ntokenizer.json\ntokenizer_config.json\n<\/code><\/pre>\n<p>Run (GPU container, network blocked, model files read-only): Python 3.12, <code>pip install torch transformers accelerate safetensors pillow torchvision<\/code>. The full code we ran in the container (<code>decision_remote.py<\/code>); the questions are the same request bodies as above:<\/p>\n<pre><code class=\"language-python\">&quot;&quot;&quot;\nModal \u306e\u30b3\u30f3\u30c6\u30ca\u306e\u4e2d\u3067\u52d5\u304f\u3001\u610f\u601d\u6c7a\u5b9a\u30e2\u30c7\u30eb\u306e\u914d\u5e03\u5143\u306e\u30b3\u30fc\u30c9\u306b\u3088\u308b\u8a55\u4fa1(2026-10-02\u3001\u30e6\u30fc\u30b6\u30fc\u306e\u5224\u65ad\u300cClef \u3092 Modal \u3067\u52d5\u304b\u3059\u300d)\u3002\nlmw\/lab\/gpu_modal.py \u306e decision_run \u304b\u3089\u547c\u3070\u308c\u308b\u3002**\u6a19\u6e96\u30e9\u30a4\u30d6\u30e9\u30ea\u3060\u3051\u3092\u5148\u982d\u3067\u8aad\u307f**\u3001torch\u30fbtransformers \u306f\u95a2\u6570\u306e\u4e2d\u3067\u8aad\u3080\u3002\n\nllama.cpp \u306b\u5224\u5b9a\u306e\u65b9\u5f0f\u304c\u7121\u3044\u610f\u601d\u6c7a\u5b9a\u30e2\u30c7\u30eb(Clef \u306f\u5224\u5b9a\u30d8\u30c3\u30c9 joint_head.safetensors \u3092\u914d\u5e03\u5143\u306e joint_schema_model.py \u3067\n\u8aad\u3080)\u3092\u3001\u914d\u5e03\u5143\u306e SystemOne \u4e92\u63db\u306e\u95a2\u6570(systemone)\u3067\u52d5\u304b\u3059\u3002\u3053\u306e\u30b3\u30f3\u30c6\u30ca\u306f**\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u906e\u65ad\u30fbModal \u306e\u6a5f\u80fd\u306a\u3057\u30fb\nVolume \u306f\u8aad\u307f\u53d6\u308a\u5c02\u7528**\u3067\u3001\u914d\u5e03\u5143\u306e\u30b3\u30fc\u30c9\u306f\u3053\u3053\u3067\u3060\u3051 import \u3059\u308b(\u30c0\u30a6\u30f3\u30ed\u30fc\u30c9\u306f remote_code_remote.download \u304c\nGPU \u306e\u7121\u3044\u5225\u306e\u30b3\u30f3\u30c6\u30ca\u3067\u884c\u3044\u3001\u305d\u3061\u3089\u3067\u306f import \u3057\u306a\u3044)\u3002\n\n\u554f\u984c(\u8981\u6c42\u306e\u672c\u6587)\u306f\u30db\u30b9\u30c8\u304c\u7d44\u3093\u3067\u6e21\u3059(lmw\/lab\/decision.py \u306e request_body\u3002llama.cpp \u3067\u6e2c\u308b\u30e2\u30c7\u30eb\u3068\u540c\u3058\u3082\u306e)\u3002\n\u6b63\u8aa4\u306f\u30db\u30b9\u30c8\u306e\u30b3\u30fc\u30c9(decision.score)\u304c\u5224\u5b9a\u3059\u308b\u3002\u3053\u3053\u3067\u306f\u5fdc\u7b54\u3092\u305d\u306e\u307e\u307e\u8fd4\u3057\u30011\u56de\u306e\u5224\u5b9a\u306b\u304b\u304b\u3063\u305f\u6642\u9593\u3092\u6e2c\u308b\u3060\u3051\u3002\n&quot;&quot;&quot;\nimport os\nimport platform\nimport sys\nimport time\nfrom datetime import datetime, timezone\n\nMOUNT = &quot;\/models&quot;\nPYTHON_VERSION = &quot;3.12&quot;\n# \u5b9f\u884c\u7528\u306e\u30b3\u30f3\u30c6\u30ca\u306b\u5165\u308c\u308b\u30d1\u30c3\u30b1\u30fc\u30b8(\u8a18\u4e8b\u306e\u624b\u9806\u306b\u3082\u540c\u3058\u5024\u3092\u8f09\u305b\u308b)\u3002Clef \u306e\u30ab\u30fc\u30c9\u306e\u8a18\u8f09\u306f torch 2.11\u30fb\n# transformers 5.10.2 \u3067\u3001\u753b\u50cf\u3092\u8aad\u3080\u51e6\u7406(AutoProcessor)\u306e\u305f\u3081\u306b pillow\u30fbtorchvision \u3092\u5165\u308c\u308b\nRUN_PACKAGES = (&quot;torch&quot;, &quot;transformers&quot;, &quot;accelerate&quot;, &quot;safetensors&quot;, &quot;pillow&quot;, &quot;torchvision&quot;)\n\ndef installed_packages() -&gt; list[str]:\n    from importlib import metadata\n    seen = {}\n    for dist in metadata.distributions():\n        name = dist.metadata[&quot;Name&quot;]\n        if name and name.lower() not in seen:\n            seen[name.lower()] = f&quot;{name}=={dist.version}&quot;\n    return sorted(seen.values(), key=str.lower)\n\ndef _keep(answer: dict) -&gt; dict:\n    return {k: answer[k] for k in (&quot;type&quot;, &quot;choice&quot;, &quot;noul&quot;, &quot;score&quot;, &quot;probabilities&quot;, &quot;confidence&quot;) if k in answer}\n\ndef _log(t0: float, message: str) -&gt; None:\n    print(f&quot;[decision {time.time() - t0:6.0f}s] {message}&quot;, flush=True)\n\ndef run(job: dict) -&gt; dict:\n    &quot;&quot;&quot;\n    job: {&quot;dir&quot;, &quot;module&quot;, &quot;loader&quot;, &quot;answer&quot;, &quot;warmup&quot;: \u8981\u6c42, &quot;requests&quot;: {lang: [[\u72b6\u6cc1\u306e id, \u8981\u6c42], ...]}}\u3002\n    \u623b\u308a\u5024\u306f\u8a18\u9332\u306e\u4e00\u90e8({&quot;results&quot;, &quot;gpu&quot;, &quot;vram_*&quot;, &quot;transformers&quot;, &quot;torch&quot;, &quot;packages&quot;, ...})\u3002\n    &quot;&quot;&quot;\n    import torch\n    import transformers\n    os.environ[&quot;HF_HUB_OFFLINE&quot;] = &quot;1&quot;\n    os.environ[&quot;TRANSFORMERS_OFFLINE&quot;] = &quot;1&quot;\n    t0 = time.time()\n    path = os.path.join(MOUNT, job[&quot;dir&quot;])\n    gpu = torch.cuda.get_device_name(0)\n    _log(t0, f&quot;GPU: {gpu}\u30fbtransformers {transformers.__version__}\u30fbtorch {torch.__version__}&quot;)\n    # \u914d\u5e03\u5143\u306e\u30b3\u30fc\u30c9\u3002\u3053\u306e\u30b3\u30f3\u30c6\u30ca(\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u906e\u65ad)\u3067\u3060\u3051\u8aad\u3080\n    sys.path.insert(0, path)\n    module = __import__(job[&quot;module&quot;])\n    model, processor = getattr(module, job[&quot;loader&quot;])(path, device=&quot;cuda&quot;)\n    answer = getattr(module, job[&quot;answer&quot;])\n    torch.cuda.synchronize()\n    load_sec = round(time.time() - t0)\n    vram_loaded = torch.cuda.memory_allocated()\n    _log(t0, f&quot;\u8aad\u307f\u8fbc\u307f {load_sec}\u79d2\u30fbVRAM {vram_loaded \/ 1024 ** 3:.1f}GB&quot;)\n    torch.cuda.reset_peak_memory_stats()\n    answer(model, processor, job[&quot;warmup&quot;])  # 1\u56de\u76ee\u306f\u6642\u9593\u3092\u6e2c\u3089\u306a\u3044\n    results: dict = {}\n    for lang, items in job[&quot;requests&quot;].items():\n        out = {}\n        for task_id, body in items:\n            torch.cuda.synchronize()\n            started = time.perf_counter()\n            resp = answer(model, processor, body)\n            torch.cuda.synchronize()\n            out[task_id] = {&quot;answers&quot;: {qid: _keep(a) for qid, a in (resp.get(&quot;answers&quot;) or {}).items()},\n                            &quot;ms&quot;: round((time.perf_counter() - started) * 1000, 1),\n                            &quot;input_tokens&quot;: (resp.get(&quot;usage&quot;) or {}).get(&quot;input_tokens&quot;)}\n        results[lang] = out\n    _log(t0, &quot;\u5168\u3066\u306e\u72b6\u6cc1\u3092\u89e3\u304d\u7d42\u3048\u307e\u3057\u305f&quot;)\n    return {&quot;results&quot;: results, &quot;gpu&quot;: gpu, &quot;vram_total&quot;: int(torch.cuda.get_device_properties(0).total_memory),\n            &quot;vram_loaded&quot;: int(vram_loaded), &quot;vram_peak&quot;: int(torch.cuda.max_memory_allocated()),\n            &quot;load_sec&quot;: load_sec, &quot;dtype&quot;: &quot;bfloat16&quot;, &quot;transformers&quot;: str(transformers.__version__),\n            &quot;torch&quot;: str(torch.__version__), &quot;python&quot;: platform.python_version(), &quot;packages&quot;: installed_packages(),\n            &quot;measured_at&quot;: datetime.now(timezone.utc).isoformat()}\n<\/code><\/pre>\n<p>Packages installed in the run container:<\/p>\n<pre><code class=\"language-text\">accelerate==1.15.0\naiohappyeyeballs==2.6.1\naiohttp==3.12.7\naiosignal==1.3.2\nannotated-doc==0.0.5\nanyio==4.15.1\nattrs==25.3.0\ncbor2==5.7.0\ncertifi==2026.7.22\nclick==8.5.0\ncuda-bindings==13.4.3\ncuda-pathfinder==1.8.3\ncuda-toolkit==13.0.3.0\nfilelock==4.0.9\nfrozenlist==1.6.0\nfsspec==2026.9.0\ngrpclib==0.4.8\nh11==0.16.0\nh2==4.2.0\nhf-xet==1.6.0\nhpack==4.1.0\nhttpcore==1.0.9\nhttpx==0.28.1\nhuggingface_hub==1.33.0\nhyperframe==6.1.0\nidna==3.20\nJinja2==3.1.6\nmarkdown-it-py==4.2.0\nMarkupSafe==3.0.3\nmdurl==0.1.2\nmpmath==1.3.0\nmultidict==6.4.4\nnetworkx==3.7\nnumpy==2.5.3\nnvidia-cublas==13.1.1.3\nnvidia-cuda-cupti==13.0.85\nnvidia-cuda-nvrtc==13.0.88\nnvidia-cuda-runtime==13.0.96\nnvidia-cudnn-cu13==9.24.0.43\nnvidia-cufft==12.0.0.61\nnvidia-cufile==1.15.1.6\nnvidia-curand==10.4.0.35\nnvidia-cusolver==12.0.4.66\nnvidia-cusparse==12.6.3.3\nnvidia-cusparselt-cu13==0.8.1\nnvidia-nccl-cu13==2.30.7\nnvidia-nvjitlink==13.4.92\nnvidia-nvshmem-cu13==3.4.5\nnvidia-nvtx==13.0.85\npackaging==26.3\npillow==12.3.0\npip==25.1.1\npropcache==0.3.1\nprotobuf==6.31.1\npsutil==7.2.2\nPygments==2.21.0\nPyYAML==6.0.3\nregex==2026.9.29\nrich==15.0.0\nsafetensors==0.8.0\nsetuptools==84.0.0\nshellingham==1.5.4\nsympy==1.14.0\ntokenizers==0.23.2\ntorch==2.14.1\ntorchvision==0.29.1\ntqdm==4.70.1\ntransformers==5.18.0\ntriton==3.8.0\ntyper==0.27.2\ntyping_extensions==4.16.0\nuv==0.7.19\nwheel==0.45.1\nyarl==1.20.0\n<\/code><\/pre>\n<p><!-- \/lmw:lab --><\/p>\n<h2>How to Get It<\/h2>\n<p>Clef-Flash is distributed in safetensors format as <code>Cloudflare\/clef-flash<\/code>. The license is apache-2.0, and it is stated that gating (agreeing to terms of use) is not required.<\/p>\n<p>The model card shows acquisition via <code>huggingface_hub<\/code> and loading using the bundled <code>joint_schema_model.py<\/code> (custom code) as follows:<\/p>\n<pre><code class=\"language-python\">from huggingface_hub import snapshot_download\n\npath = snapshot_download(&quot;Cloudflare\/clef-flash&quot;)\n<\/code><\/pre>\n<p>It is stated that verification was performed using <code>torch<\/code> 2.11, <code>transformers<\/code> 5.10.2, and a single H200 GPU. <code>pillow<\/code> is also required when using image and video inputs. Note that it uses a unique inference path (Joint schema head) loaded via the code included with the card, rather than the standard text generation pipeline.<\/p>\n<p><!-- lmw:variants --><\/p>\n<h2>Quantized and Converted Variants<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Added<\/th>\n<th>Publisher<\/th>\n<th>Format<\/th>\n<th>Repository<\/th>\n<th>Smallest VRAM tier (build, est. memory)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-10-01<\/td>\n<td>bartowski<\/td>\n<td>GGUF (imatrix)<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/Cloudflare_clef-flash-GGUF\">bartowski\/Cloudflare_clef-flash-GGUF<\/a><\/td>\n<td>IQ2_M 4.0GB (fits in 4GB VRAM)<\/td>\n<\/tr>\n<tr>\n<td>2026-10-01<\/td>\n<td>mlx-community<\/td>\n<td>MLX<\/td>\n<td><a href=\"https:\/\/huggingface.co\/mlx-community\/clef-flash-4bit\">mlx-community\/clef-flash-4bit<\/a><\/td>\n<td>MLX 4bit 6.6GB (fits in 8GB VRAM)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>File sizes of each build:<\/p>\n<ul>\n<li>Available builds in bartowski\/Cloudflare_clef-flash-GGUF: IQ2_M 3.3GB \/ Q2_K 3.4GB \/ IQ3_XXS 3.9GB \/ Q3_K_S 4.0GB \/ IQ3_XS 4.0GB \/ Q3_K_M 4.2GB \/ Q3_K_L 4.3GB \/ IQ3_M 4.5GB \/ IQ4_XS 4.9GB \/ Q4_0 5.1GB \/ Q4_K_S 5.1GB \/ IQ4_NL 5.4GB \/ Q4_K_M 5.4GB \/ Q4_1 5.5GB \/ Q4_K_L 5.8GB \/ Q5_K_S 6.0GB \/ Q5_K_M 6.4GB \/ Q6_K_S 7.0GB \/ Q6_K 7.3GB \/ Q6_K_L 7.5GB \/ Q8_0 8.9GB \/ BF16 16.7GB<\/li>\n<li>Available builds in mlx-community\/clef-flash-4bit: MLX 4bit 5.5GB<\/li>\n<\/ul>\n<p>In addition, 6 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model&#8217;s publisher or established quantization maintainers.<\/p>\n<p><em>This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:variants --><\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/apple-lensvlm-9b-2\/\">LensVLM-9B Vision-Language Model: 4GB+ VRAM, GGUF Builds<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/\">MiMo-V2.6-Distill-Qwen-9B-GGUF: Our Test Answers, 12GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-releases-clef-and-clef-flash\/\">clef Vision-Language Model: 80GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-clef-27b-multimodal-decision-model\/\">clef Structured Decision-Making Model: 80GB+ VRAM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 4GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-8gb-en\/\">Other models that run on a 8GB GPU<\/a><\/li>\n<li><strong>Engines that run this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<li><strong>What IQ2_M, Q5_K_M, Q8_0 mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-mlx-en\/\">MLX format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-alibaba-en\/\">Alibaba (Qwen): models, licenses and articles<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-vision\">Other vision-language models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/Cloudflare\/clef-flash\">https:\/\/huggingface.co\/Cloudflare\/clef-flash<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3.5-9B\">https:\/\/huggingface.co\/Qwen\/Qwen3.5-9B<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-10-02: Added converted builds to \u201cQuantized and Converted Variants\u201d: bartowski\/Cloudflare_clef-flash-GGUF, mlx-community\/clef-flash-4bit<\/li>\n<li>2026-10-02: Updated the hardware requirements table with the actual file sizes of bartowski\/Cloudflare_clef-flash-GGUF.<\/li>\n<li>2026-10-02: Changed the title to show what the article covers (VRAM requirements, file list, etc.).<\/li>\n<li>2026-10-03: Added our own measurements: decision-model evaluation on our fixed tasks.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Explore the specifications, benchmarks, strengths, and hardware requirements of Clef-Flash, a 9B structured decision-making model.<\/p>\n","protected":false},"author":1,"featured_media":8945,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[2865,2810,165,996,2867,1547,1952],"class_list":["post-8939","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-clef-flash-en","tag-cloudflare-en","tag-moe-en","tag-qwen3-5-en","tag-qwen3-5-9b-en","tag-verified","tag-vlm-en"],"lang":"en","translations":{"en":8939,"ja":8937},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8939","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=8939"}],"version-history":[{"count":7,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8939\/revisions"}],"predecessor-version":[{"id":9320,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8939\/revisions\/9320"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/8945"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=8939"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=8939"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=8939"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}