{"id":9695,"date":"2026-10-04T05:13:07","date_gmt":"2026-10-03T20:13:07","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/04\/liquidai-lfm2-5-350m-diffusion-exp\/"},"modified":"2026-10-05T01:38:16","modified_gmt":"2026-10-04T16:38:16","slug":"liquidai-lfm2-5-350m-diffusion-exp","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/04\/liquidai-lfm2-5-350m-diffusion-exp\/","title":{"rendered":"LFM2.5-350M-Diffusion-Exp Text Generation Model: 4GB+ VRAM"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-350M-Diffusion-Exp\">LiquidAI\/LFM2.5-350M-Diffusion-Exp<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-liquidai-en\/\">Liquid AI: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-03<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>lfm1.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<p><!-- lmw:lab-summary --><\/p>\n<p><strong>What we checked ourselves<\/strong><\/p>\n<ul>\n<li>The same Japanese text takes about as many tokens as with the Qwen3 tokenizer.<\/li>\n<\/ul>\n<p>Details and conditions are in \u201cOur Own Measurements\u201d below.<\/p>\n<p><!-- \/lmw:lab-summary --><\/p>\n<h2>Overview<\/h2>\n<p>LiquidAI has released an experimental model called &#8220;LFM2.5-350M-Diffusion-Exp&#8221;, which converts the company&#8217;s lightweight LLM &#8220;LFM2.5-350M&#8221; into a uniform-state block-diffusion model. Instead of outputting one token at a time in a single forward pass like traditional autoregressive models, this model adopts a method of parallel denoising for 32-token blocks. As a result, it achieves faster decoding speeds than the parent model under small batch size environments, in exchange for the number of denoising steps.<\/p>\n<p>It is reported that at a denoising setting of 8 steps per block (NFE 8), the model maintains performance equivalent to, or very close to, the original autoregressive model &#8220;LFM2.5-350M&#8221; across most benchmarks. The backbone directly reuses the same &#8220;short-convolution + attention&#8221; hybrid architecture, tokenizer, and chat template as LFM2.5.<\/p>\n<p>The license for this model is &#8220;lfm1.0&#8221;, and detailed information can be found on the <a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-350M-Diffusion-Exp\">Hugging Face model page<\/a>.<\/p>\n<h2>Specifications<\/h2>\n<p>The specifications of this model are as follows:<\/p>\n<ul>\n<li><strong>Parameters<\/strong>: 350M<\/li>\n<li><strong>Layers<\/strong>: 16 (10 short-convolution + 6 GQA, adaLN noise conditioning, self-conditioning)<\/li>\n<li><strong>Block Size<\/strong>: 32 tokens (bidirectional within blocks, causal between blocks)<\/li>\n<li><strong>Training<\/strong>: 800B pre-training + 600B mid-training + 150B SFT tokens, followed by post-training and step distillation<\/li>\n<li><strong>Vocabulary<\/strong>: 65,536 (64,400 sampled)<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Below is the performance comparison table between this model and other lightweight models, as published in the official model card. Note that &#8220;NFE&#8221; represents the number of denoising steps per 32-token block.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>NFE 32<\/th>\n<th>NFE 8<\/th>\n<th>NFE 4<\/th>\n<th>LFM2.5-350M<\/th>\n<th>Granite 4.0 H-350M<\/th>\n<th>Qwen3.5-0.8B<\/th>\n<th>Granite 4.0 H-1B<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>GPQA Diamond<\/td>\n<td>32.8<\/td>\n<td>32.5<\/td>\n<td>32.8<\/td>\n<td>29.8<\/td>\n<td>23.7<\/td>\n<td>23.8<\/td>\n<td>22.9<\/td>\n<\/tr>\n<tr>\n<td>MMLU-Pro<\/td>\n<td>18.6<\/td>\n<td>19.6<\/td>\n<td>17.6<\/td>\n<td>21.8<\/td>\n<td>10.8<\/td>\n<td>40.6<\/td>\n<td>23.6<\/td>\n<\/tr>\n<tr>\n<td>CaseReportBench<\/td>\n<td>45.5<\/td>\n<td>37.2<\/td>\n<td>13.3<\/td>\n<td>37.6<\/td>\n<td>30.2<\/td>\n<td>36.5<\/td>\n<td>35.4<\/td>\n<\/tr>\n<tr>\n<td>IFEval<\/td>\n<td>81.1<\/td>\n<td>75.7<\/td>\n<td>69.0<\/td>\n<td>77.2<\/td>\n<td>59.9<\/td>\n<td>64.0<\/td>\n<td>77.9<\/td>\n<\/tr>\n<tr>\n<td>Multi-IF<\/td>\n<td>50.1<\/td>\n<td>46.7<\/td>\n<td>40.3<\/td>\n<td>45.3<\/td>\n<td>26.9<\/td>\n<td>41.0<\/td>\n<td>47.5<\/td>\n<\/tr>\n<tr>\n<td>IFBench<\/td>\n<td>34.1<\/td>\n<td>33.6<\/td>\n<td>32.3<\/td>\n<td>38.3<\/td>\n<td>19.3<\/td>\n<td>23.6<\/td>\n<td>25.5<\/td>\n<\/tr>\n<tr>\n<td>BFCL v4<\/td>\n<td>18.7<\/td>\n<td>17.4<\/td>\n<td>14.4<\/td>\n<td>22.0<\/td>\n<td>13.2<\/td>\n<td>20.0<\/td>\n<td>27.9<\/td>\n<\/tr>\n<tr>\n<td>BFCL v3<\/td>\n<td>38.1<\/td>\n<td>35.2<\/td>\n<td>29.8<\/td>\n<td>44.4<\/td>\n<td>27.5<\/td>\n<td>38.2<\/td>\n<td>49.5<\/td>\n<\/tr>\n<tr>\n<td>\u03c4\u00b2-Bench Telecom<\/td>\n<td>12.3<\/td>\n<td>11.4<\/td>\n<td>10.5<\/td>\n<td>13.2<\/td>\n<td>4.4<\/td>\n<td>61.4<\/td>\n<td>14.9<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>According to the official measurement results, this model records high scores of 32.8 at NFE 32 and 32.5 at NFE 8 on &#8220;GPQA Diamond&#8221;, a challenging science benchmark handling PhD-level physics, chemistry, and biology questions. This is an excellent result that outperforms not only the autoregressive base model LFM2.5-350M (29.8) but also larger parameter models such as Qwen3.5-0.8B (23.8) and Granite 4.0 H-1B (22.9). Additionally, it records 81.1 at NFE 32 on &#8220;IFEval&#8221;, which measures the ability to follow formal instructions, showing the highest figure among the comparison targets.<\/p>\n<p>On the other hand, in &#8220;MMLU-Pro&#8221;, which measures general knowledge and reasoning abilities, the model&#8217;s scores remain at 18.6 for NFE 32 and 19.6 for NFE 8, falling behind Qwen3.5-0.8B (40.6) and Granite 4.0 H-1B (23.6). Furthermore, in agent-oriented metrics &#8220;BFCL v3&#8221; and &#8220;BFCL v4&#8221;, which measure the accuracy of function selection and argument assembly, the results are lower than the base model, Qwen3.5-0.8B, and Granite 4.0 H-1B.<\/p>\n<p>Regarding the impact of reducing the number of denoising steps (NFE), while reducing from NFE 32 to NFE 8 keeps performance drops slight across many benchmarks, reducing further to the fastest setting of NFE 4 results in significant performance degradation in specific tasks, such as the &#8220;CaseReportBench&#8221; score plunging from 45.5 to 13.3. Depending on the use case, it is necessary to select settings considering the balance between speed and accuracy.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>This model is an edge- and dialogue-oriented model that achieves fast text generation through a block-diffusion approach while boasting an extremely compact size of 350M parameters.<\/p>\n<p>Based on the design philosophy of the base model &#8220;LFM2.5-350M&#8221;, it is primarily suited for lightweight automation tasks such as data extraction, structured output, and tool use (function calling), as well as on-device dialogue processing. Meanwhile, the official model card explicitly states that it is not recommended for knowledge-intensive tasks or programming purposes.<\/p>\n<p>The biggest feature of this model is that the trade-off between speed and output quality can be flexibly adjusted by switching the number of denoising steps (NFE) during inference. The decoding configurations provided by the official source are the following three types:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Decode config<\/th>\n<th>Steps per block<\/th>\n<th>Use<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>decode_configs\/nfe32.yaml<\/code><\/td>\n<td>32<\/td>\n<td>Highest quality<\/td>\n<\/tr>\n<tr>\n<td><code>decode_configs\/nfe8.yaml<\/code><\/td>\n<td>8<\/td>\n<td>Recommended<\/td>\n<\/tr>\n<tr>\n<td><code>decode_configs\/nfe4.yaml<\/code><\/td>\n<td>4<\/td>\n<td>Fastest<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>In practical use, &#8220;NFE 8&#8221;, which processes with 8 steps per block, is set as the recommended configuration, allowing users to leverage fast decoding while maintaining quality equivalent to the autoregressive base model. Users can switch between configurations, such as selecting 32-step &#8220;NFE 32&#8221; when higher output accuracy is required, or 4-step &#8220;NFE 4&#8221; when prioritizing extreme processing speed over quality.<\/p>\n<p>Note that ancestral sampling using temperature annealing, which gradually lowers the temperature from 0.8 to 0.4, is specified for sampling during generation, and decoding via greedy decoding is discouraged.<\/p>\n<p>Regarding supported languages, the base model LFM2.5-350M covers multilingual capabilities including English, Arabic, Chinese, French, German, Japanese, Korean, Portuguese, and Spanish.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 425M parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>4GB (laptop iGPU \/ phone class)<\/td>\n<td>BF16<\/td>\n<td>0.8GB<\/td>\n<td>0.9GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-10-05): llama.cpp: not registered, vLLM: not registered, MLX (mlx-lm): not registered. &#8220;Not registered&#8221; means the name is absent from that registry today, not that the model cannot run.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Not usable in Ollama, LM Studio and llama.cpp yet.<\/strong><\/p>\n<p>The publisher ships safetensors only, and llama.cpp&#8217;s registry does not list this architecture. llama.cpp would need to add support before these tools can run it. Today it can be run with transformers, using the memory figures in the table above.<\/p>\n<p><strong>License \u2014 <code>lfm1.0<\/code> (Commercial use allowed with conditions):<\/strong> LiquidAI&#8217;s LFM Open License v1.0. Commercial use is allowed, except by legal entities with annual revenue of $10 million or more (Section 5). Redistribution requires the license text and notice of changes.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we measured ourselves on our server (no GPU) by actually reading and running this model&#8217;s files \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Japanese Token Efficiency<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Tokenizer<\/th>\n<th>Tokens per 1,000 Japanese characters<\/th>\n<th>Ratio to the same text in English<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>This model<\/strong><\/td>\n<td><strong>686<\/strong><\/td>\n<td><strong>1.22\u00d7<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen3<\/td>\n<td>688<\/td>\n<td>1.26\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Llama 3.2<\/td>\n<td>744<\/td>\n<td>1.36\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Gemma 3<\/td>\n<td>564<\/td>\n<td>1.03\u00d7<\/td>\n<\/tr>\n<tr>\n<td>gpt-oss<\/td>\n<td>795<\/td>\n<td>1.45\u00d7<\/td>\n<\/tr>\n<tr>\n<td>LLM-jp-3<\/td>\n<td>497<\/td>\n<td>0.85\u00d7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The same Japanese text takes about as many tokens as with the Qwen3 tokenizer.<\/p>\n<p>Counted with the <code>tokenizer.json<\/code> of <a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-350M-Base\">LiquidAI\/LFM2.5-350M-Base<\/a> on a fixed text we wrote ourselves (876 Japanese characters across news, conversation, technical docs, a formal email, travel writing and a recipe) and its English translation. Fewer tokens mean more Japanese fits in the context window.<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<p><!-- lmw:peers --><\/p>\n<h2>Recent Models in the Same Size Class<\/h2>\n<p><em>Models with <\/em><em>up to 4B<\/em><em> parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site&#8217;s estimates; licenses are as stated on the model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>togethercomputer\/Tev1-0.8B-experimental<\/td>\n<td>873M<\/td>\n<td>4GB<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/tev1-08b-experimental-2\/\">Tev1-0.8B-experimental Text Generation Model: 4GB+ VRAM<\/a> (2026-09-25)<\/td>\n<\/tr>\n<tr>\n<td>pfnet\/plamo-3-610m-fin-instruct<\/td>\n<td>890M<\/td>\n<td>4GB<\/td>\n<td>plamo-community-license<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/plamo-3-610m-fin-instruct-2\/\">plamo-3-610m-fin-instruct Text Generation Model: 4GB+ VRAM<\/a> (2026-09-24)<\/td>\n<\/tr>\n<tr>\n<td>harshatheg\/Qwen-2.5-1B-RLCD<\/td>\n<td>1.5B<\/td>\n<td>4GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/17\/qwen-25-1b-rlcd-mlx-constrained-decoding\/\">Qwen-2.5-1B-RLCD Text Generation Model: 4GB+ VRAM<\/a> (2026-09-16)<\/td>\n<\/tr>\n<tr>\n<td>tencent\/Simple-Attention-Sparsification<\/td>\n<td>4.0B<\/td>\n<td>12GB<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/tencent-simple-attention-sparsification-qwen3\/\">Simple-Attention-Sparsification Text Generation Model: 12GB+ VRAM<\/a> (2026-09-14)<\/td>\n<\/tr>\n<tr>\n<td>openbmb\/MiniCPM5-2B<\/td>\n<td>2.5B<\/td>\n<td>4GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">MiniCPM5-2B Text Generation Model: Our Test Answers, 4GB+ VRAM<\/a> (2026-09-07)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:peers --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is published on Hugging Face in &#8220;safetensors&#8221; format and is provided in an open state that does not require prior consent to usage terms.<\/p>\n<p>Dedicated code for inference, serving scripts using SGLang, and evaluation tools are published in the official repository &#8220;<a href=\"https:\/\/github.com\/Liquid4All\/lfm-diffusion\">Liquid4All\/lfm-diffusion<\/a>&#8220;.<\/p>\n<p>From a Python environment, you can load the model and execute inference using the dedicated library <code>lfm_diffusion<\/code> as follows:<\/p>\n<pre><code class=\"language-python\">from lfm_diffusion import generate, load_model\n\nmodel, tokenizer = load_model(&quot;LiquidAI\/lfm2.5-350m-diffusion-exp&quot;)\nmessages = [{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;Give three tips for getting better sleep.&quot;}]\nprint(generate(model, tokenizer, messages, config=&quot;nfe8&quot;, max_new_tokens=256))\n<\/code><\/pre>\n<p><!-- lmw:same-task --><\/p>\n<h2>Other Models for the Same Task<\/h2>\n<p><em>Recent text generation models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site&#8217;s estimate.<\/em><\/p>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/elyza-thinking-1-0-32b-33b-2\/\">ELYZA Releases ELYZA-Thinking-1.0 32B\/33B Reasoning Models<\/a> (32.1B, 80GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/orcasaq2-27b-qwen3-8-27b-quantized\/\">OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM<\/a> (27.8B, 16GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/hemmingway-1-open-27b-model-specialized-for-human-like-writing\/\">Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds<\/a> (26.9B, 12GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/xing4-0-29b-a4b-china-telecom\/\">Xing4.0-29B-A4B Text Generation Model: Our Test Answers, 24GB+ VRAM<\/a> (31.2B, 24GB)<\/li>\n<\/ul>\n<p><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-text\">See all text generation models \u2192<\/a><\/p>\n<p><!-- \/lmw:same-task --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 4GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-8gb-en\/\">Other models that run on a 8GB GPU<\/a><\/li>\n<li><strong>Formats this model is available in<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">Safetensors format guide and models<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-liquidai-en\/\">Liquid AI: models, licenses and articles<\/a><\/li>\n<li><strong>How to read GPQA Diamond, MMLU-Pro, IFEval<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-benchmarks-en\/\">Benchmark glossary<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-350M-Diffusion-Exp\">LiquidAI\/LFM2.5-350M-Diffusion-Exp (Hugging Face)<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-350M\">LiquidAI\/LFM2.5-350M (Hugging Face)<\/a><\/li>\n<li><a href=\"https:\/\/github.com\/Liquid4All\/lfm-diffusion\">Liquid4All\/lfm-diffusion (GitHub)<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-10-05: Added our own measurements: Japanese token efficiency.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>LiquidAI has released LFM2.5-350M-Diffusion-Exp, an experimental block-diffusion model converted from its lightweight LFM2.5-350M LLM.<\/p>\n","protected":false},"author":1,"featured_media":9694,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[2959,2961,2278,1547,1555,1746],"class_list":["post-9695","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-lfm2-5-en","tag-lfm2-5-350m-diffusion-exp-en","tag-liquidai-en","tag-verified","tag--en"],"lang":"en","translations":{"en":9695,"ja":9693},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9695","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9695"}],"version-history":[{"count":3,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9695\/revisions"}],"predecessor-version":[{"id":9907,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9695\/revisions\/9907"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9694"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9695"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9695"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9695"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}