{"id":4283,"date":"2026-09-25T17:59:47","date_gmt":"2026-09-25T08:59:47","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/model-liquidai-lfm2-5-vl-en\/"},"modified":"2026-09-26T04:11:23","modified_gmt":"2026-09-25T19:11:23","slug":"model-liquidai-lfm2-5-vl-en","status":"publish","type":"page","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-liquidai-lfm2-5-vl-en\/","title":{"rendered":"LFM2.5-VL: Articles and Variants"},"content":{"rendered":"<h2>About This Model<\/h2>\n<p><strong>LFM2.5-VL-3B<\/strong> is a <strong>small model that can also read images<\/strong>, from Liquid AI. The company&#8217;s LFM2.5 is a family of &#8220;hybrid&#8221; models designed to <strong>run fast on-device<\/strong>. LFM2.5-VL-3B pairs the LFM2.5-2.6B language model with an image encoder (SigLIP2 NaFlex, 400M), supports 16 languages including Japanese, and has a 32,768-token context window.<\/p>\n<p>Compared with the previous generation (LFM2-VL-3B), it improves:<\/p>\n<ul>\n<li><strong>Grounding objects from a natural-language query<\/strong> (for example, answering &#8220;where is the red car?&#8221; with coordinates)<\/li>\n<li><strong>Full-page text recognition<\/strong> with layout information<\/li>\n<\/ul>\n<h2>What Makes It Stand Out<\/h2>\n<p><strong>1. It is small and fast.<\/strong> In the publisher&#8217;s measurements it runs at 228 tokens\/s on an Apple M5 Max and 116 tokens\/s on an AMD Ryzen AI Max+ 395, <strong>in under 3.3GB of memory.<\/strong> You can use an image-reading model at practical speed without a big GPU machine.<\/p>\n<p><strong>2. A companion model, DSpark, makes it faster without changing the output.<\/strong> <strong>LFM2.5-VL-3B-DSpark<\/strong>, covered in an article on this site, is a small 279.5M-parameter draft model. It predicts several tokens ahead and the main model verifies them in one pass (speculative decoding), speeding up generation <strong>while producing the same output the main model would on its own<\/strong> (identical under greedy decoding; under matched sampling settings, the output distribution is preserved). The publisher&#8217;s measured speedups (decoding speed):<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Task<\/th>\n<th style=\"text-align: right;\">H100 (SGLang)<\/th>\n<th style=\"text-align: right;\">Apple M5 Max (MLX-VLM)<\/th>\n<th style=\"text-align: right;\">Apple M3 Ultra (llama.cpp)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>COCO (image captioning)<\/td>\n<td style=\"text-align: right;\">2.66\u00d7<\/td>\n<td style=\"text-align: right;\">3.13\u00d7<\/td>\n<td style=\"text-align: right;\">2.14\u00d7<\/td>\n<\/tr>\n<tr>\n<td>MMMU-Pro (image-based questions)<\/td>\n<td style=\"text-align: right;\">2.43\u00d7<\/td>\n<td style=\"text-align: right;\">2.93\u00d7<\/td>\n<td style=\"text-align: right;\">2.03\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Multi-turn conversation<\/td>\n<td style=\"text-align: right;\">2.04\u00d7<\/td>\n<td style=\"text-align: right;\">2.30\u00d7<\/td>\n<td style=\"text-align: right;\">1.57\u00d7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Tasks with predictable output, like captioning, speed up the most; less predictable ones, like conversation, gain less. End to end (including image processing), the speedups are smaller than these figures.<\/p>\n<p><strong>Its limits are also clearly stated.<\/strong> The publisher recommends single-turn, speed-sensitive tasks such as real-time object detection, batch OCR of scanned documents, and translating menus or road signs, and advises against long-context or reasoning-heavy tasks, such as highly technical questions about blueprints.<\/p>\n<h2>Running It Locally<\/h2>\n<ul>\n<li><strong>The article on this page and the &#8220;Our Coverage and Data&#8221; section below cover DSpark, the draft model, not the main model.<\/strong> DSpark does not run on its own; it is used together with LFM2.5-VL-3B.<\/li>\n<li>The publisher distributes <strong>GGUF (for llama.cpp), ONNX and MLX (for Mac)<\/strong> builds.<\/li>\n<li>To use DSpark you need SGLang v0.5.19 or later, MLX-VLM v0.7.2 or later (currently <code>--temperature 0<\/code> only), or llama.cpp with the DSpark GGUF build.<\/li>\n<li><strong>The license is Liquid AI&#8217;s own LFM Open License v1.0.<\/strong> Commercial use by businesses above its US$10 million annual-revenue threshold is not licensed under it (commercial use below the threshold is allowed). Check the original terms before using it for work.<\/li>\n<\/ul>\n<p><em>Sources: model cards and license for <a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B\">LiquidAI\/LFM2.5-VL-3B<\/a> and <a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B-DSpark\">LiquidAI\/LFM2.5-VL-3B-DSpark<\/a>, as of 2026-09-25. Speed figures are the publisher&#8217;s measurements.<\/em><\/p>\n<h2>Our Coverage and Data<\/h2>\n<p>Everything Local Model Watch has published about the <strong>LFM2.5-VL<\/strong> family: 1 article(s) covering the base model and its fine-tunes, plus converted builds we tracked after publication. Memory requirements below are computed by this site from file sizes, not quoted from model cards. Part of our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-en\/\">model family index<\/a>.<\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Base model(s)<\/td>\n<td><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B\">LiquidAI\/LFM2.5-VL-3B<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher<\/td>\n<td>LiquidAI<\/td>\n<\/tr>\n<tr>\n<td>Articles<\/td>\n<td>1<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Quantized and Converted Variants<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Added<\/th>\n<th>Publisher<\/th>\n<th>Format<\/th>\n<th>Repository<\/th>\n<th>Smallest VRAM tier (build, est. memory)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-25<\/td>\n<td>LiquidAI<\/td>\n<td>GGUF<\/td>\n<td><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B-DSpark-GGUF\">LiquidAI\/LFM2.5-VL-3B-DSpark-GGUF<\/a><\/td>\n<td>F16 0.6GB (fits in 4GB VRAM)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>File sizes of each build:<\/p>\n<ul>\n<li>Available builds in LiquidAI\/LFM2.5-VL-3B-DSpark-GGUF: F16 0.5GB<\/li>\n<\/ul>\n<p><em>This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files.<\/em><\/p>\n<h2>Articles (the family&#8217;s own models first, then newest)<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Published<\/th>\n<th>Model<\/th>\n<th>Type<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-25<\/td>\n<td>LiquidAI\/LFM2.5-VL-3B-DSpark<\/td>\n<td>New Models<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/liquidai-lfm2-5-vl-dspark-released\/\">LFM2.5-VL-3B-DSpark Draft Model for Vision-Language Models: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Repositories<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B\">LiquidAI\/LFM2.5-VL-3B<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B-DSpark\">LiquidAI\/LFM2.5-VL-3B-DSpark<\/a><\/li>\n<\/ul>\n<p><em>Last updated 2026-09-26 (JST). The explanation at the top of this page was written with the help of AI from the primary sources it cites. The tables and lists under &#8220;Our Coverage and Data&#8221; are assembled by code from our article log.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>About This Model LFM2.5-VL-3B is a small model that can also read images, from Liquid AI. The company&#8217;s LFM2.5 is a family of &#8220;hybrid&#8221; models designed to run fast on-device. LFM2.5-VL-3B pairs the LFM2.5-2.6B language model with an image encoder (SigLIP2 NaFlex, 400M), supports 16 languages including Japanese, and has a 32,768-token context window. Compared [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-4283","page","type-page","status-publish","hentry"],"lang":"en","translations":{"en":4283,"ja":4282},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/4283","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=4283"}],"version-history":[{"count":2,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/4283\/revisions"}],"predecessor-version":[{"id":4474,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/4283\/revisions\/4474"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=4283"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}