{"id":6230,"date":"2026-09-28T12:13:51","date_gmt":"2026-09-28T03:13:51","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/apple-lensvlm-9b-2\/"},"modified":"2026-09-28T22:38:39","modified_gmt":"2026-09-28T13:38:39","slug":"apple-lensvlm-9b-2","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/apple-lensvlm-9b-2\/","title":{"rendered":"LensVLM-9B Vision-Language Model: 4GB+ VRAM, GGUF Builds"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/apple\/LensVLM-9B\">apple\/LensVLM-9B<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-apple-en\/\">Apple: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-22<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apple-amlr<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Paper<\/td>\n<td><a href=\"https:\/\/arxiv.org\/abs\/2605.07019\">arXiv:2605.07019<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>A research team at Apple has released a vision-language model (VLM) called <strong>LensVLM-9B<\/strong> on Hugging Face, which <strong>shrinks long documents into images to have the model read them<\/strong>. Built on top of Qwen3.5-9B and fine-tuned by researchers from Apple and Duke University, the model stems from the paper &#8220;LensVLM: Selective Context Expansion for Compressed Visual Representation of Text&#8221; (arXiv:2605.07019).<\/p>\n<p>The goal is to reduce token counts and memory usage when handling long contexts. If text is tokenized as-is, a long document simply turns into a long sequence of tokens. On the other hand, when text is rendered as an image and passed to a VLM, a single image translates to a fixed number of visual tokens. Lowering the rendering resolution allows more text to be packed into fewer tokens. The problem was that shrinking it too much makes the text blur out and become unreadable.<\/p>\n<p>LensVLM solves this issue by &#8220;<strong>skimming the whole text while it is shrunk, and restoring only the pages likely to contain the answer to their original form to read them<\/strong>.&#8221; The model looks at a list of compressed page images, picks the relevant pages, calls tools learned during training to fetch the original text (or high-resolution images) for those pages, and reads them to answer. The name &#8220;Lens&#8221; comes from this behavior of magnifying a part of the compressed context to read it.<\/p>\n<p>The license is the Apple Machine Learning Research Model License, which <strong>restricts usage to non-commercial research purposes only<\/strong> (details below).<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameter count: Approx. 9B (fine-tuned from Qwen3.5-9B-Base)<\/li>\n<li>Input: Images with rendered text (page images) and questions. Can also be used with actual document images (such as PDFs)<\/li>\n<li>Compression settings: The included code allows choosing from three levels: <code>5x<\/code>, <code>10x<\/code>, and <code>15x<\/code>. According to the paper, Qwen3.5&#8217;s visual encoder converts a single image into 72, 48, or 24 visual tokens depending on the compression stage<\/li>\n<li>Tools: Uses <code>5x<\/code> to fetch original contents by specifying page numbers. Opens one page per call, and opens multiple pages in sequence when needed<\/li>\n<li>Reasoning: Reasons inside <code>&lt;think&gt;<\/code> tags before acting, then calls the tool<\/li>\n<li>Training procedure (from the paper): \u2460 Identify &#8220;difficult&#8221; questions that the original model cannot answer with compressed images alone, \u2461 Have Qwen3.5-397B synthesize demonstrations of &#8220;which pages to open and what to read&#8221; given the correct answer and rationale, followed by supervised fine-tuning (SFT), \u2462 Perform reinforcement learning (RL) through self-trial. Rewards are constructed from answer correctness and a tool-use bonus added only when correct. The visual encoder is frozen during training, updating only the language model portion<\/li>\n<li>Training data: Question answering from NQ, HotpotQA, MuSiQue, and HELMET<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>Here are the main results from the paper, presented using the publisher&#8217;s own figures.<\/p>\n<p><strong>Text Question Answering (Average of 7 Benchmarks)<\/strong>: While the accuracy of passing the full document as text (serving as an upper bound) is 72.4%, passing 5x compressed images as-is drops accuracy to 31.3%. LensVLM maintains <strong>68.9%<\/strong> under the same conditions, achieving an effective compression rate of 4.3x measured by the number of tokens actually processed by the model. The paper states that up to an effective compression rate of 10.1x, it outperformed all comparison methods across retrieval (RAG), text compression, and image compression.<\/p>\n<p><strong>Code Comprehension (Transfer without Fine-Tuning)<\/strong>: Despite no code-specific training, it outperformed comparison methods even in question answering from compressed code images.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Method<\/th>\n<th>5x<\/th>\n<th>10x<\/th>\n<th>15x<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Text (Passing full text)<\/td>\n<td>65.6<\/td>\n<td>65.6<\/td>\n<td>65.6<\/td>\n<\/tr>\n<tr>\n<td>Comp. Image (Reading compressed image as-is)<\/td>\n<td>26.7<\/td>\n<td>6.1<\/td>\n<td>3.2<\/td>\n<\/tr>\n<tr>\n<td>Glyph<\/td>\n<td>31.8<\/td>\n<td>9.7<\/td>\n<td>7.0<\/td>\n<\/tr>\n<tr>\n<td>LensVLM<\/td>\n<td>38.1<\/td>\n<td>26.1<\/td>\n<td>20.0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The gap widens with higher compression; at 15x, reading the compressed image as-is yields only 3.2% accuracy, whereas LensVLM maintains 20.0%. However, the gap from passing the full text (65.6%) is still substantial, and it cannot be claimed that accuracy is fully preserved under compression for code.<\/p>\n<p><strong>Actual Document Images (MMLongBench-Doc)<\/strong>: For document images originating as PDFs, &#8220;Zoom&#8221;\u2014which fetches contents as high-resolution images\u2014performed better than extracting text via OCR. At 15x compression, reading the compressed image as-is achieved 31.1% accuracy compared to 41.9% for the Zoom version. Conversely, at 5x compression where text remains legible, reading the compressed image as-is (51.2%) performed slightly better than the Zoom version (50.5%).<\/p>\n<p><strong>Differences by Model Size<\/strong>: Comparing models trained with the same method at 2B, 4B, and 9B scales shows large accuracy differences of 45.5%, 56.3%, and 68.9%, respectively. The paper notes that while the 2B model learns how to invoke tools, combining the ability to judge which pages to open with the understanding of opened contents only emerges at the 9B scale. For local users, it is important to note that <strong>this approach does not automatically work better on smaller models<\/strong>.<\/p>\n<p><strong>Check for Training Data Contamination<\/strong>: Re-evaluating exclusively on PubMed Central papers published after Qwen3.5&#8217;s training data cutoff (February 16, 2026) yielded an average of 39.7 points higher accuracy than reading compressed images as-is, following the same trend as the main results.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>LensVLM is useful for questions where <strong>only a small part of a long document is relevant to the answer<\/strong>. In the paper&#8217;s example, it opens only page 10 out of 15 pages of compressed images and reaches the answer in two interaction turns.<\/p>\n<p>What makes it effective for running locally is the reduction in KV cache (memory holding the context). According to the paper&#8217;s measurements, handling a 100-page input at 15x compression requires 51,273 tokens (1,602 MiB KV cache) when passing the full text, whereas LensVLM requires only 8,090 tokens (253 MiB) including the opened pages to answer, which is 84.2% less. Longer documents can thus be handled on hardware with less memory.<\/p>\n<p>On the other hand, <strong>speed is sacrificed<\/strong>. Because it alternates between generation and reading every time a page is opened, a typical case of opening a page once took 17 seconds\u2014about twice as long as the full-text passing method (8 seconds)\u2014as reported by the paper. The paper itself explicitly states that the goal of this approach is not speed, but &#8220;not losing accuracy even when compressed.&#8221;<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<ul>\n<li><strong>Differences from Glyph (Based on GLM-4.1V-9B)<\/strong>: Glyph is a model of similar scale that also has text read as images, but its research focuses on improving the ability to &#8220;read compressed images as-is&#8221; through rendering optimizations. LensVLM does not force unreadable parts to be read, fetching necessary pages instead. According to the paper, the original model is sensitive to rendering settings (fonts, image width, etc.), with accuracy fluctuating by as much as 32 points across 20 settings. When comparing three image widths matched for compression rate, the initial 18.0-point gap before training shrank to 0.5 points after training<\/li>\n<li><strong>Differences from Retrieval-Augmented Generation (RAG)<\/strong>: RAG builds an index in advance and retrieves fragments close to the question before reading. LensVLM presents the entire document compressed all at once and lets the model decide where to read. It eliminates the index-building step<\/li>\n<li><strong>Differences from the Original Qwen3.5-9B<\/strong>: Starting from the same weights, simply giving the same tools to the original model results in only 39.2% accuracy at 5x compression. The behavior of selecting and opening pages was acquired through this additional fine-tuning<\/li>\n<\/ul>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 9.4B parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>4GB (laptop iGPU \/ phone class)<\/td>\n<td>IQ2_M<\/td>\n<td>3.3GB<\/td>\n<td>4.0GB<\/td>\n<\/tr>\n<tr>\n<td>8GB (RTX 4060 \/ 3060 Ti, etc.)<\/td>\n<td>Q5_K_M<\/td>\n<td>6.4GB<\/td>\n<td>7.7GB<\/td>\n<\/tr>\n<tr>\n<td>12GB (RTX 4070 \/ 3060 12GB, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>8.9GB<\/td>\n<td>10.7GB<\/td>\n<\/tr>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>BF16<\/td>\n<td>16.7GB<\/td>\n<td>20.0GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-09-28): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. File sizes are measured from the converted build <a href=\"https:\/\/huggingface.co\/bartowski\/LensVLM-9B-GGUF\">bartowski\/LensVLM-9B-GGUF<\/a>. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Runs in Ollama, LM Studio and llama.cpp via a converted build.<\/strong><\/p>\n<p>The publisher ships safetensors, but <a href=\"https:\/\/huggingface.co\/bartowski\/LensVLM-9B-GGUF\">bartowski\/LensVLM-9B-GGUF<\/a> provides a GGUF build you can use.<\/p>\n<p><strong>License \u2014 <code>apple-amlr<\/code>:<\/strong> A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.<\/p>\n<p><strong>Compression:<\/strong> the IQ2_M build measures 3.01 bits per weight \u2014 about 19% the size of the original 16-bit weights, calculated by this site from the actual file sizes.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we measured ourselves on our server (no GPU) by actually reading and running this model&#8217;s files \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Japanese Token Efficiency<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Tokenizer<\/th>\n<th>Tokens per 1,000 Japanese characters<\/th>\n<th>Ratio to the same text in English<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>This model<\/strong><\/td>\n<td><strong>546<\/strong><\/td>\n<td><strong>0.99\u00d7<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen3<\/td>\n<td>688<\/td>\n<td>1.26\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Llama 3.2<\/td>\n<td>744<\/td>\n<td>1.36\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Gemma 3<\/td>\n<td>564<\/td>\n<td>1.03\u00d7<\/td>\n<\/tr>\n<tr>\n<td>gpt-oss<\/td>\n<td>795<\/td>\n<td>1.45\u00d7<\/td>\n<\/tr>\n<tr>\n<td>LLM-jp-3<\/td>\n<td>497<\/td>\n<td>0.85\u00d7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>It needs about 21% fewer tokens than the Qwen3 tokenizer for the same Japanese text, so about 1.26\u00d7 as much Japanese fits in the same context length, and generation is faster per character.<\/p>\n<p>Counted with the <code>tokenizer.json<\/code> of <a href=\"https:\/\/huggingface.co\/apple\/LensVLM-9B\">apple\/LensVLM-9B<\/a> on a fixed text we wrote ourselves (876 Japanese characters across news, conversation, technical docs, a formal email, travel writing and a recipe) and its English translation. Fewer tokens mean more Japanese fits in the context window.<\/p>\n<h3>Inside the GGUF File<\/h3>\n<p>File: <a href=\"https:\/\/huggingface.co\/bartowski\/LensVLM-9B-GGUF\/blob\/main\/LensVLM-9B-Q4_K_M.gguf\">LensVLM-9B-Q4_K_M.gguf<\/a> (5.44GB, <code>Q4_K_M<\/code>). We read only the header (metadata) of the file, not the weights.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Architecture (as named in the GGUF)<\/td>\n<td><code>qwen35<\/code><\/td>\n<\/tr>\n<tr>\n<td>Maximum trained context length<\/td>\n<td>262,144 tokens<\/td>\n<\/tr>\n<tr>\n<td>Layers<\/td>\n<td>32<\/td>\n<\/tr>\n<tr>\n<td>Vocabulary size<\/td>\n<td>248,320<\/td>\n<\/tr>\n<tr>\n<td>Chat template<\/td>\n<td>Included (mentions tool calls, has a thinking switch)<\/td>\n<\/tr>\n<tr>\n<td>imatrix<\/td>\n<td>Used (<code>LensVLM-9B-calibration-v6.txt<\/code>, 550 chunks)<\/td>\n<\/tr>\n<tr>\n<td>Weight types (share of parameters)<\/td>\n<td>Q4_K 72.4% \/ Q6_K 21.3% \/ Q8_0 6.2% \/ F32 0.1%<\/td>\n<\/tr>\n<tr>\n<td>Average bits per weight<\/td>\n<td>5.21 bits<\/td>\n<\/tr>\n<tr>\n<td>Embedding \/ output layer type<\/td>\n<td>Q4_K \/ Q6_K<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The quant name in the file name describes the file as a whole; in practice layers mix several types. The average bits per weight is the measured file data divided by the number of weights.<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<h2>How to Get It<\/h2>\n<ul>\n<li>Distribution format: Weights in Transformers format (safetensors) are available at <code>apple\/LensVLM-9B<\/code> on Hugging Face<\/li>\n<li>Example download command: <code>huggingface-cli download apple\/LensVLM-9B<\/code><\/li>\n<li>How to run: The model card guides users to clone <code>apple-aiml-research\/ml-lensvlm<\/code> from GitHub and run it using the provided scripts. Because this code handles rendering text into compressed images and running the page-fetching interaction loop, <strong>loading the weights into a general chat tool alone will not make the LensVLM mechanism work<\/strong>. The model card does not mention llama.cpp, Ollama, or LM Studio<\/li>\n<li>Example of provided script: <code>python demo.py --model apple\/LensVLM-9B --text_file document.txt --question \"...\" --compression 10x<\/code><\/li>\n<li>License: The weights are under the Apple Machine Learning Research Model License, which restricts usage strictly to &#8220;non-commercial scientific research and academic development,&#8221; explicitly excluding commercial product\/service use or product development. Derived models fine-tuned further are subject to the same restrictions. The included code falls under a separate Apple Sample Code License. Please review the original license terms before use.<\/li>\n<\/ul>\n<p><!-- lmw:variants --><\/p>\n<h2>Quantized and Converted Variants<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Added<\/th>\n<th>Publisher<\/th>\n<th>Format<\/th>\n<th>Repository<\/th>\n<th>Smallest VRAM tier (build, est. memory)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-28<\/td>\n<td>bartowski<\/td>\n<td>GGUF (imatrix)<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/LensVLM-9B-GGUF\">bartowski\/LensVLM-9B-GGUF<\/a><\/td>\n<td>IQ2_M 4.0GB (fits in 4GB VRAM)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-28<\/td>\n<td>mradermacher<\/td>\n<td>GGUF (imatrix)<\/td>\n<td><a href=\"https:\/\/huggingface.co\/mradermacher\/LensVLM-9B-i1-GGUF\">mradermacher\/LensVLM-9B-i1-GGUF<\/a><\/td>\n<td>IQ2_M 4.0GB (fits in 4GB VRAM)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-28<\/td>\n<td>mlx-community<\/td>\n<td>MLX<\/td>\n<td><a href=\"https:\/\/huggingface.co\/mlx-community\/LensVLM-9B-OptiQ-4bit\">mlx-community\/LensVLM-9B-OptiQ-4bit<\/a><\/td>\n<td>MLX 4bit 7.9GB (fits in 8GB VRAM)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>File sizes of each build:<\/p>\n<ul>\n<li>Available builds in bartowski\/LensVLM-9B-GGUF: IQ2_M 3.3GB \/ Q2_K 3.4GB \/ IQ3_XXS 3.9GB \/ Q3_K_S 4.0GB \/ IQ3_XS 4.0GB \/ Q3_K_M 4.2GB \/ Q3_K_L 4.3GB \/ IQ3_M 4.5GB \/ IQ4_XS 4.9GB \/ Q4_0 5.1GB \/ Q4_K_S 5.1GB \/ IQ4_NL 5.4GB \/ Q4_K_M 5.4GB \/ Q4_1 5.5GB \/ Q4_K_L 5.8GB \/ Q5_K_S 6.0GB \/ Q5_K_M 6.4GB \/ Q6_K_S 7.0GB \/ Q6_K 7.3GB \/ Q6_K_L 7.5GB \/ Q8_0 8.9GB \/ BF16 16.7GB<\/li>\n<li>Available builds in mradermacher\/LensVLM-9B-i1-GGUF: IQ1_S 2.6GB \/ IQ1_M 2.7GB \/ IQ2_XXS 2.9GB \/ IQ2_XS 3.1GB \/ IQ2_S 3.2GB \/ IQ2_M 3.4GB \/ Q2_K_S 3.4GB \/ Q2_K 3.6GB \/ IQ3_XXS 3.7GB \/ IQ3_XS 4.0GB \/ Q3_K_S 4.0GB \/ IQ3_S 4.1GB \/ IQ3_M 4.1GB \/ Q3_K_M 4.3GB \/ Q3_K_L 4.6GB \/ IQ4_XS 4.8GB \/ Q4_0 5.0GB \/ Q4_K_S 5.0GB \/ IQ4_NL 5.0GB \/ Q4_K_M 5.2GB \/ Q4_1 5.4GB \/ Q5_K_S 5.9GB \/ Q5_K_M 6.0GB \/ Q6_K 6.9GB<\/li>\n<li>Available builds in mlx-community\/LensVLM-9B-OptiQ-4bit: MLX 4bit 6.6GB<\/li>\n<\/ul>\n<p>In addition, 5 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model&#8217;s publisher or established quantization maintainers.<\/p>\n<p><em>This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:variants --><\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/\">MiMo-V2.6-Distill-Qwen-9B-GGUF Vision-Language Model: 12GB+ VRAM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 4GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-8gb-en\/\">Other models that run on a 8GB GPU<\/a><\/li>\n<li><strong>Engines that run this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<li><strong>What IQ2_M, Q5_K_M, Q8_0 mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-mlx-en\/\">MLX format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-apple-en\/\">Apple: models, licenses and articles<\/a><\/li>\n<li><strong>Other models for the same task<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-vision\">Other vision-language models<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/apple\/LensVLM-9B\">https:\/\/huggingface.co\/apple\/LensVLM-9B<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-09-28: Added converted builds to \u201cQuantized and Converted Variants\u201d: bartowski\/LensVLM-9B-GGUF, mradermacher\/LensVLM-9B-i1-GGUF, mlx-community\/LensVLM-9B-OptiQ-4bit<\/li>\n<li>2026-09-28: Updated the hardware requirements table with the actual file sizes of bartowski\/LensVLM-9B-GGUF.<\/li>\n<li>2026-09-28: Changed the title to show what the article covers (VRAM requirements, file list, etc.).<\/li>\n<li>2026-09-28: Added our own measurements: Japanese token efficiency, what is inside the GGUF file.<\/li>\n<li>2026-09-28: Added our own measurements: CPU speed and answers to Japanese questions.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Apple released LensVLM-9B, a vision-language model that compresses long documents into images to reduce token count and memory usage.<\/p>\n","protected":false},"author":1,"featured_media":6236,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[2567,2569,996,1547,1952],"class_list":["post-6230","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-apple-en","tag-lensvlm-9b-en","tag-qwen3-5-en","tag-verified","tag-vlm-en"],"lang":"en","translations":{"en":6230,"ja":6228},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6230","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=6230"}],"version-history":[{"count":7,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6230\/revisions"}],"predecessor-version":[{"id":6857,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6230\/revisions\/6857"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/6236"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=6230"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=6230"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=6230"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}