{"id":8892,"date":"2026-10-02T07:10:06","date_gmt":"2026-10-01T22:10:06","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/nvidia-releases-pixelumm\/"},"modified":"2026-10-02T07:10:06","modified_gmt":"2026-10-01T22:10:06","slug":"nvidia-releases-pixelumm","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/nvidia-releases-pixelumm\/","title":{"rendered":"PixelUMM Multimodal Model: 24GB+ VRAM"},"content":{"rendered":"<p><em>Sample outputs are available on the <a href=\"https:\/\/huggingface.co\/nvidia\/PixelUMM\">model card<\/a>.<\/em><\/p>\n<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/nvidia\/PixelUMM\">nvidia\/PixelUMM<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-nvidia-en\/\">NVIDIA: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-02<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>nvidia-one-way-noncommercial-license<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>NVIDIA has released PixelUMM, a multimodal model that integrates the mutual understanding and generation of text, images, and video into a single model. This model supports a wide range of tasks, including text-to-image generation, image-to-text generation, video-text-to-text understanding, and video generation.<\/p>\n<p>The standout feature is its adoption of an &#8220;encoder-free&#8221; structure, where RGB images are treated directly as pixel patches rather than obtaining embeddings from separate pretrained vision encoders as in traditional models. This enables both image and video understanding and generation tasks to be processed directly in pixel space while sharing a common Transformer-based representation.<\/p>\n<h2>Specifications<\/h2>\n<p>The specifications and details of PixelUMM as described in the model card and public materials are as follows:<\/p>\n<ul>\n<li>Parameter count: 15,199,672,064<\/li>\n<li>Architecture: Decoder-only Transformer with raw-pixel patch embeddings and iterative pixel generation heads<\/li>\n<li>Language backbone: Qwen\/Qwen3-8B (revision <code>b968826d9c46dd6066d109eabc6255188de91218<\/code>)<\/li>\n<li>Image representation format: RGB images split into 16&#215;16 pixel patches<\/li>\n<li>Generation method: Iterative denoising for image and video generation<\/li>\n<li>Input format: Text, RGB images, video frames<\/li>\n<li>Output format: Text, RGB images, video frames<\/li>\n<li>Terms of use: Limited to non-commercial research or evaluation purposes only<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>The published model card does not contain specific numerical benchmark results or quantitative comparison tables with other models. However, NVIDIA mentions the quality, generation characteristics, and limitations of PixelUMM in the model card as follows.<\/p>\n<p>According to NVIDIA&#8217;s reports, the model does not always reliably follow input prompts, and semantic or temporal inconsistencies may occur in the generated images and videos. It is also explained that performance and output quality may vary depending on differences in corresponding languages and visual domains, video length, resolution, aspect ratio, and the executing hardware configuration.<\/p>\n<p>Additionally, because risks exist for outputting inaccurate information or inappropriate content, safety measures such as content filtering, usage monitoring, and access restrictions are required during use. Note that the model has not been validated for commercial production environments or operational environments requiring strict accuracy.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>PixelUMM is primarily intended for technical verification and research and development related to integrated multimodal modeling, pixel-space representation learning, multimodal understanding, and image and video generation. The model card provided by NVIDIA lists the following specific use cases:<\/p>\n<ul>\n<li>Research on shared representations spanning text, images, and video<\/li>\n<li>Performance evaluation of &#8220;encoder-free&#8221; multimodal architectures that do not use external vision encoders<\/li>\n<li>Generation of images and videos using text prompts as input<\/li>\n<li>Generation of text conditioned on input images or video frames<\/li>\n<\/ul>\n<p>Since this model can handle text, image patches, and video frames within the same framework on both input and output sides, its strength lies in its ability to directly process generation and understanding on the same network without preparing separate vision encoders like traditional task-specific models.<\/p>\n<p>Furthermore, regarding the performance and features of the language model &#8220;Qwen\/Qwen3-8B&#8221; on which PixelUMM is based, capabilities include switching between a &#8220;thinking mode&#8221; that enhances logical reasoning, mathematics, and code generation and a &#8220;non-thinking mode&#8221; for general conversation, agent capabilities for tool calling, and multilingual understanding and translation supporting over 100 languages. PixelUMM builds upon this Qwen3-8B language backbone while integrating raw-pixel patch embeddings and a pixel generation head via iterative denoising to provide an environment that processes text and visual representations in a unified manner.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 8.2B parameters (taken from the base model Qwen\/Qwen3-8B)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>BF16<\/td>\n<td>15.3GB<\/td>\n<td>18.3GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>License \u2014 <code>nvidia-one-way-noncommercial-license<\/code>:<\/strong> A custom license from the publisher. Check the original terms directly, including whether commercial use is permitted.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats and the license field. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<h3>Measurements We Did Not Take<\/h3>\n<p>We have not confirmed that this site meets the commercial-use terms of this model&#8217;s license (<code>nvidia-one-way-noncommercial-license<\/code>). Because this site carries advertising, we did not run the model (no CPU run, answers, quantization comparison, conversion or generation).<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<h2>How to Get It<\/h2>\n<p>The model repository for PixelUMM is available on Hugging Face at <code>nvidia\/PixelUMM<\/code>.<\/p>\n<p>To run and verify this model, a Linux environment, an NVIDIA GPU, CUDA-compatible PyTorch, and FlashAttention are specified as mandatory requirements. The actual installation procedure involves setting up PyTorch and FlashAttention according to the CUDA environment being used, followed by installing the remaining dependency libraries.<\/p>\n<p>There is no gated license agreement required on Hugging Face for access, and it can be obtained directly from the repository. However, please note that the applicable license is the NVIDIA One-Way Non-Commercial License (<code>nvidia-one-way-noncommercial-license<\/code>), and distribution is restricted solely to non-commercial research or evaluation purposes.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/nvidia-releases-deepseek-v4-pro-nvfp4\/\">DeepSeek-V4-Pro-0813-nvfp4-DSpark: ~1005GB Memory<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 24GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-24gb-en\/\">Other models that run on a 24GB GPU<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-nvidia-en\/\">NVIDIA: models, licenses and articles<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/nvidia\/PixelUMM\">https:\/\/huggingface.co\/nvidia\/PixelUMM<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3-8B\">https:\/\/huggingface.co\/Qwen\/Qwen3-8B<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Explore NVIDIA&#8217;s PixelUMM, an encoder-free multimodal model handling text, image, and video generation and understanding in a single framework.<\/p>\n","protected":false},"author":1,"featured_media":8891,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[695,2848,375,2850,1547,896],"class_list":["post-8892","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-nvidia-en","tag-pixelumm-en","tag-qwen3-en","tag-transformer-en","tag-verified","tag--en"],"lang":"en","translations":{"en":8892,"ja":8890},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8892","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=8892"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8892\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/8891"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=8892"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=8892"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=8892"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}