{"id":458,"date":"2026-09-11T02:13:47","date_gmt":"2026-09-10T17:13:47","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/11\/nvidia-nim-optimization-nemotron-3-ultra\/"},"modified":"2026-09-18T21:41:58","modified_gmt":"2026-09-18T12:41:58","slug":"nvidia-nim-optimization-nemotron-3-ultra","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/nvidia-nim-optimization-nemotron-3-ultra\/","title":{"rendered":"NVIDIA Nemotron 3 Ultra NIM Achieves 2.5x Throughput on 4xB200"},"content":{"rendered":"<h2>Overview<\/h2>\n<p>In September 2026, NVIDIA announced that optimizations for Nemotron 3 Ultra via NIM (NVIDIA Inference Microservice) have improved system throughput on 4xB200 systems by up to 2.5x compared to the baseline. Using the NIM 2.0.12 optimized serving stack, it achieves a throughput of 1,997 tokens per second at a target of 50 TPS per user.<\/p>\n<h2>Announcement Details<\/h2>\n<p>The newly released NIM 2.0.12 optimized serving stack demonstrates significant performance improvements compared to the open-source baseline configuration (NIM Off). In tests with a native 256K maximum context, the optimized NIM stack recorded 1,997 tokens per second compared to 718 tokens per second for the baseline.<\/p>\n<p>This performance boost is achieved by combining multiple optimization layers rather than a single feature. Key optimization techniques include auto-tuned Mixture of Experts (MoE) and Mamba kernels for the Blackwell architecture, tensor parallel execution distributing the model across 4 GPUs, prefix and model state reuse (including prefix caching and partial prefix matching), scheduler, batching, and memory tuning, and MTP (Multi-Token Prediction) speculative decoding.<\/p>\n<p>Additionally, NIM packages model- and GPU-aware serving settings along with verified configurations, standard APIs, and container lifecycles. Furthermore, commercial support through NVIDIA AI Enterprise, regular inference stack updates, and CVE (Common Vulnerabilities and Exposures) responses are also provided.<\/p>\n<h2>Impact on Local LLM Users<\/h2>\n<p>For engineers operating open-weight models on their own NVIDIA GPU infrastructure, this announcement provides a concrete execution path to maximize inference throughput and effective performance.<\/p>\n<ul>\n<li>Developers can download the Nemotron 3 Ultra NIM (such as version 2.0.12) from NGC and deploy it as a container in a local NVIDIA GPU environment.<\/li>\n<li>Documentation and profile options (e.g., <code>vllm-nvidia-b200-nvfp4-tp4-pp1-throughput-90.0<\/code>) are provided, along with configuration steps such as enabling speculative decoding.<\/li>\n<li>For performance measurement in their own traffic environments, verification commands using NVIDIA AIPerf (using Mooncake format JSONL traces, etc.) are guided, enabling the selection of Pareto points tailored to latency SLOs.<\/li>\n<li>Regarding whether to focus on APIs or local execution, it is mentioned that in addition to local deployment via download from NGC, hosted APIs are also available for immediate testing.<\/li>\n<\/ul>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/developer.nvidia.com\/blog\/how-full-stack-nim-optimizations-deliver-2-5x-more-users-on-nemotron-3-ultra\/\">How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>NVIDIA announces that NIM 2.0.12 optimization for Nemotron 3 Ultra achieves up to 2.5x higher system throughput on 4xB200 systems.<\/p>\n","protected":false},"author":1,"featured_media":457,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1528],"tags":[917,968,971,165,919,921,117],"class_list":["post-458","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technical-reports","tag-b200-en","tag-blackwell-en","tag-mamba-en","tag-moe-en","tag-nemotron-3-ultra-en","tag-nvidia-nim-en","tag--en"],"lang":"en","translations":{"en":458,"ja":456},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/458","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=458"}],"version-history":[{"count":4,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/458\/revisions"}],"predecessor-version":[{"id":944,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/458\/revisions\/944"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/457"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=458"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=458"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=458"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}