{"id":462,"date":"2026-09-11T02:14:35","date_gmt":"2026-09-10T17:14:35","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/11\/nvidia-bionemo-inference-runtime-announced\/"},"modified":"2026-09-18T21:41:58","modified_gmt":"2026-09-18T12:41:58","slug":"nvidia-bionemo-inference-runtime-announced","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/nvidia-bionemo-inference-runtime-announced\/","title":{"rendered":"NVIDIA Announces BioNeMo Inference Runtime for Biomolecules"},"content":{"rendered":"<h2>Overview<\/h2>\n<p>On September 10, 2026, NVIDIA announced the &#8220;NVIDIA BioNeMo Inference Runtime (BioIR)&#8221;, which accelerates inference for biomolecular structure prediction models. While maintaining existing PyTorch workflows, this runtime speeds up processing for supported models on NVIDIA GPUs, achieving higher throughput and reduced power consumption in models such as Boltz-2.<\/p>\n<h2>Announcement Details<\/h2>\n<p>BioIR provides processor and PyTorch module integration that streamlines the entire biomolecular structure prediction pipeline. Key features and optimization mechanisms include:<\/p>\n<ul>\n<li>End-to-end execution capabilities for input parsing, tokenization, feature generation, GPU inference, and output in PDB or mmCIF formats as a pipeline process.<\/li>\n<li>Optimization across three layers: kernel selection, module optimization via CUDA Graph capture, and pipeline scaling using Ray replicas.<\/li>\n<li>Use of a Ray executor to place a full model replica on each GPU on a single node, overlapping CPU stages with GPU folding processing to improve throughput.<\/li>\n<li>In a benchmark of 1,000 human-derived dimer targets using 8 H100 GPUs, BioIR-accelerated Boltz-2 processed 58.5K folded residues per allocated GPU hour, achieving 2.90x the throughput compared to 20.2K for the torch-compiled open-source implementation.<\/li>\n<li>The estimated energy required to fold 1,000 million equivalent targets, calculated from the TDP of 8 H100 GPU nodes, is estimated at 11 MWh for BioIR and 35 MWh for the public implementation.<\/li>\n<\/ul>\n<h2>Background<\/h2>\n<p>Biomolecular structure prediction is now frequently executed at proteome scale, requiring efficient processing of large worklists. BioIR has already been utilized to accelerate the generation of approximately 31 million protein complex structures across 4,777 proteomes in recent expansions of the AlphaFold Database (AFDB), of which 1.81 million have been published as high-confidence predictions. It was developed to address processing efficiency challenges in such large-scale structure prediction workloads.<\/p>\n<h2>Impact on Local LLM Users<\/h2>\n<p>For engineers running open-weight structure prediction models (such as Boltz-2, OpenFold2, and OpenFold3) on local or on-premise NVIDIA GPU environments at scale, BioIR can be a tool that significantly improves inference efficiency. Key impacts and considerations include:<\/p>\n<ul>\n<li>Published as an official repository, developers can introduce it into their own environments and incorporate it into large-scale structure prediction workflows.<\/li>\n<li>Provided as a wheel containing pre-compiled CUBINS, meaning it does not necessarily require a full source build environment of nvcc or the CUDA toolkit at runtime.<\/li>\n<li>Execution requires Python 3.12 or later, compatible NVIDIA GPUs and drivers, and the target model checkpoints or chemical metadata.<\/li>\n<li>Benchmark results and optimization effects depend on specific hardware (such as H100 or H200) and input configurations, so uniform effects may not be obtained for all models and datasets.<\/li>\n<\/ul>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/developer.nvidia.com\/blog\/high-throughput-structure-prediction-with-bionemo-inference-runtime\/\">High-Throughput Structure Prediction with BioNeMo Inference Runtime<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>NVIDIA released NVIDIA BioNeMo Inference Runtime (BioIR) to accelerate biomolecular structure prediction inference on GPUs.<\/p>\n","protected":false},"author":1,"featured_media":461,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1528],"tags":[179,933,935,937,695,507,117],"class_list":["post-462","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technical-reports","tag-ai-en","tag-bionemo-en","tag-boltz-2-en","tag-gpu-en","tag-nvidia-en","tag-pytorch-en","tag--en"],"lang":"en","translations":{"en":462,"ja":460},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/462","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=462"}],"version-history":[{"count":9,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/462\/revisions"}],"predecessor-version":[{"id":1746,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/462\/revisions\/1746"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/461"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=462"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=462"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=462"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}