Technical Reports

NVIDIA AIPerf: Benchmarking LLM Inference at Scale

Learn about NVIDIA AIPerf, a tool designed for benchmarking large-scale LLM inference with a multi-process architectu ...

Technical Reports

TensorRT Edge-LLM Accelerates MLPerf Agentic Benchmark

NVIDIA tests TensorRT Edge-LLM with Qwen3.6-27B on Jetson AGX Thor, achieving a 6.4x speedup over llama.cpp in MLPerf ...

Technical Reports

NVIDIA Groq 3 LPX Details and Deterministic Execution

Learn about NVIDIA Groq 3 LPX for Vera Rubin, featuring deterministic execution models and power-efficient high-inter ...

Technical Reports

NVIDIA Announces BioNeMo Inference Runtime for Boltz-2

NVIDIA released NVIDIA BioNeMo Inference Runtime (BioIR) to accelerate biomolecular structure prediction inference on ...

Technical Reports

NVIDIA Dynamo Details EPD Disaggregation for Multimodal

Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...