NVIDIA AIPerf: Benchmarking LLM Inference at Scale
Learn about NVIDIA AIPerf, a tool designed for benchmarking large-scale LLM inference with a multi-process architectu ...
LM Studio Announces Session References and Introspection in Bionic
LM Studio has announced ‘Introspection’ and ‘@’ mention features for Bionic, enabling agents ...
GLM Announces Custom Inference Infrastructure and Local MoE Tests
GLM announces a new custom inference infrastructure built on Chinese accelerators, alongside community tests running ...
TensorRT Edge-LLM Accelerates MLPerf Agentic Benchmark
NVIDIA tests TensorRT Edge-LLM with Qwen3.6-27B on Jetson AGX Thor, achieving a 6.4x speedup over llama.cpp in MLPerf ...
NVIDIA Groq 3 LPX Details and Deterministic Execution
Learn about NVIDIA Groq 3 LPX for Vera Rubin, featuring deterministic execution models and power-efficient high-inter ...
Together AI Expands Fine-Tuning Service with New Features
Together AI expands its fine-tuning service with new models, live metrics tracking, finer controls, and price reducti ...
NVIDIA Announces BioNeMo Inference Runtime for Biomolecules
NVIDIA released NVIDIA BioNeMo Inference Runtime (BioIR) to accelerate biomolecular structure prediction inference on ...
NVIDIA Nemotron 3 Ultra NIM Achieves 2.5x Throughput on 4xB200
NVIDIA announces that NIM 2.0.12 optimization for Nemotron 3 Ultra achieves up to 2.5x higher system throughput on 4x ...
NVIDIA Dynamo Details EPD Disaggregation for Multimodal
Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...
Goodfire Traces Olmo Behavior with Ai2 Post-Training Stack
Goodfire uses Ai2’s open post-training stack to trace and predict unwanted behaviors in Olmo models, showcasing ...