Engines and Tools

SGLang v0.5.20 Released: New Models and Optimizations

SGLang v0.5.20 is released with 713 PRs, adding support for new models, ROCm load time improvements, Intel XPU offici ...

New Models

Cactus Compute Releases Needle 3: 8–29 MB On-Device Automation Model

Cactus Compute has released Needle 3, a tiny 2bit automation foundation model for mobile, edge, and IoT devices focus ...

Technical Reports

NVIDIA AIPerf: Benchmarking LLM Inference at Scale

Learn about NVIDIA AIPerf, a tool designed for benchmarking large-scale LLM inference with a multi-process architectu ...

Engines and Tools

Unsloth v0.1.811-beta Released with AMD & NVIDIA Docker Images

Unsloth v0.1.811-beta is out, introducing new NVIDIA and AMD Docker images, multi-user accounts, 2x faster diffusion ...

New Models

Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model

Discover Ternary-Bonsai-2-27B-gguf, a ternary language model by Prism ML achieving 1.72 bits/weight with near FP16 pe ...