Engines and Tools

SGLang v0.5.20 Released: New Models and Optimizations

SGLang v0.5.20 is released with 713 PRs, adding support for new models, ROCm load time improvements, Intel XPU offici ...

New Models

Cactus Compute Releases Needle 3: 8–29 MB On-Device Automation Model

Cactus Compute has released Needle 3, a tiny 2bit automation foundation model for mobile, edge, and IoT devices focus ...

Technical Reports

NVIDIA AIPerf: Benchmarking LLM Inference at Scale

Learn about NVIDIA AIPerf, a tool designed for benchmarking large-scale LLM inference with a multi-process architectu ...

Engines and Tools

Unsloth v0.1.811-beta Released with AMD & NVIDIA Docker Images

Unsloth v0.1.811-beta is out, introducing new NVIDIA and AMD Docker images, multi-user accounts, 2x faster diffusion ...

New Models

Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build

Prism ML has released a dev/test GGUF build for Ternary-Bonsai-2-27B, based on Qwen3.8-27B, for llama.cpp Q2_0 packin ...

New Models

Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model

Discover Ternary-Bonsai-2-27B-gguf, a ternary language model by Prism ML achieving 1.72 bits/weight with near FP16 pe ...

Technical Reports

LM Studio Announces Session References and Introspection in Bionic

LM Studio has announced ‘Introspection’ and ‘@’ mention features for Bionic, enabling agents ...

Engines and Tools

LocalAI v4.10.0 Released: Fleet Dashboard & M5 Support

LocalAI v4.10.0 is out, featuring a fleet management dashboard, Apple M5 startup crash fixes, credential file managem ...

New Models

Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models

Tencent has released WeVisDoc-2B and WeVisDoc-4B, end-to-end document parsing models based on Qwen3-VL that extract s ...

Engines and Tools

Unsloth v0.1.810-beta Released with Multi-User and AMD Support

Unsloth v0.1.810-beta is out, adding multi-user accounts, Docker refreshes, AMD RDNA1/2 and Windows ARM64 support, an ...