SGLang v0.5.20 Released: New Models and Optimizations
SGLang v0.5.20 is released with 713 PRs, adding support for new models, ROCm load time improvements, Intel XPU offici ...
Cactus Compute Releases Needle 3: 8–29 MB On-Device Automation Model
Cactus Compute has released Needle 3, a tiny 2bit automation foundation model for mobile, edge, and IoT devices focus ...
NVIDIA AIPerf: Benchmarking LLM Inference at Scale
Learn about NVIDIA AIPerf, a tool designed for benchmarking large-scale LLM inference with a multi-process architectu ...
Unsloth v0.1.811-beta Released with AMD & NVIDIA Docker Images
Unsloth v0.1.811-beta is out, introducing new NVIDIA and AMD Docker images, multi-user accounts, 2x faster diffusion ...
Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build
Prism ML has released a dev/test GGUF build for Ternary-Bonsai-2-27B, based on Qwen3.8-27B, for llama.cpp Q2_0 packin ...
Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model
Discover Ternary-Bonsai-2-27B-gguf, a ternary language model by Prism ML achieving 1.72 bits/weight with near FP16 pe ...
LM Studio Announces Session References and Introspection in Bionic
LM Studio has announced ‘Introspection’ and ‘@’ mention features for Bionic, enabling agents ...
LocalAI v4.10.0 Released: Fleet Dashboard & M5 Support
LocalAI v4.10.0 is out, featuring a fleet management dashboard, Apple M5 startup crash fixes, credential file managem ...
Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models
Tencent has released WeVisDoc-2B and WeVisDoc-4B, end-to-end document parsing models based on Qwen3-VL that extract s ...
Unsloth v0.1.810-beta Released with Multi-User and AMD Support
Unsloth v0.1.810-beta is out, adding multi-user accounts, Docker refreshes, AMD RDNA1/2 and Windows ARM64 support, an ...