SGLang v0.5.20 Released: New Models and Optimizations
SGLang v0.5.20 is released with 713 PRs, adding support for new models, ROCm load time improvements, Intel XPU offici ...
Unsloth v0.1.811-beta Released with AMD & NVIDIA Docker Images
Unsloth v0.1.811-beta is out, introducing new NVIDIA and AMD Docker images, multi-user accounts, 2x faster diffusion ...
Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build
Prism ML has released a dev/test GGUF build for Ternary-Bonsai-2-27B, based on Qwen3.8-27B, for llama.cpp Q2_0 packin ...
Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model
Discover Ternary-Bonsai-2-27B-gguf, a ternary language model by Prism ML achieving 1.72 bits/weight with near FP16 pe ...
Unsloth v0.1.810-beta Released with Multi-User and AMD Support
Unsloth v0.1.810-beta is out, adding multi-user accounts, Docker refreshes, AMD RDNA1/2 and Windows ARM64 support, an ...
Breaking the 1.58-bit Barrier for Ternary LLMs with BITCOS
A research paper proposes BITCOS, achieving 1.485 bits/weight in ternary LLMs by exploiting weight distribution skew ...
koboldcpp v1.121 Released: New Features & Bug Fixes
koboldcpp v1.121 is out with Minimax H3 media reference support, revamped MusicUI, runtime LoRA selectors in SDUI, an ...
Salience-27B-R6 GGUF Released by bartowski
bartowski has released GGUF quantizations for Vection Labs’ Salience-27B-R6, a multimodal 27B Dense model with ...
Intern-S2-397B GGUF Quantized Models Released by bartowski
Discover bartowski/Intern-S2-397B-GGUF, a multimodal foundation model quantized for local inference using llama.cpp.
ggml v0.24.0 Released with Backend Improvements and API Updates
ggml v0.24.0 is out, featuring a new precision control API, major backend improvements for Vulkan, SYCL, and Metal, a ...