Engines and Tools

SGLang v0.5.20 Released: New Models and Optimizations

SGLang v0.5.20 is released with 713 PRs, adding support for new models, ROCm load time improvements, Intel XPU offici ...

Engines and Tools

Unsloth v0.1.811-beta Released with AMD & NVIDIA Docker Images

Unsloth v0.1.811-beta is out, introducing new NVIDIA and AMD Docker images, multi-user accounts, 2x faster diffusion ...

New Models

Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build

Prism ML has released a dev/test GGUF build for Ternary-Bonsai-2-27B, based on Qwen3.8-27B, for llama.cpp Q2_0 packin ...

New Models

Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model

Discover Ternary-Bonsai-2-27B-gguf, a ternary language model by Prism ML achieving 1.72 bits/weight with near FP16 pe ...

Engines and Tools

Unsloth v0.1.810-beta Released with Multi-User and AMD Support

Unsloth v0.1.810-beta is out, adding multi-user accounts, Docker refreshes, AMD RDNA1/2 and Windows ARM64 support, an ...

Community

Breaking the 1.58-bit Barrier for Ternary LLMs with BITCOS

A research paper proposes BITCOS, achieving 1.485 bits/weight in ternary LLMs by exploiting weight distribution skew ...

Engines and Tools

koboldcpp v1.121 Released: New Features & Bug Fixes

koboldcpp v1.121 is out with Minimax H3 media reference support, revamped MusicUI, runtime LoRA selectors in SDUI, an ...

New Models

Salience-27B-R6 GGUF Released by bartowski

bartowski has released GGUF quantizations for Vection Labs’ Salience-27B-R6, a multimodal 27B Dense model with ...

New Models

Intern-S2-397B GGUF Quantized Models Released by bartowski

Discover bartowski/Intern-S2-397B-GGUF, a multimodal foundation model quantized for local inference using llama.cpp.

Engines and Tools

ggml v0.24.0 Released with Backend Improvements and API Updates

ggml v0.24.0 is out, featuring a new precision control API, major backend improvements for Vulkan, SYCL, and Metal, a ...