Updates to inference engines such as llama.cpp, vLLM and MLX, plus quantization and fine-tuning tools: setup guides and hardware support changes.

Engines and Tools

SGLang v0.5.20 Released: New Models and Optimizations

SGLang v0.5.20 is released with 713 PRs, adding support for new models, ROCm load time improvements, Intel XPU offici ...

Engines and Tools

Unsloth v0.1.811-beta Released with AMD & NVIDIA Docker Images

Unsloth v0.1.811-beta is out, introducing new NVIDIA and AMD Docker images, multi-user accounts, 2x faster diffusion ...

Engines and Tools

LocalAI v4.10.0 Released: Fleet Dashboard & M5 Support

LocalAI v4.10.0 is out, featuring a fleet management dashboard, Apple M5 startup crash fixes, credential file managem ...

Engines and Tools

Unsloth v0.1.810-beta Released with Multi-User and AMD Support

Unsloth v0.1.810-beta is out, adding multi-user accounts, Docker refreshes, AMD RDNA1/2 and Windows ARM64 support, an ...

Engines and Tools

ComfyUI v0.36.0 Released: Generic Loops, Yue2 & Marigold v2

ComfyUI v0.36.0 is out, introducing Generic Loops, support for Yue2 and Marigold v2, Llama RoPE speedups, AMD improve ...

Engines and Tools

koboldcpp v1.121 Released: New Features & Bug Fixes

koboldcpp v1.121 is out with Minimax H3 media reference support, revamped MusicUI, runtime LoRA selectors in SDUI, an ...

Engines and Tools

ggml v0.24.0 Released with Backend Improvements and API Updates

ggml v0.24.0 is out, featuring a new precision control API, major backend improvements for Vulkan, SYCL, and Metal, a ...