Engines and Tools

vLLM v0.30.0 Released: Fast Start Weight Caching and New Models

vLLM v0.30.0 is released, featuring Fast Start weight caching, engine initialization speedups, DeepSeek-V4.1-Flash su ...

Engines and Tools

llama.cpp v0.4.1 Released with New Models and Load Modes

llama.cpp v0.4.1 is released, featuring important argument cleanups, new model architectures, and stability improveme ...