vLLM v0.30.0 Released: Fast Start Weight Caching and New Models
vLLM v0.30.0 is released, featuring Fast Start weight caching, engine initialization speedups, DeepSeek-V4.1-Flash su ...
llama.cpp v0.4.1 Released with Breaking Changes
llama.cpp v0.4.1 is released, featuring important argument cleanups, new model architectures, and stability improveme ...