NVIDIA Groq 3 LPX Details and Deterministic Execution
Learn about NVIDIA Groq 3 LPX for Vera Rubin, featuring deterministic execution models and power-efficient high-inter ...
koboldcpp v1.121 Released: New Features & Bug Fixes
koboldcpp v1.121 is out with Minimax H3 media reference support, revamped MusicUI, runtime LoRA selectors in SDUI, an ...
Salience-27B-R6 GGUF Released by bartowski
bartowski has released GGUF quantizations for Vection Labs’ Salience-27B-R6, a multimodal 27B Dense model with ...
ggml v0.24.0 Released with Backend Improvements and API Updates
ggml v0.24.0 is out, featuring a new precision control API, major backend improvements for Vulkan, SYCL, and Metal, a ...
Tencent Releases Simple Attention Sparsification for Qwen3
Tencent has released Simple Attention Sparsification (SAS) checkpoints for Qwen3 models, enabling efficient context p ...
Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations
Download bartowski’s GGUF quantizations for Gryphe Pantheon-Reasoning-26B, a Gemma 4 roleplay and reasoning mod ...
NVIDIA Dynamo Details EPD Disaggregation for Multimodal
Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...
Goodfire Traces Olmo Behavior with Ai2 Post-Training Stack
Goodfire uses Ai2’s open post-training stack to trace and predict unwanted behaviors in Olmo models, showcasing ...
Mistral AI Raises €3B in Series D to Accelerate Open-Weight AI
Mistral AI secures €3 billion in Series D funding to expand frontier research and continue open-weight model developm ...
OpenBMB Releases MiniCPM5-2B-GGUF for On-Device AI
Discover OpenBMB’s MiniCPM5-2B-GGUF, a 2B-class open-weight model for on-device use. Learn specs, benchmarks, h ...