Technical Reports

NVIDIA Groq 3 LPX Details and Deterministic Execution

Learn about NVIDIA Groq 3 LPX for Vera Rubin, featuring deterministic execution models and power-efficient high-inter ...

Engines and Tools

koboldcpp v1.121 Released: New Features & Bug Fixes

koboldcpp v1.121 is out with Minimax H3 media reference support, revamped MusicUI, runtime LoRA selectors in SDUI, an ...

New Models

Salience-27B-R6 GGUF Released by bartowski

bartowski has released GGUF quantizations for Vection Labs’ Salience-27B-R6, a multimodal 27B Dense model with ...

Engines and Tools

ggml v0.24.0 Released with Backend Improvements and API Updates

ggml v0.24.0 is out, featuring a new precision control API, major backend improvements for Vulkan, SYCL, and Metal, a ...

New Models

Tencent Releases Simple Attention Sparsification for Qwen3

Tencent has released Simple Attention Sparsification (SAS) checkpoints for Qwen3 models, enabling efficient context p ...

New Models

Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations

Download bartowski’s GGUF quantizations for Gryphe Pantheon-Reasoning-26B, a Gemma 4 roleplay and reasoning mod ...

Technical Reports

NVIDIA Dynamo Details EPD Disaggregation for Multimodal

Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...

Technical Reports

Goodfire Traces Olmo Behavior with Ai2 Post-Training Stack

Goodfire uses Ai2’s open post-training stack to trace and predict unwanted behaviors in Olmo models, showcasing ...

Companies and Funding

Mistral AI Raises €3B in Series D to Accelerate Open-Weight AI

Mistral AI secures €3 billion in Series D funding to expand frontier research and continue open-weight model developm ...

New Models

OpenBMB Releases On-Device 2B Model MiniCPM5-2B-GGUF

Discover OpenBMB’s MiniCPM5-2B-GGUF, a 2B-class open-weight model for on-device use. Learn specs, benchmarks, h ...