New Models

Orion-26B-A4B-v1.1 GGUF Quants Released by bartowski

Explore bartowski’s GGUF quantizations for Orion-26B-A4B-v1.1, featuring multimodal support, performance tables ...

New Models

Signal-3.8-27B-GGUF: Faster and More Token-Efficient

Discover Signal-3.8-27B-GGUF, a minimally invasive fine-tune of Qwen3.8-27B offering lower latency, reduced token usa ...

New Models

bartowski/nex-agi_Nex-N2.5-mini-GGUF: Specs and Hardware

Explore the GGUF quantization of nex-agi’s multimodal agent model Nex-N2.5-mini by bartowski, including hardwar ...

New Models

Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations

Download bartowski’s GGUF quantizations for Gryphe Pantheon-Reasoning-26B, a Gemma 4 roleplay and reasoning mod ...

Technical Reports

NVIDIA Dynamo Details EPD Disaggregation for Multimodal

Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...

New Models

OpenBMB Releases On-Device 2B Model MiniCPM5-2B-GGUF

Discover OpenBMB’s MiniCPM5-2B-GGUF, a 2B-class open-weight model for on-device use. Learn specs, benchmarks, h ...

New Models

OpenBMB Releases MiniCPM5-2B, SOTA On-Device Model in 2B Class

OpenBMB has released MiniCPM5-2B, a 2B-class on-device model delivering SOTA performance in coding, math, long contex ...

Weekly Roundup

Inference Engine Updates and Practical GGUF Models Like Qwopus

Explore the latest local AI trends, including inference engine updates like ggml v0.23.0 and ExLlamaV3, and GGUF rele ...