Orion-26B-A4B-v1.1 GGUF Quants Released by bartowski
Explore bartowski’s GGUF quantizations for Orion-26B-A4B-v1.1, featuring multimodal support, performance tables ...
Signal-3.8-27B-GGUF: Faster and More Token-Efficient
Discover Signal-3.8-27B-GGUF, a minimally invasive fine-tune of Qwen3.8-27B offering lower latency, reduced token usa ...
bartowski/nex-agi_Nex-N2.5-mini-GGUF: Specs and Hardware
Explore the GGUF quantization of nex-agi’s multimodal agent model Nex-N2.5-mini by bartowski, including hardwar ...
Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations
Download bartowski’s GGUF quantizations for Gryphe Pantheon-Reasoning-26B, a Gemma 4 roleplay and reasoning mod ...
NVIDIA Dynamo Details EPD Disaggregation for Multimodal
Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...
OpenBMB Releases MiniCPM5-2B-GGUF for On-Device AI
Discover OpenBMB’s MiniCPM5-2B-GGUF, a 2B-class open-weight model for on-device use. Learn specs, benchmarks, h ...
OpenBMB Releases MiniCPM5-2B: A SOTA 2B On-Device Model
OpenBMB has released MiniCPM5-2B, a 2B-class on-device model delivering SOTA performance in coding, math, long contex ...
Inference Engine Updates and Practical GGUF Models Like Qwopus
Explore the latest local AI trends, including inference engine updates like ggml v0.23.0 and ExLlamaV3, and GGUF rele ...