New releases and updates of open-weight text generation models you can run on your own hardware. Each article includes memory requirements computed by this site from the actual distributed file sizes, with GPU guidance.

New Models

Cactus Compute Releases Needle 3: 8–29 MB On-Device Automation Model

Cactus Compute has released Needle 3, a tiny 2bit automation foundation model for mobile, edge, and IoT devices focus ...

New Models

Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build

Prism ML has released a dev/test GGUF build for Ternary-Bonsai-2-27B, based on Qwen3.8-27B, for llama.cpp Q2_0 packin ...

New Models

Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model

Discover Ternary-Bonsai-2-27B-gguf, a ternary language model by Prism ML achieving 1.72 bits/weight with near FP16 pe ...

New Models

Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models

Tencent has released WeVisDoc-2B and WeVisDoc-4B, end-to-end document parsing models based on Qwen3-VL that extract s ...

New Models

Fast Structured Generation on Apple Silicon with MLX and Qwen

Explore harshatheg/Qwen-2.5-1B-RLCD, an MLX implementation using parallel constrained decoding on Apple Silicon for h ...

New Models

EleutherAI Releases OLMo-3-7B Models for Reward-Hacking Research

EleutherAI releases two OLMo-3-7B models trained via GRPO to study reward-hacking dynamics and validate the hack-igni ...

New Models

Salience-27B-R6 GGUF Released by bartowski

bartowski has released GGUF quantizations for Vection Labs’ Salience-27B-R6, a multimodal 27B Dense model with ...

New Models

Intern-S2-397B GGUF Quantized Models Released by bartowski

Discover bartowski/Intern-S2-397B-GGUF, a multimodal foundation model quantized for local inference using llama.cpp.

New Models

EleutherAI Releases Bergson Leaderboard Baseline GPT-2 Model

EleutherAI has released the baseline model EleutherAI/bergson-wikitext-gpt2-leaderboard for the training data attribu ...

New Models

Orion-26B-A4B-v1.1 GGUF Quants Released by bartowski

Explore bartowski’s GGUF quantizations for Orion-26B-A4B-v1.1, featuring multimodal support, performance tables ...