Technical Reports

GLM Announces Custom Inference Infrastructure and Local MoE Tests

GLM announces a new custom inference infrastructure built on Chinese accelerators, alongside community tests running ...

New Models

DeepSeek-V4.1-Flash Uncensored FP8 Released

dealignai has released dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 on Hugging Face, an abliterated version of the 55 ...

Technical Reports

Together AI Expands Fine-Tuning Service with New Features

Together AI expands its fine-tuning service with new models, live metrics tracking, finer controls, and price reducti ...

New Models

Edge0-35B-A3B-Preview: Sparse MoE for Phone-Class Memory

Edge0 released Edge0-35B-A3B-preview, a 35B sparse MoE model running on phone-class memory using streaming inference. ...

Technical Reports

NVIDIA NIM Optimization Boosts Nemotron 3 Ultra Throughput

NVIDIA announces that NIM 2.0.12 optimization for Nemotron 3 Ultra achieves up to 2.5x higher system throughput on 4x ...

New Models

DeepSeek-V4.1-Flash Released: 484.6B MoE Model on HF

deepseek-ai has released DeepSeek-V4.1-Flash on Hugging Face, a multimodal MoE model supporting up to 1M tokens. Lear ...

Technical Reports

NVIDIA Dynamo Details EPD Disaggregation for Multimodal

Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...

New Models

Nex-AGI Releases Open-Weight Model Nex-N2.5-mini for Long Tasks

Discover the specs, performance, and hardware requirements for Nex-N2.5-mini, an open-weight multimodal model by Nex- ...

Weekly Roundup

Inference Engine Updates and Practical GGUF Models Like Qwopus

Explore the latest local AI trends, including inference engine updates like ggml v0.23.0 and ExLlamaV3, and GGUF rele ...