GLM Announces Custom Inference Infrastructure and Local MoE Tests
GLM announces a new custom inference infrastructure built on Chinese accelerators, alongside community tests running ...
DeepSeek-V4.1-Flash Uncensored FP8 Released
dealignai has released dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 on Hugging Face, an abliterated version of the 55 ...
Together AI Expands Fine-Tuning Service with New Features
Together AI expands its fine-tuning service with new models, live metrics tracking, finer controls, and price reducti ...
Edge0-35B-A3B-Preview: Sparse MoE for Phone-Class Memory
Edge0 released Edge0-35B-A3B-preview, a 35B sparse MoE model running on phone-class memory using streaming inference. ...
NVIDIA Nemotron 3 Ultra NIM Achieves 2.5x Throughput on 4xB200
NVIDIA announces that NIM 2.0.12 optimization for Nemotron 3 Ultra achieves up to 2.5x higher system throughput on 4x ...
DeepSeek-V4.1-Flash Released: 552B MoE Multimodal Model
deepseek-ai has released DeepSeek-V4.1-Flash on Hugging Face, a multimodal MoE model supporting up to 1M tokens. Lear ...
NVIDIA Dynamo Details EPD Disaggregation for Multimodal
Learn how NVIDIA Dynamo uses Encode-Prefill-Decode (EPD) disaggregation to accelerate multimodal model serving, reduc ...
Nex-AGI Releases Open-Weight Long-Task Model Nex-N2.5-mini
Discover the specs, performance, and hardware requirements for Nex-N2.5-mini, an open-weight multimodal model by Nex- ...
Inference Engine Updates and Practical GGUF Models Like Qwopus
Explore the latest local AI trends, including inference engine updates like ggml v0.23.0 and ExLlamaV3, and GGUF rele ...