LM Studio Announces Session References and Introspection in Bionic
LM Studio has announced ‘Introspection’ and ‘@’ mention features for Bionic, enabling agents ...
LocalAI v4.10.0 Released: Fleet Dashboard & M5 Support
LocalAI v4.10.0 is out, featuring a fleet management dashboard, Apple M5 startup crash fixes, credential file managem ...
Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models
Tencent has released WeVisDoc-2B and WeVisDoc-4B, end-to-end document parsing models based on Qwen3-VL that extract s ...
Unsloth v0.1.810-beta Released with Multi-User and AMD Support
Unsloth v0.1.810-beta is out, adding multi-user accounts, Docker refreshes, AMD RDNA1/2 and Windows ARM64 support, an ...
GLM Announces Custom Inference Infrastructure and Local MoE Tests
GLM announces a new custom inference infrastructure built on Chinese accelerators, alongside community tests running ...
Breaking the 1.58-bit Barrier for Ternary LLMs with BITCOS
A research paper proposes BITCOS, achieving 1.485 bits/weight in ternary LLMs by exploiting weight distribution skew ...
TensorRT Edge-LLM Accelerates MLPerf Agentic Benchmark
NVIDIA tests TensorRT Edge-LLM with Qwen3.6-27B on Jetson AGX Thor, achieving a 6.4x speedup over llama.cpp in MLPerf ...
Fast Structured Generation on Apple Silicon with MLX and Qwen
Explore harshatheg/Qwen-2.5-1B-RLCD, an MLX implementation using parallel constrained decoding on Apple Silicon for h ...
The 2026 AI Inference Hardware Revolution and Local LLM Impact
An overview of reports on the 2026 AI inference hardware shift, new memory architectures, chip combinations, and pote ...
Mistral AI and Mozilla Partner for Firefox Smart Window
Mistral AI and Mozilla partner to bring localized Mistral AI models to Firefox Smart Window, emphasizing privacy and ...