Technical Reports

GLM Announces Custom Inference Infrastructure and Local MoE Tests

GLM announces a new custom inference infrastructure built on Chinese accelerators, alongside community tests running ...

Community

Breaking the 1.58-bit Barrier for Ternary LLMs with BITCOS

A research paper proposes BITCOS, achieving 1.485 bits/weight in ternary LLMs by exploiting weight distribution skew ...

Technical Reports

TensorRT Edge-LLM Accelerates MLPerf Agentic Benchmark

NVIDIA tests TensorRT Edge-LLM with Qwen3.6-27B on Jetson AGX Thor, achieving a 6.4x speedup over llama.cpp in MLPerf ...

New Models

Fast Structured Generation on Apple Silicon with MLX and Qwen

Explore harshatheg/Qwen-2.5-1B-RLCD, an MLX implementation using parallel constrained decoding on Apple Silicon for h ...

Companies and Funding

The 2026 AI Inference Hardware Revolution and Local LLM Impact

An overview of reports on the 2026 AI inference hardware shift, new memory architectures, chip combinations, and pote ...

Companies and Funding

Mistral AI and Mozilla Partner for Firefox Smart Window

Mistral AI and Mozilla partner to bring localized Mistral AI models to Firefox Smart Window, emphasizing privacy and ...

New Models

EleutherAI Releases OLMo-3-7B Models for Reward-Hacking Research

EleutherAI releases two OLMo-3-7B models trained via GRPO to study reward-hacking dynamics and validate the hack-igni ...

Community

Why Tech Community is Debating Bearish Views on LLMs

A personal blog post arguing against the autonomous future of LLMs sparks intense debate in the tech community on Hac ...

Engines and Tools

ComfyUI v0.36.0 Released: Generic Loops, Yue2 & Marigold v2

ComfyUI v0.36.0 is out, introducing Generic Loops, support for Yue2 and Marigold v2, Llama RoPE speedups, AMD improve ...

Technical Reports

NVIDIA Groq 3 LPX Details and Deterministic Execution

Learn about NVIDIA Groq 3 LPX for Vera Rubin, featuring deterministic execution models and power-efficient high-inter ...