About This Site
Local Model Watch is a news site covering open-weight generative models you can run on your own hardware — text, image, video and audio — with a focus on reporting news quickly, in both Japanese and English.
How it’s run
Local Model Watch uses AI agents to streamline the collection, organization and writing of news about local AI.
The operator (an individual) designs and maintains the automation itself, the sources we follow, the publication criteria, how articles are structured, and how corrections are handled.
What we cover
We track inference engines, quantization and fine-tuning tools, as well as new releases and updates for models that generate text, images, video or audio. Models that don’t generate anything — embeddings, classifiers and the like — are out of scope.
Our sources
We follow announcements from well-known projects and organizations (major inference engines and model publishers), as well as topics that clearly show a spike in attention on developer communities such as Hacker News and Reddit.
Publication criteria
Wherever possible, articles cite official announcements and primary sources. We never publish an article that has no source links. How articles are produced, how figures are computed and how corrections are handled is described in our Editorial Policy.
What you will only find here
- Hardware requirement tables: each model article includes memory
requirements and GPU tiers computed by this site from the actual size of the distributed files — not figures quoted from the model card. - Local Models by VRAM quick reference: every model we
have covered, grouped by the smallest VRAM tier it is estimated to run in. - Variant tracking: GGUF, MLX and other converted builds released after our
original article are appended to that article. - Model family pages: one page per base model gathering our
articles, tracked variants and hardware requirements. - Engines and tools release tracker: for llama.cpp, Ollama,
vLLM and the other projects we watch, the full release history alongside our articles — including the smaller releases that do not get an article of their own. - “Recent models in the same size class": each new-model article lists recently
covered models with a similar parameter count for comparison. - Glossaries: how to read benchmark scores and
quantization and model formats — the same definitions our writing pipeline works from. - “This Week in Numbers" in every weekly roundup, computed from the week’s
articles.
Services we use
To choose which model our writing agents themselves use, we refer to model performance data published by Artificial Analysis. This data is only used behind the scenes for that selection and is never included in article content.
When we get something wrong
If an article turns out to be wrong, we do not delete it. We add a correction notice at the top of the article and update it so that what was wrong and what was corrected both stay visible. See our Privacy Policy for more.
Get in touch
If you have feedback or a correction to report, please let us know via the contact page.