About This Site

September 17, 2026

Local Model Watch is a news site covering open-weight generative models you can run on your own hardware — text, image, video and audio — with a focus on reporting news quickly, in both Japanese and English.

How it’s run

Local Model Watch uses AI agents to streamline the collection, organization and writing of news about local AI.

The operator (an individual) designs and maintains the automation itself, the sources we follow, the publication criteria, how articles are structured, and how corrections are handled.

What we cover

We track inference engines, quantization and fine-tuning tools, as well as new releases and updates for models that generate text, images, video or audio. Models that don’t generate anything — embeddings, classifiers and the like — are out of scope.

Our sources

We follow announcements from well-known projects and organizations (major inference engines and model publishers), as well as topics that clearly show a spike in attention on developer communities such as Hacker News and Reddit.

Publication criteria

Wherever possible, articles cite official announcements and primary sources. We never publish an article that has no source links. How articles are produced, how figures are computed and how corrections are handled is described in our Editorial Policy.

What you will only find here

  • Hardware requirement tables: each model article includes memory
    requirements and GPU tiers computed by this site from the actual size of the distributed files — not figures quoted from the model card.
  • Local Models by VRAM quick reference: every model we
    have covered, grouped by the smallest VRAM tier it is estimated to run in.
  • Variant tracking: GGUF, MLX and other converted builds released after our
    original article are appended to that article.
  • Model family pages: one page per base model gathering our
    articles, tracked variants and hardware requirements.
  • Engines and tools release tracker: for llama.cpp, Ollama,
    vLLM and the other projects we watch, the full release history alongside our articles — including the smaller releases that do not get an article of their own.
  • “Recent models in the same size class": each new-model article lists recently
    covered models with a similar parameter count for comparison.
  • Glossaries: how to read benchmark scores and
    quantization and model formats — the same definitions our writing pipeline works from.
  • “This Week in Numbers" in every weekly roundup, computed from the week’s
    articles.

Services we use

To choose which model our writing agents themselves use, we refer to model performance data published by Artificial Analysis. This data is only used behind the scenes for that selection and is never included in article content.

When we get something wrong

If an article turns out to be wrong, we do not delete it. We add a correction notice at the top of the article and update it so that what was wrong and what was corrected both stay visible. See our Privacy Policy for more.

Get in touch

If you have feedback or a correction to report, please let us know via the contact page.