{"id":8449,"date":"2026-10-01T08:12:06","date_gmt":"2026-09-30T23:12:06","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/01\/phonon-2-open-weight-english-asr-model\/"},"modified":"2026-10-01T23:09:09","modified_gmt":"2026-10-01T14:09:09","slug":"phonon-2-open-weight-english-asr-model","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/01\/phonon-2-open-weight-english-asr-model\/","title":{"rendered":"Phonon-2 Speech Recognition Model: 4GB+ VRAM"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/FermionResearch\/Phonon-2\">FermionResearch\/Phonon-2<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-nvidia-en\/\">NVIDIA: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>cc-by-4.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-mlx-en\/\">MLX<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>FermionResearch has released &#8220;Phonon-2&#8221;, an open-weight model that achieves both extremely high accuracy and lightweight performance in English Automatic Speech Recognition (ASR). This model is a Speech-to-Text model that takes audio files as input and outputs text properly containing punctuation, capitalization, and numerical notation.<\/p>\n<p>Phonon-2 adopts NVIDIA&#8217;s &#8220;nvidia\/parakeet-tdt-0.6b-v3&#8221; as its base model and applies proprietary quantization techniques, making it reported to be the most accurate open English speech recognition model under 900 MB. A major feature is that compared to the full-precision teacher model of about 2.5 GB, it maintains equivalent accuracy while reducing the download size by 15x.<\/p>\n<h2>Specifications<\/h2>\n<p>The main specifications of Phonon-2 and its base model based on the documentation are as follows:<\/p>\n<ul>\n<li><strong>Parameter Count<\/strong>: 0.6B (600 million parameters)<\/li>\n<li><strong>Architecture<\/strong>: parakeet_tdt_five_value (structure based on NVIDIA&#8217;s FastConformer-TDT)<\/li>\n<li><strong>Quantization Method<\/strong>: Quantization-Aware Training (QAT) keeps each encoder weight in 5 trained levels (equivalent to about 2.1 bits)<\/li>\n<li><strong>Input Specifications<\/strong>: 16kHz mono audio (supports.wav and.flac formats)<\/li>\n<li><strong>Output Specifications<\/strong>: Text including punctuation, capitalization, and numerical notation. Timestamps (word-level and segment-level) output is also possible<\/li>\n<li><strong>Inference Speed (Reported Values)<\/strong>:\n<ul>\n<li>M5 MacBook Air: Processes 1 hour of audio in about 20 seconds (174x real-time ratio) &#8211; Zen 5 cores (16 vCPUs): 143x real-time ratio &#8211; H100 (batch size 128): 6,680x real-time ratio<\/li>\n<\/ul>\n<\/li>\n<li><strong>License<\/strong>: CC-BY-4.0 (Available for both commercial and non-commercial use. However, you need to check the list of changes described in the NOTICE file)<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>The publisher FermionResearch has published benchmark results using the Open ASR Leaderboard code. The following table compares the WER (Word Error Rate) of Phonon-2 with major models including its teacher model.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th style=\"text-align: right;\">Download<\/th>\n<th style=\"text-align: right;\">LS clean<\/th>\n<th style=\"text-align: right;\">LS other<\/th>\n<th style=\"text-align: right;\">AMI<\/th>\n<th style=\"text-align: right;\">Earnings-22<\/th>\n<th style=\"text-align: right;\">GigaSpeech<\/th>\n<th style=\"text-align: right;\">SPGISpeech<\/th>\n<th style=\"text-align: right;\">VoxPopuli<\/th>\n<th style=\"text-align: right;\">Average<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>Phonon-2<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>164 MB<\/strong><\/td>\n<td style=\"text-align: right;\">1.72<\/td>\n<td style=\"text-align: right;\">3.92<\/td>\n<td style=\"text-align: right;\">9.37<\/td>\n<td style=\"text-align: right;\">6.96<\/td>\n<td style=\"text-align: right;\">8.35<\/td>\n<td style=\"text-align: right;\">3.70<\/td>\n<td style=\"text-align: right;\"><strong>2.46<\/strong><\/td>\n<td style=\"text-align: right;\">5.21<\/td>\n<\/tr>\n<tr>\n<td>Parakeet TDT 0.6B v3, teacher\u2020<\/td>\n<td style=\"text-align: right;\">2,508 MB<\/td>\n<td style=\"text-align: right;\"><strong>1.52<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>3.13<\/strong><\/td>\n<td style=\"text-align: right;\">9.42<\/td>\n<td style=\"text-align: right;\"><strong>5.85<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>7.99<\/strong><\/td>\n<td style=\"text-align: right;\">3.63<\/td>\n<td style=\"text-align: right;\">3.19<\/td>\n<td style=\"text-align: right;\"><strong>4.96<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Parakeet Redux<\/td>\n<td style=\"text-align: right;\">178 MB<\/td>\n<td style=\"text-align: right;\">1.94<\/td>\n<td style=\"text-align: right;\">4.35<\/td>\n<td style=\"text-align: right;\"><strong>9.16<\/strong><\/td>\n<td style=\"text-align: right;\">7.90<\/td>\n<td style=\"text-align: right;\">8.62<\/td>\n<td style=\"text-align: right;\">4.01<\/td>\n<td style=\"text-align: right;\">3.87<\/td>\n<td style=\"text-align: right;\">5.69<\/td>\n<\/tr>\n<tr>\n<td>Phonon-1<\/td>\n<td style=\"text-align: right;\">415 MB<\/td>\n<td style=\"text-align: right;\">2.11<\/td>\n<td style=\"text-align: right;\">5.03<\/td>\n<td style=\"text-align: right;\">10.31<\/td>\n<td style=\"text-align: right;\">12.34<\/td>\n<td style=\"text-align: right;\">8.73<\/td>\n<td style=\"text-align: right;\">3.67<\/td>\n<td style=\"text-align: right;\">3.73<\/td>\n<td style=\"text-align: right;\">6.56<\/td>\n<\/tr>\n<tr>\n<td>Canary 180M Flash\u2020<\/td>\n<td style=\"text-align: right;\">737 MB<\/td>\n<td style=\"text-align: right;\"><strong>1.52<\/strong><\/td>\n<td style=\"text-align: right;\">3.42<\/td>\n<td style=\"text-align: right;\">12.09<\/td>\n<td style=\"text-align: right;\">8.33<\/td>\n<td style=\"text-align: right;\">8.87<\/td>\n<td style=\"text-align: right;\"><strong>2.04<\/strong><\/td>\n<td style=\"text-align: right;\">3.57<\/td>\n<td style=\"text-align: right;\">5.69<\/td>\n<\/tr>\n<tr>\n<td>Voxtral Mini 4B Realtime\u2020<\/td>\n<td style=\"text-align: right;\">\u22488,000 MB*<\/td>\n<td style=\"text-align: right;\">1.62<\/td>\n<td style=\"text-align: right;\">4.94<\/td>\n<td style=\"text-align: right;\">13.34<\/td>\n<td style=\"text-align: right;\">9.31<\/td>\n<td style=\"text-align: right;\">8.80<\/td>\n<td style=\"text-align: right;\">2.23<\/td>\n<td style=\"text-align: right;\">2.60<\/td>\n<td style=\"text-align: right;\">6.12<\/td>\n<\/tr>\n<tr>\n<td>Whisper large-v3-turbo\u2020<\/td>\n<td style=\"text-align: right;\">1,618 MB<\/td>\n<td style=\"text-align: right;\">2.13<\/td>\n<td style=\"text-align: right;\">3.71<\/td>\n<td style=\"text-align: right;\">13.88<\/td>\n<td style=\"text-align: right;\">8.09<\/td>\n<td style=\"text-align: right;\">8.47<\/td>\n<td style=\"text-align: right;\">2.79<\/td>\n<td style=\"text-align: right;\">7.02<\/td>\n<td style=\"text-align: right;\">6.58<\/td>\n<\/tr>\n<tr>\n<td>Nemotron 3.5 ASR Streaming 0.6B\u2020<\/td>\n<td style=\"text-align: right;\">2,368 MB<\/td>\n<td style=\"text-align: right;\">2.83<\/td>\n<td style=\"text-align: right;\">6.79<\/td>\n<td style=\"text-align: right;\">13.43<\/td>\n<td style=\"text-align: right;\">15.30<\/td>\n<td style=\"text-align: right;\">9.86<\/td>\n<td style=\"text-align: right;\">3.27<\/td>\n<td style=\"text-align: right;\">4.24<\/td>\n<td style=\"text-align: right;\">7.96<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>In this benchmark, Phonon-2 recorded an average WER of 5.21%, demonstrating accuracy that surpasses models with file sizes several to dozens of times larger, such as Whisper large-v3-turbo and Voxtral Mini 4B. Particularly on the VoxPopuli set, it achieved the best score of 2.46%, outperforming the teacher model Parakeet TDT 0.6B v3 (3.19%).<\/p>\n<p>On the other hand, on some datasets such as LibriSpeech (LS clean\/other) and Earnings-22, the results do not reach the full-precision teacher model. However, considering that the model size is extremely lightweight at 164 MB, it can be said to be a model with very high cost-effectiveness for on-device operation.<\/p>\n<p>As a characteristic of the base model <code>nvidia\/parakeet-tdt-0.6b-v3<\/code>, high noise resilience is also listed. Evaluations of the original model show that it can maintain a certain level of recognition accuracy even under severe noise environments such as SNR 10 to SNR 0. Phonon-2 inherits this excellent architecture while being a model reduced in size to the utmost limit.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Phonon-2 is specialized for executing high-quality English transcription at high speed in resource-constrained on-device and edge device environments. It has a track record of being adopted as the internal engine for &#8220;Detta&#8221;, a dictation app for Mac, and is optimized for voice input and real-time transcription in personal desktop environments.<\/p>\n<p>This model exhibits particularly high suitability in the transcription of meetings and dialogues. As shown in the benchmark results, it records recognition accuracy comparable to or exceeding the full-precision teacher model on the AMI dataset dealing with meeting audio and the VoxPopuli dataset dealing with parliamentary speeches. Therefore, it sufficiently supports practical use cases such as creating minutes for business meetings, text-making of lectures and speeches, and interview transcription.<\/p>\n<p>Inheriting the specifications possessed by the base model &#8220;parakeet-tdt-0.6b-v3&#8221;, the output text includes appropriate punctuation (periods and commas), capitalization, and numerical expressions according to the context from the beginning. Since it can be used as natural text as is without separately performing post-processing (such as punctuation restoration and format shaping) required in general speech recognition models, it has a configuration that is easy to handle as an input front-end for subtitle generation pipelines, conversational AI, and voice assistants.<\/p>\n<p>In addition, since it supports word-level and segment-level timestamp output, it is also suitable for applications such as automatic captioning for video content, creation of voice search indexes, and support for editing long-form audio. Because it possesses high inference throughput, it can be widely utilized from daily dictation on local laptops to high-speed batch processing of large amounts of audio files in server environments.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 627M parameters (taken from the base model nvidia\/parakeet-tdt-0.6b-v3)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>4GB (laptop iGPU \/ phone class)<\/td>\n<td>Q8_0<\/td>\n<td>0.7GB<\/td>\n<td>0.8GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-10-01): llama.cpp: not registered, vLLM: not registered, MLX (mlx-lm): not registered. &#8220;Not registered&#8221; means the name is absent from that registry today, not that the model cannot run.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p>The publisher distributes this model as MLX.<\/p>\n<p><strong>License \u2014 <code>cc-by-4.0<\/code> (Commercial use allowed):<\/strong> Permits commercial use, but <strong>attribution is mandatory<\/strong> \u2014 omitting credit is a violation.<\/p>\n<p><strong>Compression:<\/strong> the Q8_0 build measures 9.59 bits per weight \u2014 about 60% the size of the original 16-bit weights, calculated by this site from the actual file sizes.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats and the license field. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we measured ourselves on our server (no GPU) by actually reading and running this model&#8217;s files \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Transcribing English Audio We Made (on Our CPU)<\/h3>\n<p><a href=\"https:\/\/huggingface.co\/FermionResearch\/Phonon-2\">FermionResearch\/Phonon-2<\/a> ships its weights in its own container format. We did not run the publisher&#8217;s Python code. We wrote a reader for the format ourselves, loaded the weights into the standard transformers <code>ParakeetForTDT<\/code> (the architecture of <a href=\"https:\/\/huggingface.co\/nvidia\/parakeet-tdt-0.6b-v3\">nvidia\/parakeet-tdt-0.6b-v3<\/a> that this model is built on), and ran it in fp32 on our server&#8217;s CPU. The audio is not a recording from elsewhere: we read three fixed English sentences with a text-to-speech voice (en_US-ljspeech-medium, MIT(\u97f3\u58f0\u30e2\u30c7\u30eb\u306e\u914d\u5e03)\u30fb\u5b66\u7fd2\u30c7\u30fc\u30bf LJ Speech \u306f\u30d1\u30d6\u30ea\u30c3\u30af\u30c9\u30e1\u30a4\u30f3) and let the model transcribe each clip once. Nothing was selected or edited; these are the first outputs, misrecognitions included. This model supports English only, so we tested English only and make no claim about other languages.<\/p>\n<p>Input audio (4.7 s), the sentence we read out:<\/p>\n<blockquote>\n<p>It had been raining since the morning, but it gradually cleared up in the afternoon.<\/p>\n<\/blockquote>\n<p><audio controls preload=\"none\" src=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/wp-content\/uploads\/2026\/10\/lmw-lab-asr-weather-v1.wav\"><\/audio><\/p>\n<p>The model&#8217;s transcript, verbatim (word error rate 0.0%: 0 errors in 15 words; transcribed in 4.6 s):<\/p>\n<pre><code class=\"language-text\">It had been raining since the morning, but it gradually cleared up in the afternoon.\n<\/code><\/pre>\n<p>Input audio (4.7 s), the sentence we read out:<\/p>\n<blockquote>\n<p>The next meeting will be held from 2 p.m. on March 15 in meeting room B.<\/p>\n<\/blockquote>\n<p><audio controls preload=\"none\" src=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/wp-content\/uploads\/2026\/10\/lmw-lab-asr-numbers-v1.wav\"><\/audio><\/p>\n<p>The model&#8217;s transcript, verbatim (word error rate 12.5%: 2 errors in 16 words; transcribed in 4.63 s):<\/p>\n<pre><code class=\"language-text\">The next meeting will be held from two PM on march fifteenth in meeting room B.\n<\/code><\/pre>\n<p>Input audio (3.61 s), the sentence we read out:<\/p>\n<blockquote>\n<p>Oh, it&#8217;s already finished? That was much quicker than I expected.<\/p>\n<\/blockquote>\n<p><audio controls preload=\"none\" src=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/wp-content\/uploads\/2026\/10\/lmw-lab-asr-surprise-v1.wav\"><\/audio><\/p>\n<p>The model&#8217;s transcript, verbatim (word error rate 0.0%: 0 errors in 11 words; transcribed in 4.17 s):<\/p>\n<pre><code class=\"language-text\">Oh, it's already finished. That was much quicker than I expected.\n<\/code><\/pre>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Result<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Word error rate over all three clips<\/td>\n<td>4.8% (2 errors in 42 words)<\/td>\n<\/tr>\n<tr>\n<td>Speed on this CPU<\/td>\n<td>1.0\u00d7 real time (13.0 s of audio in 13.4 s)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Conditions: CPU Neoverse-N1 with 3 threads (no GPU), fp32, transformers 5.13.0, PyTorch 2.14.1+cpu, weights at revision <code>429f35b9c8<\/code> (base config at <code>541d1f99c6<\/code>), greedy decoding, measured 2026-10-01 (JST). The first decode was run once beforehand and not timed.<\/p>\n<p>The word error rate is counted by our code after lowercasing and removing punctuation; numerals are not converted, so &#8220;2 p.m.&#8221; and &#8220;two p.m.&#8221; count as different words. Three short clips of one synthetic voice say nothing about accuracy on real recordings, accents or noise. The speed is for these short clips on this CPU and says nothing about a GPU. The weights are stored in a compressed format, and we expanded them to fp32 to run them, so the memory use here is not the size of the distributed file.<\/p>\n<h4>Setup, Download and Run Steps (What We Actually Ran)<\/h4>\n<p>We ran this on our own server. We downloaded the weights container and the base model&#8217;s configuration files from Hugging Face at pinned revisions and checked the SHA-256 of the container before using it. The script below (<code>asr_cpu_run.py<\/code>, written by us; full text) takes out only the one weights file from the container, checks its SHA-256, reads the bytes with numpy (no pickle), loads them into the standard <code>ParakeetForTDT<\/code> with strict key matching, and transcribes our audio. The weights are deleted afterwards. These are the packages we installed:<\/p>\n<pre><code class=\"language-bash\"># Make an environment (no GPU needed)\npython3.12 -m venv asr-venv\nasr-venv\/bin\/python -m pip install 'torch' 'transformers==5.13.0' 'numpy' 'zstandard' 'librosa'\n<\/code><\/pre>\n<pre><code class=\"language-python\">&quot;&quot;&quot;\n\u72ec\u81ea\u5f62\u5f0f\u306e\u91cd\u307f(Phonon-2 \u306e fermion-five-value-parakeet-v1)\u3092\u3001**\u5f53\u30b5\u30a4\u30c8\u304c\u66f8\u3044\u305f\u8aad\u307f\u53d6\u308a\u51e6\u7406**\u3067\u6a19\u6e96\u306e transformers \u306e\nParakeetForTDT \u306b\u8aad\u307e\u305b\u3066\u3001CPU \u3067\u6587\u5b57\u8d77\u3053\u3057\u3059\u308b(lmw\/lab\/asr_cpu.py \u304b\u3089\u3001\u5c02\u7528\u306e venv \u306e Python \u3067\u547c\u3070\u308c\u308b)\u3002\n\u914d\u5e03\u5143\u306e Python \u30d5\u30a1\u30a4\u30eb(fermion_container.py\u30fbreference_transformers.py)\u306f\u8aad\u307f\u8fbc\u307e\u306a\u3044\u30fb\u5b9f\u884c\u3057\u306a\u3044\u3002\n\u91cd\u307f\u306f pickle \u3092\u4f7f\u308f\u305a\u3001\u30d0\u30a4\u30c8\u5217\u3092 numpy \u3067\u8aad\u3080\u3002\u8a18\u4e8b\u306e\u624b\u9806\u306b\u3053\u306e\u30d5\u30a1\u30a4\u30eb\u306e\u5168\u6587\u304c\u8f09\u308b\u3002\n\n    python asr_cpu_run.py &lt;job.json&gt; &lt;out.json&gt;\n&quot;&quot;&quot;\nimport hashlib\nimport json\nimport os\nimport platform\nimport sys\nimport tarfile\nimport time\nimport wave\nfrom datetime import datetime, timezone\n\nimport numpy as np\n\nMAX_INNER_BYTES = 1024 ** 3\n\ndef sha256_of(path: str) -&gt; str:\n    h = hashlib.sha256()\n    with open(path, &quot;rb&quot;) as f:\n        for chunk in iter(lambda: f.read(1 &lt;&lt; 20), b&quot;&quot;):\n            h.update(chunk)\n    return h.hexdigest()\n\ndef extract_inner(archive: str, inner: str, dest: str) -&gt; str:\n    &quot;&quot;&quot;tar.zst \u304b\u3089\u3001\u540d\u524d\u304c inner \u306e\u901a\u5e38\u306e\u30d5\u30a1\u30a4\u30eb1\u3064\u3060\u3051\u3092\u53d6\u308a\u51fa\u3059(\u30ea\u30f3\u30af\u30fb\u307b\u304b\u306e\u30d5\u30a1\u30a4\u30eb\u306f\u8aad\u307e\u306a\u3044)\u3002&quot;&quot;&quot;\n    import zstandard\n    with open(archive, &quot;rb&quot;) as raw, zstandard.ZstdDecompressor().stream_reader(raw) as stream:\n        with tarfile.open(fileobj=stream, mode=&quot;r|&quot;) as tar:\n            for member in tar:\n                if member.isfile() and os.path.basename(member.name) == inner:\n                    if member.size &gt; MAX_INNER_BYTES:\n                        raise ValueError(&quot;\u5c55\u958b\u5f8c\u306e\u30d5\u30a1\u30a4\u30eb\u304c\u5927\u304d\u3059\u304e\u307e\u3059&quot;)\n                    path = os.path.join(dest, inner)\n                    with open(path, &quot;wb&quot;) as f:\n                        f.write(tar.extractfile(member).read())\n                    return path\n    raise ValueError(f&quot;\u30a2\u30fc\u30ab\u30a4\u30d6\u306b {inner} \u304c\u3042\u308a\u307e\u305b\u3093&quot;)\n\ndef _trits(buf: bytes, rows: int, cols: int) -&gt; np.ndarray:\n    &quot;&quot;&quot;1\u30d0\u30a4\u30c8\u306b3\u5024(0\u301c2)\u30925\u3064\u8a70\u3081\u305f\u7b26\u53f7\u3092\u3001(rows, cols) \u306e 0\/1\/2 \u306b\u623b\u3059\u3002&quot;&quot;&quot;\n    row_bytes = (cols + 4) \/\/ 5\n    x = np.frombuffer(buf, dtype=np.uint8, count=rows * row_bytes).reshape(rows, row_bytes).astype(np.uint16)\n    out = np.empty((rows, row_bytes, 5), dtype=np.uint8)\n    for k in range(5):\n        out[:, :, k] = (x \/\/ (3 ** k)) % 3\n    return out.reshape(rows, row_bytes * 5)[:, :cols]\n\ndef five_value(blob: bytes, shape: tuple) -&gt; np.ndarray:\n    &quot;&quot;&quot;\u7b26\u53f7(-1\/0\/+1)\u3068\u3001\u884c\u3054\u3068\u306e2\u3064\u306e\u5927\u304d\u3055(lo\/hi)\u304b\u3089\u91cd\u307f\u3092\u4f5c\u308b\u3002\u7b26\u53f7\u304c0\u3067\u306a\u3044\u8981\u7d20\u306f1\u30d3\u30c3\u30c8\u3067 lo \u304b hi \u3092\u9078\u3076\u3002&quot;&quot;&quot;\n    rows, cols = shape\n    row_bytes = (cols + 4) \/\/ 5\n    codes = _trits(blob[: rows * row_bytes], rows, cols)\n    nonzero = codes != 1\n    count = int(nonzero.sum())\n    bit_bytes = (count + 7) \/\/ 8\n    off = rows * row_bytes\n    if off + bit_bytes + 4 * rows != len(blob):\n        raise ValueError(&quot;five_value \u306e\u5927\u304d\u3055\u304c\u5408\u3044\u307e\u305b\u3093&quot;)\n    bits = np.unpackbits(np.frombuffer(blob[off: off + bit_bytes], dtype=np.uint8), bitorder=&quot;little&quot;)[:count].astype(bool)\n    off += bit_bytes\n    lo = np.frombuffer(blob[off: off + 2 * rows], dtype=np.float16)\n    hi = np.frombuffer(blob[off + 2 * rows: off + 4 * rows], dtype=np.float16)\n    is_hi = np.zeros((rows, cols), dtype=bool)\n    is_hi[nonzero] = bits\n    magnitude = np.where(is_hi, hi[:, None], lo[:, None])\n    sign = codes.astype(np.int8) - 1\n    return (sign.astype(np.float16) * magnitude).astype(np.float16)\n\ndef int_n(blob: bytes, shape: tuple, bits: int) -&gt; np.ndarray:\n    &quot;&quot;&quot;\u884c\u3054\u3068\u306e\u5c3a\u5ea6(fp16)\u3092\u639b\u3051\u308b 6bit \/ 8bit \u306e\u6574\u6570\u306e\u91cd\u307f\u3002&quot;&quot;&quot;\n    rows = shape[0]\n    total = int(np.prod(shape))\n    cols = total \/\/ rows\n    body, scales = blob[: -2 * rows], np.frombuffer(blob[-2 * rows:], dtype=np.float16)\n    if bits == 8:\n        q = np.frombuffer(body, dtype=np.int8).astype(np.int32)[:total]\n    elif bits == 6:\n        b = np.frombuffer(body, dtype=np.uint8).reshape(-1, 3).astype(np.uint32)\n        packed = b[:, 0] | (b[:, 1] &lt;&lt; 8) | (b[:, 2] &lt;&lt; 16)\n        q = np.stack([(packed &gt;&gt; s) &amp; 0x3F for s in (0, 6, 12, 18)], axis=1).ravel()[:total].astype(np.int32) - 32\n    else:\n        raise ValueError(f&quot;\u5bfe\u5fdc\u3057\u3066\u3044\u306a\u3044\u30d3\u30c3\u30c8\u6570\u3067\u3059: {bits}&quot;)\n    if q.size != total:\n        raise ValueError(&quot;int \u306e\u5927\u304d\u3055\u304c\u5408\u3044\u307e\u305b\u3093&quot;)\n    return (q.reshape(rows, cols).astype(np.float32) * scales.astype(np.float32)[:, None]).reshape(shape)\n\ndef read_container(path: str, fmt: str) -&gt; dict:\n    &quot;&quot;&quot;\u5148\u982d8\u30d0\u30a4\u30c8\u304c\u898b\u51fa\u3057\u306e\u9577\u3055\u3001\u7d9a\u304f JSON \u306e index \u306b\u5f93\u3063\u3066\u3001\u5404\u30c6\u30f3\u30bd\u30eb\u306e\u5024\u3092\u4e26\u3079\u305f\u30d5\u30a1\u30a4\u30eb\u3092\u8aad\u3080\u3002&quot;&quot;&quot;\n    tensors = {}\n    with open(path, &quot;rb&quot;) as f:\n        header = json.loads(f.read(int.from_bytes(f.read(8), &quot;little&quot;)))\n        if header.get(&quot;format&quot;) != fmt:\n            raise ValueError(f&quot;\u5f62\u5f0f\u304c\u9055\u3044\u307e\u3059: {header.get('format')}&quot;)\n        for e in header[&quot;index&quot;]:\n            blob = f.read(e[&quot;b&quot;])\n            if len(blob) != e[&quot;b&quot;]:\n                raise ValueError(&quot;\u30d5\u30a1\u30a4\u30eb\u304c\u9014\u4e2d\u3067\u7d42\u308f\u3063\u3066\u3044\u307e\u3059&quot;)\n            kind, shape = e[&quot;k&quot;], tuple(e[&quot;shape&quot;])\n            if kind == &quot;five_value&quot;:\n                tensors[e[&quot;n&quot;] + &quot;.weight&quot;] = five_value(blob, shape)\n            elif kind in (&quot;int6&quot;, &quot;int8&quot;):\n                tensors[e[&quot;n&quot;]] = int_n(blob, shape, int(kind[3:]))\n            elif kind == &quot;fp16&quot;:\n                tensors[e[&quot;n&quot;]] = np.frombuffer(blob, dtype=np.float16).reshape(shape).copy()\n            else:\n                raise ValueError(f&quot;\u5bfe\u5fdc\u3057\u3066\u3044\u306a\u3044\u7a2e\u985e\u3067\u3059: {kind}&quot;)\n        if f.read(1):\n            raise ValueError(&quot;\u672b\u5c3e\u306b\u4f59\u308a\u306e\u30d0\u30a4\u30c8\u304c\u3042\u308a\u307e\u3059&quot;)\n    return tensors\n\ndef state_dict(tensors: dict):\n    import re\n\n    import torch\n    out = {}\n    for name, value in tensors.items():\n        if name.endswith(&quot;num_batches_tracked&quot;):\n            count = float(np.asarray(value, dtype=np.float32).reshape(-1)[0])\n            out[name] = torch.tensor(int(count) if np.isfinite(count) else 0, dtype=torch.int64)\n            continue\n        arr = np.ascontiguousarray(value.astype(np.float32))\n        if arr.ndim == 2 and re.search(r&quot;\\.conv\\.pointwise_conv[12]\\.weight$&quot;, name):\n            arr = arr[:, :, None]\n        out[name] = torch.from_numpy(arr)\n    return out\n\ndef waveform(path: str) -&gt; np.ndarray:\n    with wave.open(path, &quot;rb&quot;) as r:\n        if (r.getframerate(), r.getnchannels(), r.getsampwidth()) != (16000, 1, 2):\n            raise ValueError(&quot;16kHz\u30fb\u30e2\u30ce\u30e9\u30eb\u30fb16bit \u306e WAV \u3067\u306f\u3042\u308a\u307e\u305b\u3093&quot;)\n        pcm = r.readframes(r.getnframes())\n    return np.frombuffer(pcm, dtype=np.int16).astype(np.float32) \/ 32768.0\n\ndef installed_packages() -&gt; list:\n    from importlib import metadata\n    seen = {}\n    for dist in metadata.distributions():\n        name = dist.metadata[&quot;Name&quot;]\n        if name and name.lower() not in seen:\n            seen[name.lower()] = f&quot;{name}=={dist.version}&quot;\n    return sorted(seen.values(), key=str.lower)\n\ndef main(job_path: str, out_path: str) -&gt; None:\n    import torch\n    import transformers\n    from transformers import AutoProcessor, GenerationConfig, ParakeetForTDT, ParakeetTDTConfig\n    with open(job_path, encoding=&quot;utf-8&quot;) as f:\n        job = json.load(f)\n    os.environ[&quot;HF_HUB_OFFLINE&quot;] = &quot;1&quot;\n    os.environ[&quot;TRANSFORMERS_OFFLINE&quot;] = &quot;1&quot;\n    torch.set_num_threads(job[&quot;threads&quot;])\n    t0 = time.time()\n    record = {&quot;transformers&quot;: transformers.__version__, &quot;torch&quot;: torch.__version__, &quot;dtype&quot;: &quot;float32&quot;,\n              &quot;threads&quot;: job[&quot;threads&quot;], &quot;python&quot;: platform.python_version(), &quot;packages&quot;: installed_packages(),\n              &quot;measured_at&quot;: datetime.now(timezone.utc).isoformat()}\n    inner = extract_inner(job[&quot;archive&quot;], job[&quot;inner&quot;], job[&quot;work&quot;])\n    if sha256_of(inner) != job[&quot;inner_sha256&quot;]:\n        raise ValueError(&quot;\u5c55\u958b\u3057\u305f\u30d5\u30a1\u30a4\u30eb\u306e SHA-256 \u304c\u56fa\u5b9a\u3057\u305f\u5024\u3068\u9055\u3044\u307e\u3059&quot;)\n    tensors = read_container(inner, job[&quot;format&quot;])\n    model = ParakeetForTDT(ParakeetTDTConfig.from_pretrained(job[&quot;base_dir&quot;]))\n    missing, unexpected = model.load_state_dict(state_dict(tensors), strict=False)\n    if missing or unexpected:\n        raise ValueError(f&quot;\u30c6\u30f3\u30bd\u30eb\u306e\u540d\u524d\u304c\u5408\u3044\u307e\u305b\u3093: \u4e0d\u8db3 {list(missing)[:5]}\u30fb\u4f59\u308a {list(unexpected)[:5]}&quot;)\n    record[&quot;tensors&quot;] = len(tensors)\n    record[&quot;parameters&quot;] = sum(p.numel() for p in model.parameters())\n    model.eval()\n    model.generation_config = GenerationConfig.from_pretrained(job[&quot;base_dir&quot;])\n    processor = AutoProcessor.from_pretrained(job[&quot;base_dir&quot;])\n    record[&quot;load_sec&quot;] = round(time.time() - t0, 1)\n\n    def decode(item: dict):\n        wav = waveform(item[&quot;path&quot;])\n        started = time.time()\n        inputs = processor([wav], sampling_rate=16000, return_tensors=&quot;pt&quot;, padding=True)\n        with torch.no_grad():\n            out = model.generate(input_features=inputs[&quot;input_features&quot;], attention_mask=inputs.get(&quot;attention_mask&quot;))\n        text = processor.batch_decode(getattr(out, &quot;sequences&quot;, out), skip_special_tokens=True)[0].strip()\n        return text, time.time() - started, wav.size \/ 16000\n\n    decode(job[&quot;inputs&quot;][0])  # \u6700\u521d\u306e1\u56de\u306f\u521d\u671f\u5316\u306e\u6642\u9593\u304c\u5165\u308b\u306e\u3067\u3001\u6e2c\u3089\u305a\u306b\u6d41\u3059\n    outputs = []\n    for item in job[&quot;inputs&quot;]:\n        text, seconds, audio_seconds = decode(item)\n        outputs.append({&quot;id&quot;: item[&quot;id&quot;], &quot;language&quot;: item[&quot;language&quot;], &quot;hypothesis&quot;: text,\n                        &quot;decode_sec&quot;: round(seconds, 2), &quot;audio_sec&quot;: round(audio_seconds, 2)})\n        print(f&quot;[asr-cpu] {item['id']}: {audio_seconds:.1f}\u79d2\u306e\u97f3\u58f0\u3092 {seconds:.1f}\u79d2\u3067\u6587\u5b57\u8d77\u3053\u3057&quot;, flush=True)\n    record[&quot;outputs&quot;] = outputs\n    with open(out_path, &quot;w&quot;, encoding=&quot;utf-8&quot;) as f:\n        json.dump(record, f, ensure_ascii=False)\n\nif __name__ == &quot;__main__&quot;:\n    main(sys.argv[1], sys.argv[2])\n<\/code><\/pre>\n<p>We ran it as <code>asr-venv\/bin\/python asr_cpu_run.py job.json out.json<\/code>. The <code>job.json<\/code> we passed (as used; the audio is shown by its SHA-256):<\/p>\n<pre><code class=\"language-json\">{\n  &quot;archive&quot;: &quot;\/tmp\/lmw-asr-cpu-lcq29pwq\/phonon-2.bps.tar.zst&quot;,\n  &quot;inner&quot;: &quot;model.fermion&quot;,\n  &quot;inner_sha256&quot;: &quot;4b6bfa3a12cc3c4e0a54f2ab3ec4ca7a842b09e5c7ecfc8e7ca0ac6cc8c11468&quot;,\n  &quot;format&quot;: &quot;fermion-five-value-parakeet-v1&quot;,\n  &quot;work&quot;: &quot;\/tmp\/lmw-asr-cpu-lcq29pwq&quot;,\n  &quot;base_dir&quot;: &quot;\/tmp\/lmw-asr-cpu-lcq29pwq\/base&quot;,\n  &quot;threads&quot;: 3,\n  &quot;inputs&quot;: [\n    {\n      &quot;id&quot;: &quot;weather&quot;,\n      &quot;language&quot;: &quot;en&quot;,\n      &quot;wav_sha256&quot;: &quot;1258dcd9f83234db2601a928bd8facddc1c0c2751d7a90765375cc9197dc0651&quot;\n    },\n    {\n      &quot;id&quot;: &quot;numbers&quot;,\n      &quot;language&quot;: &quot;en&quot;,\n      &quot;wav_sha256&quot;: &quot;ce0b8d12768dd1d9f5fdbf5fd0a0a98c35b5f5b7c522b814c5da01579812f430&quot;\n    },\n    {\n      &quot;id&quot;: &quot;surprise&quot;,\n      &quot;language&quot;: &quot;en&quot;,\n      &quot;wav_sha256&quot;: &quot;fcd7c8d3f8cce918a56d959071a298d3a994d29c78586a2868b2a0a0c7c56e40&quot;\n    }\n  ]\n}\n<\/code><\/pre>\n<details>\n<summary>Files we downloaded (SHA-256)<\/summary>\n<pre><code>98125795b6dda72f5c6eee9ba33d19815df65dcb18b50a357bf9f73c9935309e  phonon-2.bps.tar.zst None  None 4b6bfa3a12cc3c4e0a54f2ab3ec4ca7a842b09e5c7ecfc8e7ca0ac6cc8c11468  model.fermion<\/code><\/pre>\n<\/details>\n<details>\n<summary>Package versions in the environment (full list)<\/summary>\n<pre><code>annotated-doc==0.0.5 annotated-types==0.8.0 anyio==4.15.1 certifi==2026.7.22 cffi==2.1.1 charset-normalizer==3.5.2 click==8.5.0 cloudpickle==3.1.2 decorator==5.3.1 filelock==3.32.3 fsspec==2026.7.0 h11==0.16.0 hf-xet==1.6.0 httpcore2==2.13.1 httpcore==1.0.9 httpx2==2.13.1 httpx==0.28.1 huggingface_hub==1.33.0 idna==3.20 Jinja2==3.1.6 joblib==1.6.0 lazy-loader==0.6 librosa==1.0.0 llvmlite==0.50.0 markdown-it-py==4.2.0 MarkupSafe==3.0.3 mdurl==0.1.2 mpmath==1.3.0 msgpack==1.2.3 narwhals==2.26.0 networkx==3.6.1 numba==0.68.0 numpy==2.5.3 packaging==26.3 pip==24.0 platformdirs==4.12.2 pooch==1.9.0 pycparser==3.0 pydantic==2.13.5 pydantic_core==2.46.5 Pygments==2.21.0 PyYAML==6.0.3 regex==2026.9.29 requests==2.34.2 rich==15.0.0 safetensors==0.8.0 scikit-learn==1.9.1 scipy==1.18.1 setuptools==78.1.0 shellingham==1.5.4 sniffio==1.3.1 soundfile==0.14.0 soxr==1.1.0 sympy==1.14.0 threadpoolctl==3.7.0 tokenizers==0.22.2 torch==2.14.1+cpu tqdm==4.70.1 transformers==5.13.0 truststore==0.10.4 typer==0.27.2 typing-inspection==0.4.4 typing_extensions==4.16.0 urllib3==2.8.0 zstandard==0.25.0<\/code><\/pre>\n<\/details>\n<p><!-- \/lmw:lab --><\/p>\n<h2>How to Get It<\/h2>\n<p>Phonon-2 is available as open-weight on Hugging Face and is not a gated model, so it can be downloaded and used without prior application for use or license agreement procedures. As a distribution format, MLX-compatible weights optimized for Apple Silicon are provided.<\/p>\n<p>In Mac environments equipped with Apple Silicon, transcription can be executed directly from the command line by installing the officially provided Python package and MLX-related libraries. Run the following commands in the terminal to set up the necessary dependencies:<\/p>\n<pre><code class=\"language-bash\">pip install fermion-research\npip install mlx mlx-audio mlx-lm soundfile scipy zstandard\n<\/code><\/pre>\n<p>Once the installation is complete, transcribe an audio file (such as.wav format) using one of the following commands:<\/p>\n<pre><code class=\"language-bash\">phonon transcribe recording.wav\nfermion transcribe phonon-2 recording.wav\n<\/code><\/pre>\n<p>For Linux (x86-64 and Arm architectures) and Windows CPU environments, a dedicated execution engine is published in the GitHub repository (github.com\/fermionresearch\/phonon). In addition, for developers who want to simplify environment construction or use GPU acceleration, official Docker images are also provided.<\/p>\n<p>When running the Docker container in a CPU environment, use the following command:<\/p>\n<pre><code class=\"language-bash\">docker run --rm -v &quot;$PWD&quot;:\/audio -v phonon-cache:\/home\/phonon\/.cache ghcr.io\/fermionresearch\/phonon-cpu:2.0.3 transcribe phonon-2 \/audio\/recording.wav\n<\/code><\/pre>\n<p>When running in a CUDA environment equipped with an NVIDIA GPU, start the container by specifying the GPU option as follows:<\/p>\n<pre><code class=\"language-bash\">docker run --rm --gpus all -v &quot;$PWD&quot;:\/audio -v phonon-cache:\/home\/phonon\/.cache ghcr.io\/fermionresearch\/phonon-cuda:1.0.4 transcribe phonon-2 \/audio\/recording.wav\n<\/code><\/pre>\n<p>Regarding the license aspect, the model weights themselves are released under the same &#8220;CC-BY-4.0&#8221; as the base model, and changes are described in the NOTICE file within the repository. In addition, the provided command-line tools and the code in the repository are applied with the &#8220;Apache 2.0&#8221; license.<\/p>\n<p><!-- lmw:same-task --><\/p>\n<h2>Other Models for the Same Task<\/h2>\n<p><em>Recent audio models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site&#8217;s estimate.<\/em><\/p>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/edge0-audio8-asr-infinite-2\/\">Audio8-ASR-Infinite Speech Recognition Model: 12GB+ VRAM, File List<\/a> (4.1B, 12GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/yue2-3b-music-generation-model\/\">YuE2-3B Music Generation Model for Lyrics and Style: 12GB+ VRAM<\/a> (3.6B, 12GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/07\/irodori-tts-v41-anime-released\/\">Irodori-TTS-v4.1-Anime: Our Generated Audio, File List<\/a> (766M, 4GB)<\/li>\n<\/ul>\n<p><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-audio\">See all audio models \u2192<\/a><\/p>\n<p><!-- \/lmw:same-task --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model runs from the 4GB tier) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-8gb-en\/\">Other models that run on a 8GB GPU<\/a><\/li>\n<li><strong>What Q8_0 mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-mlx-en\/\">MLX format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-nvidia-en\/\">NVIDIA: models, licenses and articles<\/a><\/li>\n<li><strong>How to read WER<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-benchmarks-en\/\">Benchmark glossary<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/FermionResearch\/Phonon-2\">FermionResearch\/Phonon-2 (Hugging Face)<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/nvidia\/parakeet-tdt-0.6b-v3\">nvidia\/parakeet-tdt-0.6b-v3 (Hugging Face)<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-10-01: Added our own measurements: transcripts of English audio we made.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Discover Phonon-2 by FermionResearch, an open-weight English ASR model offering high accuracy and low memory usage for on-device speech-to-text.<\/p>\n","protected":false},"author":1,"featured_media":8448,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[2773,977,2775,2777,2779,1547,1017],"class_list":["post-8449","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-fermionresearch-en","tag-mlx-en","tag-parakeet-tdt-v3-en","tag-phonon-2-en","tag-speech-to-text-en","tag-verified","tag--en"],"lang":"en","translations":{"en":8449,"ja":8447},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8449","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=8449"}],"version-history":[{"count":2,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8449\/revisions"}],"predecessor-version":[{"id":8754,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8449\/revisions\/8754"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/8448"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=8449"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=8449"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=8449"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}