Phonon-2 Speech Recognition Model: 4GB+ VRAM

At a Glance
| Item | Value |
|---|---|
| Repository | FermionResearch/Phonon-2 |
| Publisher guide | NVIDIA: models and licenses |
| Published | 2026-09-28 |
| License | cc-by-4.0 |
| Formats | MLX |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
FermionResearch has released “Phonon-2", an open-weight model that achieves both extremely high accuracy and lightweight performance in English Automatic Speech Recognition (ASR). This model is a Speech-to-Text model that takes audio files as input and outputs text properly containing punctuation, capitalization, and numerical notation.
Phonon-2 adopts NVIDIA’s “nvidia/parakeet-tdt-0.6b-v3" as its base model and applies proprietary quantization techniques, making it reported to be the most accurate open English speech recognition model under 900 MB. A major feature is that compared to the full-precision teacher model of about 2.5 GB, it maintains equivalent accuracy while reducing the download size by 15x.
Specifications
The main specifications of Phonon-2 and its base model based on the documentation are as follows:
- Parameter Count: 0.6B (600 million parameters)
- Architecture: parakeet_tdt_five_value (structure based on NVIDIA’s FastConformer-TDT)
- Quantization Method: Quantization-Aware Training (QAT) keeps each encoder weight in 5 trained levels (equivalent to about 2.1 bits)
- Input Specifications: 16kHz mono audio (supports.wav and.flac formats)
- Output Specifications: Text including punctuation, capitalization, and numerical notation. Timestamps (word-level and segment-level) output is also possible
- Inference Speed (Reported Values):
- M5 MacBook Air: Processes 1 hour of audio in about 20 seconds (174x real-time ratio) – Zen 5 cores (16 vCPUs): 143x real-time ratio – H100 (batch size 128): 6,680x real-time ratio
- License: CC-BY-4.0 (Available for both commercial and non-commercial use. However, you need to check the list of changes described in the NOTICE file)
Performance
The publisher FermionResearch has published benchmark results using the Open ASR Leaderboard code. The following table compares the WER (Word Error Rate) of Phonon-2 with major models including its teacher model.
→ Scroll horizontally to see all columns
| Model | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Average |
|---|---|---|---|---|---|---|---|---|---|
| Phonon-2 | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 |
| Parakeet TDT 0.6B v3, teacher† | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 |
| Parakeet Redux | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 |
| Phonon-1 | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 |
| Canary 180M Flash† | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 |
| Voxtral Mini 4B Realtime† | ≈8,000 MB* | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 |
| Whisper large-v3-turbo† | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 |
| Nemotron 3.5 ASR Streaming 0.6B† | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 |
In this benchmark, Phonon-2 recorded an average WER of 5.21%, demonstrating accuracy that surpasses models with file sizes several to dozens of times larger, such as Whisper large-v3-turbo and Voxtral Mini 4B. Particularly on the VoxPopuli set, it achieved the best score of 2.46%, outperforming the teacher model Parakeet TDT 0.6B v3 (3.19%).
On the other hand, on some datasets such as LibriSpeech (LS clean/other) and Earnings-22, the results do not reach the full-precision teacher model. However, considering that the model size is extremely lightweight at 164 MB, it can be said to be a model with very high cost-effectiveness for on-device operation.
As a characteristic of the base model nvidia/parakeet-tdt-0.6b-v3, high noise resilience is also listed. Evaluations of the original model show that it can maintain a certain level of recognition accuracy even under severe noise environments such as SNR 10 to SNR 0. Phonon-2 inherits this excellent architecture while being a model reduced in size to the utmost limit.
Strengths and Use Cases
Phonon-2 is specialized for executing high-quality English transcription at high speed in resource-constrained on-device and edge device environments. It has a track record of being adopted as the internal engine for “Detta", a dictation app for Mac, and is optimized for voice input and real-time transcription in personal desktop environments.
This model exhibits particularly high suitability in the transcription of meetings and dialogues. As shown in the benchmark results, it records recognition accuracy comparable to or exceeding the full-precision teacher model on the AMI dataset dealing with meeting audio and the VoxPopuli dataset dealing with parliamentary speeches. Therefore, it sufficiently supports practical use cases such as creating minutes for business meetings, text-making of lectures and speeches, and interview transcription.
Inheriting the specifications possessed by the base model “parakeet-tdt-0.6b-v3", the output text includes appropriate punctuation (periods and commas), capitalization, and numerical expressions according to the context from the beginning. Since it can be used as natural text as is without separately performing post-processing (such as punctuation restoration and format shaping) required in general speech recognition models, it has a configuration that is easy to handle as an input front-end for subtitle generation pipelines, conversational AI, and voice assistants.
In addition, since it supports word-level and segment-level timestamp output, it is also suitable for applications such as automatic captioning for video content, creation of voice search indexes, and support for editing long-form audio. Because it possesses high inference throughput, it can be widely utilized from daily dictation on local laptops to high-speed batch processing of large amounts of audio files in server environments.
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 627M parameters (taken from the base model nvidia/parakeet-tdt-0.6b-v3)
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 4GB (laptop iGPU / phone class) | Q8_0 | 0.7GB | 0.8GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-10-01): llama.cpp: not registered, vLLM: not registered, MLX (mlx-lm): not registered. “Not registered" means the name is absent from that registry today, not that the model cannot run.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Can You Run It Locally?
The publisher distributes this model as MLX.
License — cc-by-4.0 (Commercial use allowed): Permits commercial use, but attribution is mandatory — omitting credit is a violation.
Compression: the Q8_0 build measures 9.59 bits per weight — about 60% the size of the original 16-bit weights, calculated by this site from the actual file sizes.
Compiled by this site’s code from the published formats and the license field. License summaries are not legal advice — check the publisher’s original terms before relying on them.
Our Own Measurements
Values we measured ourselves on our server (no GPU) by actually reading and running this model’s files — not figures copied from the model card. How we measure, and the results for every model: Our Measurements.
Transcribing English Audio We Made (on Our CPU)
FermionResearch/Phonon-2 ships its weights in its own container format. We did not run the publisher’s Python code. We wrote a reader for the format ourselves, loaded the weights into the standard transformers ParakeetForTDT (the architecture of nvidia/parakeet-tdt-0.6b-v3 that this model is built on), and ran it in fp32 on our server’s CPU. The audio is not a recording from elsewhere: we read three fixed English sentences with a text-to-speech voice (en_US-ljspeech-medium, MIT(音声モデルの配布)・学習データ LJ Speech はパブリックドメイン) and let the model transcribe each clip once. Nothing was selected or edited; these are the first outputs, misrecognitions included. This model supports English only, so we tested English only and make no claim about other languages.
Input audio (4.7 s), the sentence we read out:
It had been raining since the morning, but it gradually cleared up in the afternoon.
The model’s transcript, verbatim (word error rate 0.0%: 0 errors in 15 words; transcribed in 4.6 s):
It had been raining since the morning, but it gradually cleared up in the afternoon.
Input audio (4.7 s), the sentence we read out:
The next meeting will be held from 2 p.m. on March 15 in meeting room B.
The model’s transcript, verbatim (word error rate 12.5%: 2 errors in 16 words; transcribed in 4.63 s):
The next meeting will be held from two PM on march fifteenth in meeting room B.
Input audio (3.61 s), the sentence we read out:
Oh, it’s already finished? That was much quicker than I expected.
The model’s transcript, verbatim (word error rate 0.0%: 0 errors in 11 words; transcribed in 4.17 s):
Oh, it's already finished. That was much quicker than I expected.
| Item | Result |
|---|---|
| Word error rate over all three clips | 4.8% (2 errors in 42 words) |
| Speed on this CPU | 1.0× real time (13.0 s of audio in 13.4 s) |
Conditions: CPU Neoverse-N1 with 3 threads (no GPU), fp32, transformers 5.13.0, PyTorch 2.14.1+cpu, weights at revision 429f35b9c8 (base config at 541d1f99c6), greedy decoding, measured 2026-10-01 (JST). The first decode was run once beforehand and not timed.
The word error rate is counted by our code after lowercasing and removing punctuation; numerals are not converted, so “2 p.m." and “two p.m." count as different words. Three short clips of one synthetic voice say nothing about accuracy on real recordings, accents or noise. The speed is for these short clips on this CPU and says nothing about a GPU. The weights are stored in a compressed format, and we expanded them to fp32 to run them, so the memory use here is not the size of the distributed file.
Setup, Download and Run Steps (What We Actually Ran)
We ran this on our own server. We downloaded the weights container and the base model’s configuration files from Hugging Face at pinned revisions and checked the SHA-256 of the container before using it. The script below (asr_cpu_run.py, written by us; full text) takes out only the one weights file from the container, checks its SHA-256, reads the bytes with numpy (no pickle), loads them into the standard ParakeetForTDT with strict key matching, and transcribes our audio. The weights are deleted afterwards. These are the packages we installed:
# Make an environment (no GPU needed)
python3.12 -m venv asr-venv
asr-venv/bin/python -m pip install 'torch' 'transformers==5.13.0' 'numpy' 'zstandard' 'librosa'
"""
独自形式の重み(Phonon-2 の fermion-five-value-parakeet-v1)を、**当サイトが書いた読み取り処理**で標準の transformers の
ParakeetForTDT に読ませて、CPU で文字起こしする(lmw/lab/asr_cpu.py から、専用の venv の Python で呼ばれる)。
配布元の Python ファイル(fermion_container.py・reference_transformers.py)は読み込まない・実行しない。
重みは pickle を使わず、バイト列を numpy で読む。記事の手順にこのファイルの全文が載る。
python asr_cpu_run.py <job.json> <out.json>
"""
import hashlib
import json
import os
import platform
import sys
import tarfile
import time
import wave
from datetime import datetime, timezone
import numpy as np
MAX_INNER_BYTES = 1024 ** 3
def sha256_of(path: str) -> str:
h = hashlib.sha256()
with open(path, "rb") as f:
for chunk in iter(lambda: f.read(1 << 20), b""):
h.update(chunk)
return h.hexdigest()
def extract_inner(archive: str, inner: str, dest: str) -> str:
"""tar.zst から、名前が inner の通常のファイル1つだけを取り出す(リンク・ほかのファイルは読まない)。"""
import zstandard
with open(archive, "rb") as raw, zstandard.ZstdDecompressor().stream_reader(raw) as stream:
with tarfile.open(fileobj=stream, mode="r|") as tar:
for member in tar:
if member.isfile() and os.path.basename(member.name) == inner:
if member.size > MAX_INNER_BYTES:
raise ValueError("展開後のファイルが大きすぎます")
path = os.path.join(dest, inner)
with open(path, "wb") as f:
f.write(tar.extractfile(member).read())
return path
raise ValueError(f"アーカイブに {inner} がありません")
def _trits(buf: bytes, rows: int, cols: int) -> np.ndarray:
"""1バイトに3値(0〜2)を5つ詰めた符号を、(rows, cols) の 0/1/2 に戻す。"""
row_bytes = (cols + 4) // 5
x = np.frombuffer(buf, dtype=np.uint8, count=rows * row_bytes).reshape(rows, row_bytes).astype(np.uint16)
out = np.empty((rows, row_bytes, 5), dtype=np.uint8)
for k in range(5):
out[:, :, k] = (x // (3 ** k)) % 3
return out.reshape(rows, row_bytes * 5)[:, :cols]
def five_value(blob: bytes, shape: tuple) -> np.ndarray:
"""符号(-1/0/+1)と、行ごとの2つの大きさ(lo/hi)から重みを作る。符号が0でない要素は1ビットで lo か hi を選ぶ。"""
rows, cols = shape
row_bytes = (cols + 4) // 5
codes = _trits(blob[: rows * row_bytes], rows, cols)
nonzero = codes != 1
count = int(nonzero.sum())
bit_bytes = (count + 7) // 8
off = rows * row_bytes
if off + bit_bytes + 4 * rows != len(blob):
raise ValueError("five_value の大きさが合いません")
bits = np.unpackbits(np.frombuffer(blob[off: off + bit_bytes], dtype=np.uint8), bitorder="little")[:count].astype(bool)
off += bit_bytes
lo = np.frombuffer(blob[off: off + 2 * rows], dtype=np.float16)
hi = np.frombuffer(blob[off + 2 * rows: off + 4 * rows], dtype=np.float16)
is_hi = np.zeros((rows, cols), dtype=bool)
is_hi[nonzero] = bits
magnitude = np.where(is_hi, hi[:, None], lo[:, None])
sign = codes.astype(np.int8) - 1
return (sign.astype(np.float16) * magnitude).astype(np.float16)
def int_n(blob: bytes, shape: tuple, bits: int) -> np.ndarray:
"""行ごとの尺度(fp16)を掛ける 6bit / 8bit の整数の重み。"""
rows = shape[0]
total = int(np.prod(shape))
cols = total // rows
body, scales = blob[: -2 * rows], np.frombuffer(blob[-2 * rows:], dtype=np.float16)
if bits == 8:
q = np.frombuffer(body, dtype=np.int8).astype(np.int32)[:total]
elif bits == 6:
b = np.frombuffer(body, dtype=np.uint8).reshape(-1, 3).astype(np.uint32)
packed = b[:, 0] | (b[:, 1] << 8) | (b[:, 2] << 16)
q = np.stack([(packed >> s) & 0x3F for s in (0, 6, 12, 18)], axis=1).ravel()[:total].astype(np.int32) - 32
else:
raise ValueError(f"対応していないビット数です: {bits}")
if q.size != total:
raise ValueError("int の大きさが合いません")
return (q.reshape(rows, cols).astype(np.float32) * scales.astype(np.float32)[:, None]).reshape(shape)
def read_container(path: str, fmt: str) -> dict:
"""先頭8バイトが見出しの長さ、続く JSON の index に従って、各テンソルの値を並べたファイルを読む。"""
tensors = {}
with open(path, "rb") as f:
header = json.loads(f.read(int.from_bytes(f.read(8), "little")))
if header.get("format") != fmt:
raise ValueError(f"形式が違います: {header.get('format')}")
for e in header["index"]:
blob = f.read(e["b"])
if len(blob) != e["b"]:
raise ValueError("ファイルが途中で終わっています")
kind, shape = e["k"], tuple(e["shape"])
if kind == "five_value":
tensors[e["n"] + ".weight"] = five_value(blob, shape)
elif kind in ("int6", "int8"):
tensors[e["n"]] = int_n(blob, shape, int(kind[3:]))
elif kind == "fp16":
tensors[e["n"]] = np.frombuffer(blob, dtype=np.float16).reshape(shape).copy()
else:
raise ValueError(f"対応していない種類です: {kind}")
if f.read(1):
raise ValueError("末尾に余りのバイトがあります")
return tensors
def state_dict(tensors: dict):
import re
import torch
out = {}
for name, value in tensors.items():
if name.endswith("num_batches_tracked"):
count = float(np.asarray(value, dtype=np.float32).reshape(-1)[0])
out[name] = torch.tensor(int(count) if np.isfinite(count) else 0, dtype=torch.int64)
continue
arr = np.ascontiguousarray(value.astype(np.float32))
if arr.ndim == 2 and re.search(r"\.conv\.pointwise_conv[12]\.weight$", name):
arr = arr[:, :, None]
out[name] = torch.from_numpy(arr)
return out
def waveform(path: str) -> np.ndarray:
with wave.open(path, "rb") as r:
if (r.getframerate(), r.getnchannels(), r.getsampwidth()) != (16000, 1, 2):
raise ValueError("16kHz・モノラル・16bit の WAV ではありません")
pcm = r.readframes(r.getnframes())
return np.frombuffer(pcm, dtype=np.int16).astype(np.float32) / 32768.0
def installed_packages() -> list:
from importlib import metadata
seen = {}
for dist in metadata.distributions():
name = dist.metadata["Name"]
if name and name.lower() not in seen:
seen[name.lower()] = f"{name}=={dist.version}"
return sorted(seen.values(), key=str.lower)
def main(job_path: str, out_path: str) -> None:
import torch
import transformers
from transformers import AutoProcessor, GenerationConfig, ParakeetForTDT, ParakeetTDTConfig
with open(job_path, encoding="utf-8") as f:
job = json.load(f)
os.environ["HF_HUB_OFFLINE"] = "1"
os.environ["TRANSFORMERS_OFFLINE"] = "1"
torch.set_num_threads(job["threads"])
t0 = time.time()
record = {"transformers": transformers.__version__, "torch": torch.__version__, "dtype": "float32",
"threads": job["threads"], "python": platform.python_version(), "packages": installed_packages(),
"measured_at": datetime.now(timezone.utc).isoformat()}
inner = extract_inner(job["archive"], job["inner"], job["work"])
if sha256_of(inner) != job["inner_sha256"]:
raise ValueError("展開したファイルの SHA-256 が固定した値と違います")
tensors = read_container(inner, job["format"])
model = ParakeetForTDT(ParakeetTDTConfig.from_pretrained(job["base_dir"]))
missing, unexpected = model.load_state_dict(state_dict(tensors), strict=False)
if missing or unexpected:
raise ValueError(f"テンソルの名前が合いません: 不足 {list(missing)[:5]}・余り {list(unexpected)[:5]}")
record["tensors"] = len(tensors)
record["parameters"] = sum(p.numel() for p in model.parameters())
model.eval()
model.generation_config = GenerationConfig.from_pretrained(job["base_dir"])
processor = AutoProcessor.from_pretrained(job["base_dir"])
record["load_sec"] = round(time.time() - t0, 1)
def decode(item: dict):
wav = waveform(item["path"])
started = time.time()
inputs = processor([wav], sampling_rate=16000, return_tensors="pt", padding=True)
with torch.no_grad():
out = model.generate(input_features=inputs["input_features"], attention_mask=inputs.get("attention_mask"))
text = processor.batch_decode(getattr(out, "sequences", out), skip_special_tokens=True)[0].strip()
return text, time.time() - started, wav.size / 16000
decode(job["inputs"][0]) # 最初の1回は初期化の時間が入るので、測らずに流す
outputs = []
for item in job["inputs"]:
text, seconds, audio_seconds = decode(item)
outputs.append({"id": item["id"], "language": item["language"], "hypothesis": text,
"decode_sec": round(seconds, 2), "audio_sec": round(audio_seconds, 2)})
print(f"[asr-cpu] {item['id']}: {audio_seconds:.1f}秒の音声を {seconds:.1f}秒で文字起こし", flush=True)
record["outputs"] = outputs
with open(out_path, "w", encoding="utf-8") as f:
json.dump(record, f, ensure_ascii=False)
if __name__ == "__main__":
main(sys.argv[1], sys.argv[2])
We ran it as asr-venv/bin/python asr_cpu_run.py job.json out.json. The job.json we passed (as used; the audio is shown by its SHA-256):
{
"archive": "/tmp/lmw-asr-cpu-lcq29pwq/phonon-2.bps.tar.zst",
"inner": "model.fermion",
"inner_sha256": "4b6bfa3a12cc3c4e0a54f2ab3ec4ca7a842b09e5c7ecfc8e7ca0ac6cc8c11468",
"format": "fermion-five-value-parakeet-v1",
"work": "/tmp/lmw-asr-cpu-lcq29pwq",
"base_dir": "/tmp/lmw-asr-cpu-lcq29pwq/base",
"threads": 3,
"inputs": [
{
"id": "weather",
"language": "en",
"wav_sha256": "1258dcd9f83234db2601a928bd8facddc1c0c2751d7a90765375cc9197dc0651"
},
{
"id": "numbers",
"language": "en",
"wav_sha256": "ce0b8d12768dd1d9f5fdbf5fd0a0a98c35b5f5b7c522b814c5da01579812f430"
},
{
"id": "surprise",
"language": "en",
"wav_sha256": "fcd7c8d3f8cce918a56d959071a298d3a994d29c78586a2868b2a0a0c7c56e40"
}
]
}
Files we downloaded (SHA-256)
98125795b6dda72f5c6eee9ba33d19815df65dcb18b50a357bf9f73c9935309e phonon-2.bps.tar.zst None None 4b6bfa3a12cc3c4e0a54f2ab3ec4ca7a842b09e5c7ecfc8e7ca0ac6cc8c11468 model.fermion
Package versions in the environment (full list)
annotated-doc==0.0.5 annotated-types==0.8.0 anyio==4.15.1 certifi==2026.7.22 cffi==2.1.1 charset-normalizer==3.5.2 click==8.5.0 cloudpickle==3.1.2 decorator==5.3.1 filelock==3.32.3 fsspec==2026.7.0 h11==0.16.0 hf-xet==1.6.0 httpcore2==2.13.1 httpcore==1.0.9 httpx2==2.13.1 httpx==0.28.1 huggingface_hub==1.33.0 idna==3.20 Jinja2==3.1.6 joblib==1.6.0 lazy-loader==0.6 librosa==1.0.0 llvmlite==0.50.0 markdown-it-py==4.2.0 MarkupSafe==3.0.3 mdurl==0.1.2 mpmath==1.3.0 msgpack==1.2.3 narwhals==2.26.0 networkx==3.6.1 numba==0.68.0 numpy==2.5.3 packaging==26.3 pip==24.0 platformdirs==4.12.2 pooch==1.9.0 pycparser==3.0 pydantic==2.13.5 pydantic_core==2.46.5 Pygments==2.21.0 PyYAML==6.0.3 regex==2026.9.29 requests==2.34.2 rich==15.0.0 safetensors==0.8.0 scikit-learn==1.9.1 scipy==1.18.1 setuptools==78.1.0 shellingham==1.5.4 sniffio==1.3.1 soundfile==0.14.0 soxr==1.1.0 sympy==1.14.0 threadpoolctl==3.7.0 tokenizers==0.22.2 torch==2.14.1+cpu tqdm==4.70.1 transformers==5.13.0 truststore==0.10.4 typer==0.27.2 typing-inspection==0.4.4 typing_extensions==4.16.0 urllib3==2.8.0 zstandard==0.25.0
How to Get It
Phonon-2 is available as open-weight on Hugging Face and is not a gated model, so it can be downloaded and used without prior application for use or license agreement procedures. As a distribution format, MLX-compatible weights optimized for Apple Silicon are provided.
In Mac environments equipped with Apple Silicon, transcription can be executed directly from the command line by installing the officially provided Python package and MLX-related libraries. Run the following commands in the terminal to set up the necessary dependencies:
pip install fermion-research
pip install mlx mlx-audio mlx-lm soundfile scipy zstandard
Once the installation is complete, transcribe an audio file (such as.wav format) using one of the following commands:
phonon transcribe recording.wav
fermion transcribe phonon-2 recording.wav
For Linux (x86-64 and Arm architectures) and Windows CPU environments, a dedicated execution engine is published in the GitHub repository (github.com/fermionresearch/phonon). In addition, for developers who want to simplify environment construction or use GPU acceleration, official Docker images are also provided.
When running the Docker container in a CPU environment, use the following command:
docker run --rm -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cpu:2.0.3 transcribe phonon-2 /audio/recording.wav
When running in a CUDA environment equipped with an NVIDIA GPU, start the container by specifying the GPU option as follows:
docker run --rm --gpus all -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cuda:1.0.4 transcribe phonon-2 /audio/recording.wav
Regarding the license aspect, the model weights themselves are released under the same “CC-BY-4.0" as the base model, and changes are described in the NOTICE file within the repository. In addition, the provided command-line tools and the code in the repository are applied with the “Apache 2.0" license.
Other Models for the Same Task
Recent audio models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site’s estimate.
- Audio8-ASR-Infinite Speech Recognition Model: 12GB+ VRAM, File List (4.1B, 12GB)
- YuE2-3B Music Generation Model for Lyrics and Style: 12GB+ VRAM (3.6B, 12GB)
- Irodori-TTS-v4.1-Anime: Our Generated Audio, File List (766M, 4GB)
What to Read Next
- Find models by VRAM (This model runs from the 4GB tier) → Other models that run on a 8GB GPU
- What Q8_0 mean and where to get this model → MLX format guide and models / Quantization and model-format glossary
- Learn about the publisher → NVIDIA: models, licenses and articles
- How to read WER → Benchmark glossary
Sources
Update History
- 2026-10-01: Added our own measurements: transcripts of English audio we made.

