{"id":114,"date":"2026-09-07T17:15:59","date_gmt":"2026-09-07T08:15:59","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/07\/irodori-tts-v41-anime-released\/"},"modified":"2026-09-23T07:39:01","modified_gmt":"2026-09-22T22:39:01","slug":"irodori-tts-v41-anime-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/07\/irodori-tts-v41-anime-released\/","title":{"rendered":"Irodori-TTS-v4.1-Anime Released for Local Voice Generation"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/phasefield-audio\/Irodori-TTS-v4.1-Anime\">phasefield-audio\/Irodori-TTS-v4.1-Anime<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-04<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>mit<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>safetensors<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>Irodori-TTS-v4.1-Anime, a Text-to-Speech (TTS) model that generates speech from Japanese text fine-tuned using anime-style voice data, has been released. This model is adjusted based on the base model &#8220;Aratako\/Irodori-TTS-v4.1-Small&#8221; using anime-style voice data.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Base model: Aratako\/Irodori-TTS-v4.1-Small<\/li>\n<li>License: MIT License<\/li>\n<li>Distribution format: safetensors<\/li>\n<li>Quantization variants: int8-weight-only, int8-dynamic, int4-weight-only, float8-weight-only, float8-dynamic<\/li>\n<\/ul>\n<p>Note that because the base model&#8217;s annotation pipeline has not been published, the fine-tuning data for this model was independently annotated. Therefore, the behavior of caption-based conditioning and emoji-based control may differ from the base model.<\/p>\n<h2>Performance and Quality<\/h2>\n<p>This model is a derivative of &#8220;Aratako\/Irodori-TTS-v4.1-Small&#8221;, and the performance of the original model is demonstrated by the following benchmarks.<\/p>\n<h3>Joyo Kanji Yomi Benchmark<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Model<\/th>\n<th style=\"text-align: right;\">Kana-CER \u2193<\/th>\n<th style=\"text-align: right;\">Kana-CER clipped \u2193<\/th>\n<th style=\"text-align: right;\">Sentence Kana-CER \u2193<\/th>\n<th style=\"text-align: right;\">Standard CER \u2193<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\">Irodori-TTS-600M-v3-VoiceDesign<\/td>\n<td style=\"text-align: right;\">8.49 \u00b1 0.21%<\/td>\n<td style=\"text-align: right;\">5.59 \u00b1 0.09%<\/td>\n<td style=\"text-align: right;\">2.45 \u00b1 0.02%<\/td>\n<td style=\"text-align: right;\">4.88 \u00b1 0.21%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Irodori-TTS-v4-Small (original)<\/td>\n<td style=\"text-align: right;\">7.43 \u00b1 0.17%<\/td>\n<td style=\"text-align: right;\">5.08 \u00b1 0.03%<\/td>\n<td style=\"text-align: right;\">2.89 \u00b1 0.03%<\/td>\n<td style=\"text-align: right;\">5.35 \u00b1 0.18%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><strong>Irodori-TTS-v4.1-Small<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>7.29 \u00b1 0.13%<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>5.03 \u00b1 0.03%<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>2.36 \u00b1 0.01%<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>4.69 \u00b1 0.02%<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>JSUT BASIC5000<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Model<\/th>\n<th style=\"text-align: right;\">Sentence Kana-CER \u2193<\/th>\n<th style=\"text-align: right;\">Standard CER \u2193<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\">Irodori-TTS-600M-v3-VoiceDesign<\/td>\n<td style=\"text-align: right;\">3.62 \u00b1 0.03%<\/td>\n<td style=\"text-align: right;\"><strong>7.19 \u00b1 0.05%<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Irodori-TTS-v4-Small (original)<\/td>\n<td style=\"text-align: right;\">3.49 \u00b1 0.02%<\/td>\n<td style=\"text-align: right;\">7.32 \u00b1 0.09%<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\"><strong>Irodori-TTS-v4.1-Small<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>3.43 \u00b1 0.01%<\/strong><\/td>\n<td style=\"text-align: right;\">7.22 \u00b1 0.12%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Compared to the traditional &#8220;Irodori-TTS-v4-Small&#8221;, the original model &#8220;Irodori-TTS-v4.1-Small&#8221; shows improved (lower) scores across all metrics in the Joyo Kanji Yomi Benchmark (Kana-CER, Kana-CER clipped, Sentence Kana-CER, and Standard CER). In addition, for JSUT BASIC5000, improvements are seen in Sentence Kana-CER, while Standard CER shows a slight increase. These improvements result from retraining the duration predictor individually in the original model.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Since this model is fine-tuned using anime-style voice data, it specializes in generating anime-style character voices. It is tagged with <code>text-to-speech<\/code> and generates speech from Japanese text input.<\/p>\n<p>Inheriting the features of the base model &#8220;Aratako\/Irodori-TTS-v4.1-Small&#8221;, the following capabilities are available:<\/p>\n<ul>\n<li>Voice cloning: Voice cloning using reference audio<\/li>\n<li>Voice Design: Designing voice quality<\/li>\n<li>Long-reference conditioning: Conditioning using long-duration reference audio<\/li>\n<li>Emoji-based style control: Style control using emojis (however, because this model uses independent annotations, behavior may differ from the base model)<\/li>\n<\/ul>\n<p>Additionally, as a constraint of the base model, only Japanese text input is currently supported. Furthermore, to stabilize voice cloning accuracy, clean reference audio of 30 seconds or longer is recommended.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 766M parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>4GB (laptop iGPU \/ phone class)<\/td>\n<td>F32<\/td>\n<td>2.9GB<\/td>\n<td>3.4GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is distributed in a Hugging Face repository.<\/p>\n<ul>\n<li>Repository: <a href=\"https:\/\/huggingface.co\/phasefield-audio\/Irodori-TTS-v4.1-Anime\">phasefield-audio\/Irodori-TTS-v4.1-Anime<\/a><\/li>\n<li>Distribution format: safetensors<\/li>\n<li>Quantized versions: Available in the <code>int8-weight-only<\/code>, <code>int8-dynamic<\/code>, <code>int4-weight-only<\/code>, <code>float8-weight-only<\/code>, and <code>float8-dynamic<\/code> directories within the repository.<\/li>\n<\/ul>\n<p>For inference code and installation instructions, please refer to the original &#8220;Irodori-TTS&#8221; repository. <a href=\"https:\/\/github.com\/Aratako\/Irodori-TTS\">GitHub: Aratako\/Irodori-TTS<\/a><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/phasefield-audio\/Irodori-TTS-v4.1-Anime\">phasefield-audio\/Irodori-TTS-v4.1-Anime<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-09-20: Rewrote the article from re-collected sources and restored it from draft to published.<\/li>\n<li>2026-09-23: Rebuilt the article (details are in the Japanese edition).<\/li>\n<li>2026-09-23: Rebuilt the article (details are in the Japanese edition).<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Discover Irodori-TTS-v4.1-Anime, an anime-style fine-tuned Text-to-Speech model based on Aratako\/Irodori-TTS-v4.1-Small. Available now with quantization op<\/p>\n","protected":false},"author":1,"featured_media":297,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[1761,209,211,1547,1764,234],"class_list":["post-114","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-irodori-tts-v4-1-anime-en","tag-safetensors-en","tag-tts-en","tag-verified","tag--en"],"lang":"en","translations":{"en":114,"ja":113},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/114","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=114"}],"version-history":[{"count":10,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/114\/revisions"}],"predecessor-version":[{"id":2778,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/114\/revisions\/2778"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/297"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=114"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=114"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=114"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}