{"id":8886,"date":"2026-10-02T05:10:46","date_gmt":"2026-10-01T20:10:46","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/alania-synthetic-speech-tr-dataset\/"},"modified":"2026-10-02T05:10:46","modified_gmt":"2026-10-01T20:10:46","slug":"alania-synthetic-speech-tr-dataset","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/alania-synthetic-speech-tr-dataset\/","title":{"rendered":"Alania Synthetic Speech TR: Turkish Speech Dataset Released"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/datasets\/cloud0day3\/alania-synthetic-speech-tr\">cloud0day3\/alania-synthetic-speech-tr<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-30<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Unverified (not confirmed by a primary source)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>It is reported that a Turkish synthetic speech dataset named &#8220;alania-synthetic-speech-tr&#8221; has been released. However, please note that at the time of writing this article, official announcements have not been confirmed, making this unconfirmed information. The dataset is reported to be a large-scale Turkish speech data collection spanning a total of 3.451 hours, aimed at training text-to-speech (TTS) and automatic speech recognition (ASR) models.<\/p>\n<p>All of the recorded speech is generated by AI, and it is reported that no real human voices are included in the dataset. To improve the current situation where open Turkish TTS datasets are extremely scarce, it is reported that part of the data created during the development of PatientDesk AI&#8217;s Turkish voice model, &#8220;Alania-2,&#8221; has been made widely available for the community to use.<\/p>\n<h2>Specifications<\/h2>\n<p>The main specifications of this dataset confirmed from the materials are as follows:<\/p>\n<ul>\n<li><strong>Total Recording Time<\/strong>: 3.451 hours (2,051,810 records)<\/li>\n<li><strong>Sampling Rate<\/strong>: 48 kHz<\/li>\n<li><strong>Voice Composition<\/strong>: 2,752 designed voice IDs and one-off designed voices<\/li>\n<li><strong>Gender Balance<\/strong>: Female approx. 53,5%, Male approx. 46,5%<\/li>\n<li><strong>Generation Model<\/strong>: <code>openbmb\/VoxCPM2<\/code> (Apache-2.0)<\/li>\n<li><strong>Generation Settings<\/strong>: 10 diffusion steps, 2,0 guidance scale (as of September 2026)<\/li>\n<li><strong>Reading Formats<\/strong>: <\/li>\n<li><code>reference<\/code>: Normal reading using designed voices &#8211; <code>reference+instruction<\/code>: Reading following style instructions such as &#8220;Slowly and softly, like a calm nurse&#8221; &#8211; <code>voice-design<\/code>: New voices generated solely from text descriptions<\/li>\n<li><strong>Recorded Texts<\/strong>: Everyday service terminology such as reservations, banking, delivery, and support, numbers, dates, times, addresses, names, and web\/encyclopedia texts<\/li>\n<li><strong>License Conditions<\/strong>: <code>cc-by<\/code> and <code>voices<\/code> configurations are under CC BY 4.0, while the <code>cc-by-sa<\/code> configuration is under CC BY-SA 4.0 because it includes the Turkish Wikipedia. It can be used for any purpose including commercial use and model training, but attribution to PatientDesk AI is required.<\/li>\n<\/ul>\n<h2>Performance and Quality<\/h2>\n<p>As a result of the dataset composition and quality control by the publishers, the following has been reported:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Configuration<\/th>\n<th>Text Content<\/th>\n<th>License<\/th>\n<th style=\"text-align: right;\">Record Count<\/th>\n<th style=\"text-align: right;\">Time<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>cc-by<\/code> (Default)<\/td>\n<td>Sentences for voice assistants and FineWeb-2 Turkish<\/td>\n<td>CC BY 4.0<\/td>\n<td style=\"text-align: right;\">1,671,085<\/td>\n<td style=\"text-align: right;\">2,592<\/td>\n<\/tr>\n<tr>\n<td><code>cc-by-sa<\/code><\/td>\n<td>Turkish Wikipedia<\/td>\n<td>CC BY-SA 4.0<\/td>\n<td style=\"text-align: right;\">380,725<\/td>\n<td style=\"text-align: right;\">859<\/td>\n<\/tr>\n<tr>\n<td><code>voices<\/code><\/td>\n<td>Reference recordings, texts, and descriptions for each designed voice<\/td>\n<td>CC BY 4.0<\/td>\n<td style=\"text-align: right;\">2,752 Voices<\/td>\n<td style=\"text-align: right;\"><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>To ensure data quality, the following automated quality checks were reportedly conducted on all records:<\/p>\n<ul>\n<li>Verification with <code>Whisper large-v3 (turbo)<\/code>: Character Error Rate (CER) within a maximum of 5% or only a single correction<\/li>\n<li><code>UTMOS22<\/code>: Score of 2,5 or higher<\/li>\n<li>Voice similarity and gender consistency checks: Similarity confirmation with reference voice via <code>ECAPA<\/code>, and gender determination by pitch<\/li>\n<li>Speech Rate: Within the range of 7\u201325 characters per second<\/li>\n<li>Filtering: Removal of slight noise in the 7\u201312 kHz band originating from the generator<\/li>\n<\/ul>\n<p>As a result of these automated checks, it is reported that high quality scores are secured with a median CER of 0,0 and a UTMOS of 3,59. Note that CER is an index indicating the character-level error rate of speech recognition, meaning lower numbers represent better performance.<\/p>\n<p>On the other hand, as a known issue (v1.0) identified on October 1, 2026, it is reported that 140,338 records (approximately 343 hours, about 7% of the total) contain glitches where abbreviations are read out by their alphabet names (e.g., reading &#8220;KDV&#8221; as &#8220;ke de ve&#8221; instead of &#8220;kadeve&#8221;) or Roman numerals are read as characters. Filtering lists for these records are provided, and they are scheduled to be fixed in the next version, v1.1.<\/p>\n<p>Additionally, since this dataset is purely synthetic speech, it contains accents and intonations specific to the generation model <code>VoxCPM2<\/code>, and tends to be more uniform and cleaner than actual human voices. Therefore, for practical model training, it is recommended to use this dataset in combination with actual human voice data rather than by itself.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>This dataset is envisioned primarily for training foundational and derivative models for text-to-speech (TTS) and speech recognition (ASR) in Turkish. In particular, it is expected to be utilized as a large-scale training resource in areas where high-quality Turkish speech corpora available under open licenses are lacking.<\/p>\n<p>Specifically, it is said to be suited for the following developments and experiments:<\/p>\n<ul>\n<li><strong>Building TTS that supports diverse speakers and style instructions<\/strong>: Suitable for training speech style control via natural language instructions such as &#8220;like a calm nurse&#8221; in addition to 2,752 fixed voice IDs, and voice design capabilities that synthesize new voices from text descriptions.<\/li>\n<li><strong>Conversational agents specialized in everyday service domains<\/strong>: Richly contains expressions requiring normalization\u2014such as numbers, dates\/times, amounts, phone numbers, addresses, and personal names\u2014that frequently appear in reservation handling, banking operations, delivery status checks, and customer support, aiding in the development of practical voice assistants.<\/li>\n<li><strong>Learning ASR and robust TTS assuming real-world acoustic environments<\/strong>: Approximately 45% of all records include <code>audio_channel<\/code> data applied with simulated room reverberation, various microphones (headset, lavalier, laptop, smartphone), background noise, and volume fluctuations. This can be used to strengthen the resilience of recognition and generation models not only in clean audio but also in noisy environments.<\/li>\n<li><strong>Data augmentation with mitigated license concerns<\/strong>: Because it is completely synthetic data containing no real human voices, it can be utilized as augmented data to make up for shortages in existing human voice datasets while avoiding restrictions regarding portrait rights and voice rights.<\/li>\n<\/ul>\n<p>For responsible use, the publishers request that the generated voices not be used to impersonate real individuals or falsely claimed as human voices, and that it be clearly stated that they are AI-generated when required by laws and regulations (such as the EU AI Act).<\/p>\n<h2>How to Get It<\/h2>\n<p>It is reported that this dataset is available in Parquet format in the Hugging Face dataset repository <code>cloud0day3\/alania-synthetic-speech-tr<\/code>.<\/p>\n<p>Since it is not designated as a gated dataset (restrictions requiring prior agreement to terms of use), direct downloading and loading are possible with a Hugging Face account.<\/p>\n<p>From a Python environment, it can be loaded using Hugging Face&#8217;s <code>datasets<\/code> library, as well as <code>polars<\/code> or <code>dask<\/code>. For example, when using the <code>datasets<\/code> library, data can be retrieved by specifying the configuration as follows:<\/p>\n<pre><code class=\"language-python\">from datasets import load_dataset\n\n# To load the default cc-by configuration\ndataset = load_dataset(&quot;cloud0day3\/alania-synthetic-speech-tr&quot;, &quot;cc-by&quot;)\n<\/code><\/pre>\n<p>Depending on your use case, it is also possible to load by specifying the <code>cc-by-sa<\/code> configuration containing Wikipedia-derived text, or the <code>voices<\/code> configuration summarizing reference information for each speaker. Furthermore, a filtering file <code>known_issues\/initialism_readings_v1.0.csv.gz<\/code> to avoid reading errors of abbreviations included in v1.0 is also provided within the repository.<\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>How to read CER<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-benchmarks-en\/\">Benchmark glossary<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/datasets\/cloud0day3\/alania-synthetic-speech-tr\">cloud0day3\/alania-synthetic-speech-tr<\/a><\/li>\n<\/ul>\n<blockquote>\n<p><strong>This article contains unverified information.<\/strong> We will append an update note once it is confirmed by a primary source.<\/p>\n<\/blockquote>\n","protected":false},"excerpt":{"rendered":"<p>Explore alania-synthetic-speech-tr, a large-scale synthetic Turkish speech dataset for TTS and ASR training, including specs and performance.<\/p>\n","protected":false},"author":1,"featured_media":8885,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[203],"tags":[2833,2835,1565,2837,2839,213,1017],"class_list":["post-8886","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-image-video-and-audio","tag-alania-synthetic-speech-tr-en","tag-cloud0day3-en","tag-unverified","tag-voxcpm2-en","tag--en"],"lang":"en","translations":{"en":8886,"ja":8884},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8886","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=8886"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/8886\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/8885"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=8886"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=8886"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=8886"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}