Alania Synthetic Speech TR: Turkish Speech Dataset Released

At a Glance
| Item | Value |
|---|---|
| Repository | cloud0day3/alania-synthetic-speech-tr |
| Published | 2026-09-30 |
| Source type | Unverified (not confirmed by a primary source) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
It is reported that a Turkish synthetic speech dataset named “alania-synthetic-speech-tr" has been released. However, please note that at the time of writing this article, official announcements have not been confirmed, making this unconfirmed information. The dataset is reported to be a large-scale Turkish speech data collection spanning a total of 3.451 hours, aimed at training text-to-speech (TTS) and automatic speech recognition (ASR) models.
All of the recorded speech is generated by AI, and it is reported that no real human voices are included in the dataset. To improve the current situation where open Turkish TTS datasets are extremely scarce, it is reported that part of the data created during the development of PatientDesk AI’s Turkish voice model, “Alania-2," has been made widely available for the community to use.
Specifications
The main specifications of this dataset confirmed from the materials are as follows:
- Total Recording Time: 3.451 hours (2,051,810 records)
- Sampling Rate: 48 kHz
- Voice Composition: 2,752 designed voice IDs and one-off designed voices
- Gender Balance: Female approx. 53,5%, Male approx. 46,5%
- Generation Model:
openbmb/VoxCPM2(Apache-2.0) - Generation Settings: 10 diffusion steps, 2,0 guidance scale (as of September 2026)
- Reading Formats:
reference: Normal reading using designed voices –reference+instruction: Reading following style instructions such as “Slowly and softly, like a calm nurse" –voice-design: New voices generated solely from text descriptions- Recorded Texts: Everyday service terminology such as reservations, banking, delivery, and support, numbers, dates, times, addresses, names, and web/encyclopedia texts
- License Conditions:
cc-byandvoicesconfigurations are under CC BY 4.0, while thecc-by-saconfiguration is under CC BY-SA 4.0 because it includes the Turkish Wikipedia. It can be used for any purpose including commercial use and model training, but attribution to PatientDesk AI is required.
Performance and Quality
As a result of the dataset composition and quality control by the publishers, the following has been reported:
→ Scroll horizontally to see all columns
| Configuration | Text Content | License | Record Count | Time |
|---|---|---|---|---|
cc-by (Default) |
Sentences for voice assistants and FineWeb-2 Turkish | CC BY 4.0 | 1,671,085 | 2,592 |
cc-by-sa |
Turkish Wikipedia | CC BY-SA 4.0 | 380,725 | 859 |
voices |
Reference recordings, texts, and descriptions for each designed voice | CC BY 4.0 | 2,752 Voices |
To ensure data quality, the following automated quality checks were reportedly conducted on all records:
- Verification with
Whisper large-v3 (turbo): Character Error Rate (CER) within a maximum of 5% or only a single correction UTMOS22: Score of 2,5 or higher- Voice similarity and gender consistency checks: Similarity confirmation with reference voice via
ECAPA, and gender determination by pitch - Speech Rate: Within the range of 7–25 characters per second
- Filtering: Removal of slight noise in the 7–12 kHz band originating from the generator
As a result of these automated checks, it is reported that high quality scores are secured with a median CER of 0,0 and a UTMOS of 3,59. Note that CER is an index indicating the character-level error rate of speech recognition, meaning lower numbers represent better performance.
On the other hand, as a known issue (v1.0) identified on October 1, 2026, it is reported that 140,338 records (approximately 343 hours, about 7% of the total) contain glitches where abbreviations are read out by their alphabet names (e.g., reading “KDV" as “ke de ve" instead of “kadeve") or Roman numerals are read as characters. Filtering lists for these records are provided, and they are scheduled to be fixed in the next version, v1.1.
Additionally, since this dataset is purely synthetic speech, it contains accents and intonations specific to the generation model VoxCPM2, and tends to be more uniform and cleaner than actual human voices. Therefore, for practical model training, it is recommended to use this dataset in combination with actual human voice data rather than by itself.
Strengths and Use Cases
This dataset is envisioned primarily for training foundational and derivative models for text-to-speech (TTS) and speech recognition (ASR) in Turkish. In particular, it is expected to be utilized as a large-scale training resource in areas where high-quality Turkish speech corpora available under open licenses are lacking.
Specifically, it is said to be suited for the following developments and experiments:
- Building TTS that supports diverse speakers and style instructions: Suitable for training speech style control via natural language instructions such as “like a calm nurse" in addition to 2,752 fixed voice IDs, and voice design capabilities that synthesize new voices from text descriptions.
- Conversational agents specialized in everyday service domains: Richly contains expressions requiring normalization—such as numbers, dates/times, amounts, phone numbers, addresses, and personal names—that frequently appear in reservation handling, banking operations, delivery status checks, and customer support, aiding in the development of practical voice assistants.
- Learning ASR and robust TTS assuming real-world acoustic environments: Approximately 45% of all records include
audio_channeldata applied with simulated room reverberation, various microphones (headset, lavalier, laptop, smartphone), background noise, and volume fluctuations. This can be used to strengthen the resilience of recognition and generation models not only in clean audio but also in noisy environments. - Data augmentation with mitigated license concerns: Because it is completely synthetic data containing no real human voices, it can be utilized as augmented data to make up for shortages in existing human voice datasets while avoiding restrictions regarding portrait rights and voice rights.
For responsible use, the publishers request that the generated voices not be used to impersonate real individuals or falsely claimed as human voices, and that it be clearly stated that they are AI-generated when required by laws and regulations (such as the EU AI Act).
How to Get It
It is reported that this dataset is available in Parquet format in the Hugging Face dataset repository cloud0day3/alania-synthetic-speech-tr.
Since it is not designated as a gated dataset (restrictions requiring prior agreement to terms of use), direct downloading and loading are possible with a Hugging Face account.
From a Python environment, it can be loaded using Hugging Face’s datasets library, as well as polars or dask. For example, when using the datasets library, data can be retrieved by specifying the configuration as follows:
from datasets import load_dataset
# To load the default cc-by configuration
dataset = load_dataset("cloud0day3/alania-synthetic-speech-tr", "cc-by")
Depending on your use case, it is also possible to load by specifying the cc-by-sa configuration containing Wikipedia-derived text, or the voices configuration summarizing reference information for each speaker. Furthermore, a filtering file known_issues/initialism_readings_v1.0.csv.gz to avoid reading errors of abbreviations included in v1.0 is also provided within the repository.
What to Read Next
- How to read CER → Benchmark glossary
Sources
This article contains unverified information. We will append an update note once it is confirmed by a primary source.
