Dual Transcripts Expand Minnan Speech Training Data
WenetSpeech-Min pairs Minnan and Mandarin text for around 10,000 hours of online speech, creating shared training and evaluation material for ASR and TTS.
Underlying Paper
WenetSpeech-Min: A Large-Scale Minnan Speech Corpus with Dual Transcriptions for Dialectal Speech Processing
Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce WenetSpeech-Min, an open-source corpus comprising around 10,000 hours of Minnan speech collected from diverse online media, with paired Minnan and Mandarin transcripts for every utterance. We further establish an automatic speech recognition (ASR) benchmark covering both Minnan and Mandarin transcripts and a text-to-speech synthesis (TTS) benchmark using Minnan transcripts, with manually verified evaluation sets for both tasks. To assess the effectiveness of the corpus, we train ASR and TTS models on WenetSpeech-Min and compare them with representative systems on the proposed benchmarks. The resulting models outperform the evaluated open-source models on most metrics and achieve competitive performance against commercial systems. We will release the corpus, benchmarks, and models to facilitate reproducible research on Minnan speech technology.
Minnan speech technology has had a data problem: available corpora are relatively small, and paired Minnan–Mandarin transcriptions are rare at a scale useful for training modern speech models. That pairing matters because dialectal speech systems must contend with both what speakers say in Minnan and how the same content is represented in Mandarin-oriented text workflows. WenetSpeech-Min addresses that gap with an open corpus assembled from diverse online media, providing paired Minnan and Mandarin transcripts for each utterance.
Core Contribution
The central contribution is the dataset rather than a new model architecture. The authors present around 10,000 hours of Minnan speech with two transcription views, then build task-specific benchmarks on top of it: automatic speech recognition for both Minnan and Mandarin transcripts, and text-to-speech synthesis from Minnan transcripts. The paper also describes manually verified evaluation sets for both tasks.
That framing is useful. A single Minnan transcript corpus would primarily support dialect recognition or synthesis; a paired corpus can instead support models and evaluations that cross the dialect–Mandarin boundary. The authors position the release as reproducible infrastructure: corpus, benchmarks, and trained models are intended to be released together rather than leaving later users to reconstruct preprocessing and test sets.
Technical Approach
The supplied figure inventory identifies a data-construction pipeline and reports distributions across content domains, source-group WV-MOS scores, signal-to-noise ratio, and utterance duration. This indicates that the collection process treats source diversity and recording conditions as dataset properties to be measured, rather than presenting the hours as a uniform studio-quality collection. Figure 1 provides the paper's overview of that construction process.
The second figure characterizes the resulting material by domain and acoustic properties. For practitioners, these distributions are as consequential as the headline duration: a corpus drawn from online media can expose ASR and TTS systems to the channel variation and uneven utterance lengths that occur outside controlled recordings. Paired transcripts for every utterance are the differentiating annotation choice, while the two benchmark tracks make the dataset usable for both recognition and synthesis work.
The paper evaluates models trained on WenetSpeech-Min against representative open-source systems and commercial systems. The supplied abstract reports that the corpus-trained systems outperform the evaluated open-source models on most metrics and are competitive with commercial systems. That is a meaningful claim, but the available material here does not provide the individual table values, test-set sizes, model configurations, or commercial-system scores needed to establish the margin of improvement.
Results and Analysis
The directly supported quantitative result is scale: around 10,000 hours of speech, with two transcript forms attached to each utterance. The benchmark scope is also concrete: two ASR transcription targets and one Minnan TTS setting, backed by manually verified evaluation sets. Relative to prior Minnan resources described by the authors as limited, that combination should reduce a practical bottleneck for groups that need aligned dialect and Mandarin supervision without building a corpus from scratch.
The reported comparison outcome supports treating WenetSpeech-Min as a substantial dataset-and-benchmark contribution, not yet as proof that one training recipe dominates across Minnan speech processing. “Most metrics” leaves open where the open-source systems remain stronger, and “competitive” against commercial systems does not establish a consistent win. The paper's value is therefore likely to be highest for teams building ASR, transcription normalization, or Minnan TTS systems that need a common public evaluation basis and paired text supervision.
Scope and Caveats
The evidence is centered on ASR and TTS. It does not, from the supplied information, establish effectiveness for dialogue, speech translation, speaker-related tasks, or dialects beyond Minnan. Collection from diverse online media also makes the quality distributions important: performance averaged across a benchmark may not transfer equally across source domains, noise levels, or short and long utterances. Finally, the abstract-level result summary does not expose per-condition metrics, so readers should inspect the paper's evaluation tables before selecting a model on the basis of a small reported advantage.
Evidence Box
moderateKey Claims
- •Paired Minnan and Mandarin transcripts enable dialectal speech modeling
- •A public corpus and benchmarks support reproducible Minnan ASR and TTS research
- •Models trained on the corpus outperform evaluated open-source systems on most metrics
- •Corpus-trained systems are competitive with evaluated commercial systems
Key Results
- •Around 10,000 hours of Minnan speech from diverse online media
- •2 paired transcription forms for every utterance: Minnan and Mandarin
- •2 ASR transcription targets evaluated with manually verified test sets
- •1 Minnan-transcript TTS benchmark with a manually verified evaluation set
Limitations & Caveats
- •Evaluation is restricted to 2 task families: ASR and TTS
- •Audio originates from online media with varying WV-MOS, SNR, and utterance-duration distributions
- •No evidence presented here for tasks beyond Minnan–Mandarin speech processing
- •Abstract-level comparison claims omit per-system metric values and margins