Multilingual KWS Dataset Generation with Voice Design and Cloning
A completed synthetic speech data pipeline for keyword spotting across seven languages, combining voice design, identity-preserving cloning, ASR verification, and recoverable dataset assembly.

- My contribution
- Multilingual vocabulary mapping, voice-design and cloning adapters, speaker qualification, quality control, and resumable production orchestration.
- Outcome
- A completed seven-language pipeline for verified command, wake-word, and negative speech datasets, integrating voice design, reference cloning, and reproducible dataset assembly.
Project overview
I developed a multilingual synthetic speech dataset pipeline at Enhanced Robotics from July to September 2026. It creates training data for keyword spotting (KWS), including spoken commands, wake words, and negative phrases that a recognizer should treat as unknown speech.
The pipeline separates voice design from voice cloning. Voice-design models create candidate synthetic speaker identities. After those reference voices pass quality checks, cloning models reproduce each accepted identity across the required phrases. This preserves a consistent speaker identity across commands while allowing several synthesis engines to contribute to the corpus.
The engineering challenge was to make generation repeatable and auditable across languages and models: preserve vocabulary labels, reject incorrect speech, qualify speakers before expensive production, and recover missing samples without relaxing acceptance criteria.
Language coverage and vocabulary
The completed pipeline supports English, Chinese, Japanese, Korean, Spanish, French, and German. Voice design, reference cloning, vocabulary validation, and dataset assembly are integrated across all seven languages for command, wake-word, and negative speech generation.
| Language | Vocabulary and generation |
|---|---|
| English | 54 phrase classes in the command vocabulary |
| Chinese | 54 phrase classes, with language-specific class IDs and reviewed pronunciation aliases |
| Japanese | 52 phrase classes, with reviewed readings and renderer-specific pronunciation handling |
| Korean | Command, wake-word, and negative generation with language-specific vocabularies |
| Spanish | Command, wake-word, and negative generation with language-specific vocabularies |
| French | Command, wake-word, and negative generation with language-specific vocabularies |
| German | Command, wake-word, and negative generation with language-specific vocabularies |
The English, Chinese, and Japanese vocabularies connect to 50 shared semantic intents. A phrase class identifies one spoken form, while an intent identifies the device action it represents. Multiple phrase variants can therefore share an intent without being collapsed into one class. Validation rejects mismatched text and IDs before models are loaded.
Command positives and wake-word positives share the same identity, pronunciation, and completeness checks. Their negative datasets use separate reviewed phrase lists and are audited against positive forms both before generation and after transcription. A rejected positive clip is never automatically relabeled as a negative.
Voice design and cloning
I integrated model adapters around a common speaker-profile schema, deterministic seeds, and canonical 16 kHz reference audio. The model mix follows language and class capabilities rather than assuming every backend supports the same controls.
| Role | Model families |
|---|---|
| Synthetic voice design | Qwen3-TTS VoiceDesign, MOSS VoiceGenerator, VoxCPM2, Breeze-TTS-2, Irodori, FireRedTTS3 Instruct, and OmniVoice, selected by language |
| Reference voice cloning | Qwen3-TTS Base, CosyVoice3, VoxCPM2 Clone, Breeze Clone, Irodori Clone, FireRedTTS3 Base, OmniVoice Clone, and Chatterbox Multilingual V3, selected by language and dataset type |
| Speech verification | Faster-Whisper Small, with Large-v3 verification on a primary mismatch |
| Speaker verification | SpeechBrain ECAPA speaker embeddings for deduplication, identity checks, and qualification |
English and Chinese use six default voice-design sources and seven positive renderers. Japanese uses four voice-design sources and seven positive renderers. FireRed voice design is restricted to English and Chinese, while its reference cloning remains available for Japanese. Japanese negative generation uses six eligible renderers, with Irodori excluded by its renderer-specific capability policy.
Korean, Spanish, French, and German use Qwen and OmniVoice for voice design, with Qwen Base, CosyVoice3, FireRedTTS3 Base, OmniVoice Clone, and Chatterbox Multilingual V3 for reference rendering.
Positive generation retains 11 total takes per speaker and phrase. With seven eligible renderers, Qwen receives two slots and the remaining six share nine rotating slots. Additional models broaden the available synthesis paths without multiplying the planned corpus size. Negative generation uses one rotating renderer take per speaker and phrase.
Reference conditioning is adapted to each model’s audio and transcript requirements. Requested speaking styles are recorded according to what the backend actually applies; unsupported style requests become independently seeded natural variants rather than being reported as successful style control.
Quality checks and speaker qualification
The pipeline qualifies a reference voice before expanding it into a full command dataset:
- Validate and deduplicate anchors. Check audio validity, transcript agreement, and speaker embeddings; reject overly similar candidates and expand source-specific candidate pools when needed.
- Probe across renderers. Test every accepted anchor with every enabled renderer eligible for the selected probe classes. Check both phrase correctness and preservation of the reference identity.
- Qualify each branch independently. Require each renderer to meet command-coverage and pass-rate gates. Rank speakers using the weakest renderer’s pass rate, accepted-cell coverage, and speaker similarity.
- Select primary and reserve speakers. Establish a qualified, source-balanced pool before committing to the production plan.
Whisper Small performs the primary phrase check; mismatches receive a Large-v3 check before rejection. Reviewed Chinese Pinyin and Japanese Kana aliases remain tied to their phrase classes. Extra speech, repetitions, changed onsets, and unrelated homophones still fail verification.
Per-renderer ECAPA summaries include median, 10th-percentile, minimum, and low-score-ratio measurements. These support identity analysis without assuming that short-clip similarity distributions are identical across synthesis engines. A strong renderer cannot conceal a weak branch during qualification.
Production and recovery
An immutable production plan fixes the speaker, source, phrase class, take, renderer, shard, and role. Saved plans and resume fingerprints prevent a resumed run from silently changing its source allocation.
Each shard generates and verifies its planned cells. Only clips passing audio, transcript, and speaker checks enter assembly. Failed cells retain their rejection details and proceed through bounded recovery:
- Try distinct enabled alternatives that support the cell’s language and class.
- Use fresh deterministic seeds for bounded repair attempts after ordinary alternatives are exhausted.
- Preserve already accepted audio when resuming or repairing a run.
- Promote complete reserve speakers from the appropriate source when primary speakers remain incomplete.
- Keep unresolved cells explicit and write a missing-cell report if balanced final quotas cannot be met.
Recovered cells retain both the successful renderer’s quality results and the planned renderer’s failed checks. This makes recovery traceable rather than hiding failures behind the final audio files.
Outputs and observability
Positive datasets are exported as class-organized WAV files with a command mapping. Negative exports remain separate and receive an additional transcript-leakage audit before becoming the prepared unknown corpus. Optional dataset preprocessing and long-clip filtering happen after generation as separate operations.
The pipeline also records speaker manifests, immutable plans, quality reports, recovery summaries, model inventories, adapter hashes, source revisions, reference policies, and generation settings. Per-renderer listening montages support manual review.
Stage-tagged logs and timing records track vocabulary validation, voice generation, qualification, sharded production, recovery, and final validation. Append-only segment history preserves failed and resumed work, while invocation summaries expose the current run’s progress. Separate model runtimes and bounded generation processes accommodate conflicting dependencies and release generation resources before shared verification loads.
Outcome and evaluation scope
The completed work provides a repeatable pipeline for synthetic command, wake-word, and negative speech data across all seven supported languages, with explicit vocabulary identity, speaker qualification, recoverable production, and source provenance.
The downstream training, quantization, and device integration are covered in Multilingual TC-ResNet Keyword Spotting on ESP32-S3.
The model allocation and take counts describe the generation protocol, not a measured final corpus size. Broader model coverage supplies additional voice candidates and recovery options. Downstream KWS accuracy improvements are not quantified in this project record; they require separate evaluation on real speech.
Technology
Python, shell orchestration, multilingual text-to-speech, synthetic voice design, reference voice cloning, Faster-Whisper, SpeechBrain ECAPA, deterministic generation, sharded processing, manifest-based recovery, and stage-level observability.