Multilingual TC-ResNet Keyword Spotting on ESP32-S3
Trained and deployed seven-language TC-ResNet14 command models with 2–3-second audio windows and a shared TC-ResNet8 wake-word model with a 1-second window on ESP32-S3, using an INMP441 microphone.

- My contribution
- TC-ResNet training and evaluation, audio preprocessing, ONNX and INT8 export, INMP441 capture, ESP32-S3 inference, and latency optimization.
- Outcome
- On-device command recognition across seven languages and a shared wake-word model, with a repeatable training-to-firmware workflow and measured live-microphone processing optimizations.
Project overview
I trained and deployed compact keyword-spotting models on an ESP32-S3 with an INMP441 digital microphone at Enhanced Robotics from July to October 2026. The project covers command recognition in English, Chinese, Japanese, Korean, Spanish, French, and German, alongside one shared wake-word model for “Hi Tenniix.”
The command models use TC-ResNet14 with 2- or 3-second audio windows. The separate wake-word model uses TC-ResNet8 with a 1-second window. Each model has its own input duration and labels, carried through training, export, and firmware generation.
This work connects the multilingual speech dataset pipeline to embedded inference. My scope included model training, evaluation, quantization, audio capture, feature extraction, device integration, and profiling the complete microphone-to-prediction path.
Models and audio inputs
| Task | Architecture | Audio window | Coverage |
|---|---|---|---|
| Spoken commands | TC-ResNet14 | 2 or 3 seconds | Language-specific models across all seven languages |
| Wake-word detection | TC-ResNet8 | 1 second | One shared “Hi Tenniix” wake-word model |
Both architectures consume 40-bin log-Mel features from 16 kHz mono audio. The frontend uses a 30 ms Hann window, a 10 ms frame hop, and a 512-point FFT. The corresponding feature grids are 40 × 101 for one second, 40 × 201 for two seconds, and 40 × 301 for three seconds.
TC-ResNet treats the feature bands as channels and applies residual convolutions along time. A shared implementation supports the two network depths, with global average pooling and a compact classifier. Longer command windows accommodate multi-word phrases, while the wake-word model retains its shorter input window.
The checkpoint records the architecture, input duration, preprocessing configuration, and ordered labels. Validation, inference, export, and firmware assets restore these settings so a model update does not silently change the feature shape or class mapping.
Training and real-speech evaluation
I built the training workflow around the multilingual synthetic speech data and labeled real recordings. Dataset manifests keep training, validation, and test entries together with their class order and preprocessing settings. Speaker-aware splitting assigns a recognized speaker consistently across classes, reducing speaker leakage between splits.
Streaming recognition also needs explicit non-command examples. The datasets include _silence_ for background-only windows and _unknown_ for non-command words and continuous speech. Wake-word and command datasets retain separate label spaces.
The training augmentation pipeline addresses the difference between generated speech and microphone recordings:
- SNR-controlled background noise, gain variation, and speed perturbation.
- Complete commands placed at different positions within the model window.
- Room-tone and adjacent unknown-speech context from training-only sources.
- Reverberation, microphone filtering, distance and off-axis effects, and clipping or codec artifacts.
- Partial-command negatives and conservative time/frequency masking.
Validation and test preprocessing remain deterministic. Real recordings can be added through a derived validation manifest with matching labels and duration, allowing checkpoint selection to account for real microphone speech alongside synthetic data.
Quantization and firmware export
The deployment path is PyTorch checkpoint → validated ONNX graph → ESP-DL INT8 model → ESP32-S3 firmware assets.
The ONNX exporter checks graph operators, input/output shapes, metadata, and numerical agreement with PyTorch. Quantization uses exported calibration features and produces a float-versus-INT8 comparison report. Export checks cover ESP-DL exponent compatibility and a configurable maximum accuracy drop; the default accuracy gate is one percentage point. That threshold is an export criterion, not a claim that every model has the same measured quantization loss.
The firmware asset generator packages the quantized model with its ordered labels, audio duration, Hann window, and sparse Mel-filterbank coefficients. It checks the duration-derived feature shape against the exported model when metadata is available, and the firmware validates its input shape and INT8 tensor at startup.
This makes the model and frontend a coordinated deployment package. A change to the input duration, feature settings, or labels requires regenerating the corresponding firmware assets.
Live INMP441 inference
The ESP32-S3 captures the INMP441 through I2S into a rolling PSRAM buffer. A dedicated capture task continues reading microphone samples while preprocessing and inference run. Once a complete model window is available, the inference task uses the latest audio rather than waiting for capture to restart.
The embedded path performs DC removal and gain adjustment, computes the log-Mel representation, quantizes features into the model’s INT8 input tensor, and executes the ESP-DL graph. It emits the predicted class, confidence, top-ranked alternatives, and microphone-level diagnostics over serial.
The live inference implementation targets a 100 ms start-to-start prediction interval. If processing takes longer, the next prediction starts immediately. The audio-window length and the prediction interval are separate settings: a two-second input model can process overlapping windows without collecting a fresh two-second recording for every prediction.
Latency optimization and measured results
Profiling showed that microphone-buffer copies, audio conversion, and feature handling contributed substantially to the live processing time. I optimized those stages while preserving the model and frontend contracts:
- Replaced full-window raw-audio snapshots with direct conversion from the ring buffer.
- Maintained a rolling sample sum for DC removal and used single-precision audio math.
- Used internal RAM for the normalized waveform when available, with a PSRAM fallback.
- Evaluated only the nonzero portions of the Mel filters.
- Fused feature generation and quantization directly into the ESP-DL input tensor.
- Kept microphone capture on a dedicated task and scheduled predictions by their start time.
The repository’s hardware report records the following results for the two-second, 56-label deployment on an ESP32-S3 at 240 MHz with 2 MB PSRAM:
| Measurement | Recorded result |
|---|---|
| Mean active processing before optimization | 179.93 ms |
| Mean active processing after optimization | 72.96 ms |
| Active processing reduction | 59.4%, or approximately 2.47× faster |
| Optimized neural-network inference | Approximately 20.80 ms |
| Optimized active processing P95, second capture | 73.80 ms across 288 predictions |
These timings cover conversion, feature extraction and quantization, model execution, and output reading after a complete audio window is available. They exclude audio collection time, scheduling delay, and serial printing. The first capture’s clean statistics excluded one task-watchdog interruption; the second capture contained 288 predictions. The figures describe that measured two-second configuration, rather than a latency guarantee for every language or input duration.
Outcome
The project delivered seven-language command models and one shared wake-word model with a consistent path from dataset preparation to live ESP32-S3 microphone inference. The deployment includes model-specific assets, numerical and quantization checks, continuous audio capture, and diagnostics for firmware integration.
The measured optimizations brought the documented two-second live processing path below its 100 ms scheduling target. Per-language recognition accuracy and false-trigger rates are not quantified in this project record; those are separate from the hardware processing measurements.
Technology
Python, PyTorch, torchaudio, TC-ResNet8, TC-ResNet14, ONNX, ESP-PPQ, ESP-DL, INT8 quantization, ESP-IDF, C++, FreeRTOS, ESP32-S3, INMP441, I2S, log-Mel features, and streaming audio inference.
