A complete open-source system that runs general-purpose English speech-to-text and on-device neural text-to-speech entirely offline on ESP32-S3 and ESP32-P4 modules, with bindings for C, MicroPython and AtomVM (Erlang/Elixir). It ships a finetuned Citrinet-256 STT compressed to 4- or 8-bit (6.0 MB int4 or 9.8 MB int8 model files) plus a noisy-room variant, and tiny neural TTS voices (nano: 294k params, ~1.9 MB; heart: 2.27M params, additional ~2.4 MB for P4). STT needs about 6 MB flash and 84.5 KB internal RAM (plus a few hundred KB PSRAM during inference); TTS adds ~1.9 MB flash and ~0.5 MB working PSRAM. Models are memory-efficiently read in-place from flash (not copied into RAM). The ESP32-P4 runs faster (~0.30 s to transcribe 2 s of speech vs ~0.51 s on S3) and can host the larger heart voice; both boards are supported with drivers (ES7210/ES8311, PDM) and example hardware.
Model results and tooling are practical and specific: word-error-rate comparisons on LibriSpeech clean and 82 noisy-board recordings show tradeoffs (noisy-finetuned models reduce WER in noisy rooms at the cost of some clean-speech accuracy). The repo includes training/field-capture/datagen/finetuning pipelines, optimized 4/8-bit integer kernels for Xtensa/RISC-V vector units, build instructions (ESP-IDF v5.5.1, MicroPython v1.27.0, optional AtomVM), and examples for LoRa messaging, offline privacy-first voice control, and headless voice nodes. This targets projects needing local, private speech I/O rather than cloud accuracy or multilingual coverage.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.