Whistle is a 16.9 MB speech-recognition model that runs entirely on CPU inside the same C++ engine as Needle, targeting mobiles, wearables, robots, smart-home devices and microcontrollers. It transcribes up to 30 seconds of 16 kHz mono audio in seven languages, emits word-level timestamps with probabilities, and provides per-frame speech embeddings. The frontend uses 80-bin log-mel features and a 3-stage convolutional stem to produce 375 frames (one per 80 ms). The encoder has eight Simple Attention blocks with a Monarch Hadamard MLP; the decoder is a laddered stack of eight Simple Attention blocks augmented by a per-layer gated cross-attention to the encoder. Decoding uses 5-beam search, Aho-Corasick keyword biasing, a 8,192-piece vocabulary plus seven language tokens, and an audio-depth option that selects decoder depth at load time while always running the full encoder.
Benchmarks on an Apple M4 Pro show very low latency and competitive accuracy: Whistle reaches first token in ~11 ms and decodes ~1,319 tokens/s, while fitting in 16.9 MB (vs Whisper base at 145.3 MB). It outperforms Whisper on LibriSpeech, SPGISpeech, Earnings-22 and FLEURS averages, while Whisper leads on TED-LIUM, AMI and MLS. The model integrates into Needle with three load modes (speech-only, text-only, or combined tool-call flow), supports CLI and Python APIs, silence detection, word timestamps, and multiplatform prebuilt binaries; weights and source are published on Hugging Face and GitHub.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.