hn.today

Whistle: Speech to Text in 16.9 MB

cactuscompute.com293 points73 comments
Screenshot of Whistle: Speech to Text in 16.9 MB

Whistle is a 16.9 MB speech-recognition model that runs entirely on CPU inside the same C++ engine as Needle, targeting mobiles, wearables, robots, smart-home devices and microcontrollers. It transcribes up to 30 seconds of 16 kHz mono audio in seven languages, emits word-level timestamps with probabilities, and provides per-frame speech embeddings. The frontend uses 80-bin log-mel features and a 3-stage convolutional stem to produce 375 frames (one per 80 ms). The encoder has eight Simple Attention blocks with a Monarch Hadamard MLP; the decoder is a laddered stack of eight Simple Attention blocks augmented by a per-layer gated cross-attention to the encoder. Decoding uses 5-beam search, Aho-Corasick keyword biasing, a 8,192-piece vocabulary plus seven language tokens, and an audio-depth option that selects decoder depth at load time while always running the full encoder.

Benchmarks on an Apple M4 Pro show very low latency and competitive accuracy: Whistle reaches first token in ~11 ms and decodes ~1,319 tokens/s, while fitting in 16.9 MB (vs Whisper base at 145.3 MB). It outperforms Whisper on LibriSpeech, SPGISpeech, Earnings-22 and FLEURS averages, while Whisper leads on TED-LIUM, AMI and MLS. The model integrates into Needle with three load modes (speech-only, text-only, or combined tool-call flow), supports CLI and Python APIs, silence detection, word timestamps, and multiplatform prebuilt binaries; weights and source are published on Hugging Face and GitHub.

Read on cactuscompute.com73 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.