Cactus releases Whistle for local speech recognition
The compact CPU model handles seven languages and clips up to 30 seconds.

Cactus has released Whistle, a 16.9 MB speech recognition model for CPUs. Its weights appeared September 30, and the company detailed the model on October 2, 2026.
The model accepts 16 kHz mono audio in clips up to 30 seconds. Supported languages are English, German, French, Spanish, Italian, Dutch and Polish. Developers can use word timestamps to navigate transcripts and keyword biasing to favor names or product terms.
A compact model for short clips
Cactus reports an 11.1 millisecond time to first token for ten seconds of audio on an Apple M4 Pro CPU, using its C++ engine with five beams. Results on other devices may differ.
The weights carry Apache 2.0, and source code is available. Whistle shares Needleโs runtime, enabling transcription and tool calls together. Longer recordings need separate handling beyond the stated clip limit.
Microphone by freestocks.org, dedicated to the public domain under CC0 1.0. Compressed for publication. This 2016 photograph illustrates audio capture and does not depict a device running Whistle.




[…] Related coverage looks at Cactus Whistle for local speech recognition. […]