Whistle, a 16.9MB Speech-to-Text Model: Nearly 9x Smaller Than Whisper Base, More Than 6x Faster to First Token, 7 Languages, Runs on Watches and in Browsers

On-device AI company Cactus Compute released Whistle, an open speech-to-text model in a single 16.9MB file, on Hugging Face under Apache 2.0; it drew strong attention on Hacker News this week. It supports English, German, French, Spanish, Italian, Dutch and Polish with automatic language detection, plus word-level timestamps, keyword biasing and silence detection. In Cactus's test on an Apple M4 Pro CPU with 10 seconds of audio, Whistle reaches its first token in 11.1 ms and decodes 1,319 tokens per second, versus 73.2 ms and 266 tokens per second for the 145.3MB Whisper base. Cactus says its word error rate beats Whisper base on LibriSpeech, Earnings-22 and others, but trails on TED-LIUM, AMI and the multilingual MLS average. Each pass handles at most 30 seconds of audio, with transcripts capped at 320 tokens.

How small

If you want offline speech transcription on a phone, a watch or in a browser, even the commonly used Whisper base is 145MB. Whistle, open-sourced at the end of September by on-device AI company Cactus Compute, is a single 16.9MB file under Apache 2.0, and it drew strong attention on Hacker News this week.

It uses an 8-block encoder and an 8-block decoder, supports seven languages (English, German, French, Spanish, Italian, Dutch and Polish) with automatic language detection, and adds word-level timestamps, per-frame speech embeddings, keyword biasing (making proper nouns easier to get right) and silence detection. Its runtime engine ships prebuilt for 17 targets, from macOS, Linux, Android, iOS and watchOS to Windows on ARM, RISC-V, browsers and WASI, and in Python it is one pip install cactus-needle away.

The speed–accuracy trade-off

Cactus's comparison on an Apple M4 Pro CPU with 10 seconds of audio:

  • Whistle (16.9MB): 11.1 ms to first token, 1,319 tokens per second;
  • Whisper base (145.3MB): 73.2 ms, 266 tokens per second;
  • Moonshine tiny v2 (41.9MB): 22.8 ms, 262 tokens per second.

Accuracy is mixed: per Cactus's charts, Whistle's word error rate is lower than Whisper base and Moonshine tiny v2 on LibriSpeech, SPGISpeech, Earnings-22 and the FLEURS average, but it trails Whisper base on TED-LIUM, the AMI meeting recordings and the multilingual MLS average. The Whisper and Moonshine numbers are their authors' published results, not re-runs in the same environment.

Where it fits

A maximum of 30 seconds of audio per pass and a 320-token transcript cap make it better suited to voice commands, short message dictation and post-wake-word recognition on watches and earbuds than to full meeting recordings; long audio has to be chunked yourself. It supports only seven European languages, so Chinese, Japanese and others are out. For developers who want voice input in an app without uploading recordings to a server, it is one of the lightest options available, but test word error rates on real recordings from your own users before shipping.

via: Cactus blog, Hugging Face model page, GitHub runtime engine