Whistle: Speech To Text In 16.9 MB
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on networking and server gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run locally on a CPU. The company reports support for seven languages, word-level timestamps and speech embeddings, and says its benchmark results vary by dataset against Whisper base and Moonshine tiny v2.

Cactus Compute announced Whistle on October 2, describing it as a 16.9 MB speech-recognition model that runs on a CPU without external dependencies. The company says it transcribes speech in seven languages on the device, a design aimed at phones, wearables, robots, smart-home equipment, vehicles and microcontrollers.

According to Cactus Compute, Whistle accepts 16 kHz mono audio clips up to 30 seconds and supports English, German, French, Spanish, Italian, Dutch and Polish. It detects the language unless the user specifies one. The company says its browser demonstration downloads the model on the first use and processes audio locally, so audio does not leave the device in that demo.

The model provides word-level timestamps, including start and end times and a probability value, as well as speech embeddings: encoder output organized as one row per 80-millisecond frame. The company says its engine can also apply keyword biasing to phrases supplied by the user and return an empty transcript for audio below a loudness threshold.

Whistle uses the same C++ engine and some of the same model blocks as Cactus Compute’s Needle model. The company says its decoder can be configured to run at different depths, while the audio encoder always runs all eight blocks. Its published description specifies five-beam decoding and caps transcripts at 320 tokens.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact, on-device speech-to-text model that runs in the same C++ engine as its Needle model.

Local Transcription on Small Devices

A 16.9 MB model that can run locally could make speech recognition more practical on devices with limited storage or unreliable internet access. Local processing can also reduce the need to send recorded speech to a remote service, although actual privacy protections depend on how a product uses and stores audio beyond the model itself.

Cactus Compute also presents Whistle as a component for developers building speech-driven systems. Its shared engine with Needle and its speech-embedding output may let applications combine transcription with other on-device model functions. The company’s description says one binary can turn an audio clip into tool calls, but that is a product capability claim; the report does not provide an independent demonstration of performance across devices or applications.

The announced size and speed figures are relevant to deployment, but they do not by themselves establish recognition quality. Accuracy can vary by language, accent, recording conditions and vocabulary. The company’s own results show different systems leading on different evaluation datasets.

Amazon

on-device speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Was Evaluated

Cactus Compute compares Whistle with Whisper base and Moonshine tiny v2. Its report says Whistle has lower word error rates on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. It also says Whisper base performs better on TED-LIUM, AMI and the MLS average. The report warns that some comparisons are unavailable because a model’s authors did not publish a result for that test; it also says Whisper’s AMI result uses a different subset from the AMI results for the other systems.

For a separate speed comparison, the company reports tests using ten seconds of audio on an Apple M4 Pro CPU, with each system running on its official runtime and default settings. It reports 11.1 milliseconds to first token for Whistle, compared with 73.2 ms for Whisper base and 22.8 ms for Moonshine tiny v2. The report also lists decode rates of 1,319, 266 and 262 tokens per second, respectively, and model sizes of 16.9 MB, 145.3 MB and 41.9 MB.

Those figures come from Cactus Compute’s report, not an independent evaluation. Its methodology defines time to first token as the time from audio input to the first token and excludes the encoder from the decode-rate calculation. The company also says Whisper pads each input to 30 seconds, while Whistle’s reported time to first token varies with clip length.

““Speak, and Whistle transcribes it on your device.””

— Cactus Compute

Amazon

portable speech to text device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests Still Needed

The release report does not establish how Whistle performs across a broad range of phones, embedded processors or other CPU hardware. Its speed figures are from one named test platform, and no independent benchmark results are provided in the supplied material. Performance on different devices may not match the Apple M4 Pro measurements.

The report gives accuracy comparisons for several datasets, but some results are missing, and the company notes that the AMI comparison uses different subsets for Whisper and the other models. The supplied material does not provide enough detail to independently verify every result or determine how accuracy changes across all seven supported languages and real-world noise conditions.

It is also not clear from the announcement what license applies to the model, how the model was trained, or whether the browser demo’s local-processing behavior applies to every deployment. The company describes Whistle as an open model, but the release material provided here does not specify the full terms for use or redistribution.

Amazon

multilingual speech recognition app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Availability and Testing Ahead

Cactus Compute’s release and browser demonstration make Whistle available for users to try, with the company saying the first use downloads the 16.9 MB model. The announcement identifies potential uses in mobile, wearable, automotive, robotics and smart-home products, but does not name a shipping product or deployment schedule.

The next useful evidence will come from testing the model on additional hardware and across its supported languages, alongside clearer information about its license and training data. Until then, the published size, speed and accuracy figures should be read as company-reported results, rather than a settled comparison across devices and use cases.

Amazon

privacy-focused speech transcription device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is a speech-recognition model from Cactus Compute. The company says it runs on a CPU and can transcribe clips locally.

Which languages does it support?

Cactus Compute lists English, German, French, Spanish, Italian, Dutch and Polish. The model detects the language unless the user specifies one.

Does Whistle send audio to a server?

The company says audio in its browser demo stays on the device after the model downloads. The supplied report does not establish how every third-party deployment handles audio.

How does it compare with Whisper base?

Cactus Compute reports lower word error rates for Whistle on some datasets and better results for Whisper base on others. Its speed and size comparisons are company-reported tests on an Apple M4 Pro CPU, not independent evaluations across devices.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Has Shopify Dropped React Native?

Shopify says AI coding tools have changed the trade-offs around mobile development. Its Shop app rewrite to native took 12 weeks, according to The Pragmatic Engineer.

Rails World 2026 Opening Keynote [Video]

Search and coverage interest in the Rails World 2026 opening keynote video is spiking. What is confirmed, what drives the interest, and what remains unknown.

Googlebooks Might Be The Real Deal

Online interest in Google Books is spiking, but no new announcement has been confirmed. Here is what is known and what remains unclear.

Digital Deli, 1984 Book By Early PC Hackers And Enthusiasts

A 1984 book by early PC hackers and enthusiasts titled ‘Digital Deli’ has recently gained renewed attention, revealing insights into early computer hacking culture.