TL;DR
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to run locally on a CPU and support seven languages. The company reports low latency and competitive results on several speech benchmarks, but those figures come from its own comparisons and do not establish performance across devices or real-world conditions.
Cactus Compute announced Whistle on October 2, describing it as a 16.9 MB speech-recognition model that runs on a CPU without external dependencies and transcribes speech on the device. The model supports English, German, French, Spanish, Italian, Dutch and Polish, positioning it for phones, wearables, vehicles and other devices where sending audio to a server may be undesirable.
The company says Whistle processes 16 kHz mono audio clips of up to 30 seconds, detects the spoken language unless the user specifies one, and can return word-level start and end times with probability values. It also exposes speech embeddings: encoder outputs organized as one row per 80-millisecond audio frame. Cactus says audio stays on the device in its browser demonstration; the first use downloads the 16.9 MB model.
Whistle uses the same C++ CPU engine and shared model components as Needle, another Cactus model. Its published design describes an eight-block encoder and an eight-block decoder, with gated cross-attention connecting them. The decoder uses five-beam search, supports keyword biasing, and caps transcripts at 320 text tokens. Users can choose a decoder depth at load time; the encoder still runs all eight blocks.
In company-reported tests on an Apple M4 Pro CPU using ten seconds of audio, Whistle reached its first token in 11.1 milliseconds and decoded at 1,319 tokens per second. Cactus compared those measurements with Whisper base and Moonshine tiny v2 on their official runtimes and default settings. It also reports lower word error rates than the compared models on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average; Whisper base scored better on TED-LIUM, AMI and the MLS average.
Local Transcription on Smaller Devices
A 16.9 MB model that runs locally could make speech recognition more practical on devices with limited storage, intermittent connectivity or constraints on sending recordings to remote services. Cactus names mobile devices, wearables, robots, smart-home products, automotive systems and microcontrollers as intended settings. Local processing can also reduce the need to transmit audio, although the announcement does not provide an independent privacy or security evaluation.
The shared engine with Needle is another part of the pitch: Cactus says the two models can load from the same container and use the same quantization, enabling a speech-to-text stage to connect with a model that handles tool-oriented tasks. That may simplify an on-device voice interface, but the announcement does not report an end-to-end system test or demonstrate how reliably transcribed speech becomes tool calls.
The benchmark results suggest a trade-off rather than universal superiority. Whistle is reported to lead on several datasets, while Whisper base leads on others. For developers, the results are a reason to test the model against their own languages, accents, microphones and background noise—not a guarantee of accuracy in every deployment.
on-device speech recognition software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Whistle Processes Audio
Whistle converts audio into 80 log-mel features using 25-millisecond windows and a 10-millisecond hop, then reduces the frame count with a convolutional stem. Cactus says a 30-second clip becomes 375 encoder frames, each representing about 80 milliseconds. The encoder produces representations that the decoder reads while generating text.
The release draws on the architecture of Needle, but adds speech-specific cross-attention in each decoder layer. Cactus says key and value projections for the audio are computed once per clip and reused during beam search. A silence check also runs before decoding; according to the company, audio below its threshold returns an empty transcript without starting the search.
The company’s speed comparison uses different official runtimes: Whistle’s C++ engine at five beams, OpenAI Whisper, and Moonshine Voice in non-streaming mode over the full clip. Cactus notes that Whisper pads every input to 30 seconds, while Whistle’s reported time to first token varies with clip length: 5.9 milliseconds for five seconds, 11.1 for ten seconds and 36.3 for 30 seconds on the stated test system.
““Speak, and Whistle transcribes it on your device.””
— Cactus Compute
As an affiliate, we earn on qualifying purchases.
Independent Testing Is Still Needed
The figures cited here come from Cactus Compute’s own report. The supplied material does not identify an independent benchmark, provide confidence intervals, or establish whether the comparisons were repeated across multiple machines. Results from the Apple M4 Pro may not predict speed, battery use or memory demands on phones, microcontrollers or other CPUs.
Accuracy also depends on the dataset and evaluation setup. Cactus says missing chart entries mean the model authors did not publish results for those tests, and it flags that Whisper’s AMI number uses the AMI-IHM subset rather than the AMI subset used for the other models. The material does not give enough detail here to assess every scoring choice or reproduce the reported comparisons.
The announcement does not specify Whistle’s license, model-weight availability, supported hardware beyond the CPU description, or the exact silence threshold. It also leaves unanswered how the model performs with noisy recordings, varied accents, overlapping speakers or clips longer than 30 seconds. Those details matter before developers can judge suitability for a particular product.
As an affiliate, we earn on qualifying purchases.
Availability and Deployment Details
Cactus has published an interactive browser demonstration and says its first use downloads the 16.9 MB model. The company describes the release as intended for a range of edge devices, but the supplied announcement does not set out a broader product schedule, a hardware compatibility list or a formal support plan.
Developers evaluating Whistle will need to check the model’s licensing and deployment terms, measure performance on their target CPUs, and test accuracy using representative audio. Further information from Cactus on those points—and independent replication of the benchmark results—would clarify how broadly the reported size and speed advantages hold.
multilingual speech transcription app
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Whistle?
Whistle is a speech-recognition model released by Cactus Compute. The company says it runs on a CPU and can transcribe audio locally.
Which languages does it support?
Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. The model detects the language unless the user names it.
How large are the model and supported clips?
The release describes one 16.9 MB model file that handles 16 kHz mono audio clips of up to 30 seconds in one pass.
Are the reported speed and accuracy results independently verified?
The supplied results are reported by Cactus Compute. The material does not cite independent testing, so performance should be checked on the hardware and audio conditions relevant to a deployment.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
