docs/moonshine-vs-whisper.md
TL;DR - When you're working with live speech.
| Model | WER | # Parameters | MacBook Pro | Linux x86 | R. Pi 5 | Pixel 10a | iPad (A16) |
|---|---|---|---|---|---|---|---|
| Moonshine Medium Streaming | 6.65% | 245 million | 59ms | 269ms | 802ms | 420ms | 174ms |
| Whisper Large v3 | 7.44% | 1.5 billion | 11,286ms | 16,919ms | N/A | — | — |
| Moonshine Small Streaming | 7.84% | 123 million | 38ms | 165ms | 527ms | 234ms | 95ms |
| Whisper Small | 8.59% | 244 million | 1940ms | 3,425ms | 10,397ms | — | — |
| Moonshine Tiny Streaming | 12.00% | 34 million | 18ms | 69ms | 237ms | 92ms | 37ms |
| Whisper Tiny | 12.81% | 39 million | 277ms | 1,141ms | 5,863ms | — | — |
See benchmarks for how these numbers were measured.
OpenAI's release of their Whisper family of models was a massive step forward for open-source speech to text. They offered a range of sizes, allowing developers to trade off compute and storage space against accuracy to fit their applications. Their biggest models, like Large v3, also gave accuracy scores that were higher than anything available outside of large tech companies like Google or Apple. At Moonshine we were early and enthusiastic adopters of Whisper, and we still remain big fans of the models and the great frameworks like FasterWhisper and others that have been built around them.
However, as we built applications that needed a live voice interface we found we needed features that weren't available through Whisper:
82 languages are listed, but only 33 have sub-20% WER (what we consider usable). For the Base model size commonly used on edge devices, only 5 languages are under 20% WER. Asian languages like Korean and Japanese stand out as the native tongue of large markets with a lot of tech innovation, but Whisper doesn't offer good enough accuracy to use in most applications The proprietary in-house versions of Whisper that are available through OpenAI's cloud API seem to offer better accuracy, but aren't available as open models.
All these limitations drove us to create our own family of models that better meet the needs of live voice interfaces. It took us some time since the combined size of the open speech datasets available is tiny compared to the amount of web-derived text data, but after extensive data-gathering work, we were able to release the first generation of Moonshine models. These removed the fixed-input window limitation along with some other architectural improvements, and gave significantly lower latency than Whisper in live speech applications, often running 5x faster or more.
However we kept encountering applications that needed even lower latencies on even more constrained platforms. We also wanted to offer higher accuracy than the Base-equivalent that was the top end of the initial models. That led us to this second generation of Moonshine models, which offer:
Hopefully this gives you a good idea of how Moonshine compares to Whisper. If you're working with GPUs in the cloud on data in bulk where throughput is most important then Whisper (or Nvidia alternatives like Parakeet) offer advantages like batch processing, but we believe we can't be beat for live speech. We've built the framework and models we wished we'd had when we first started building applications with voice interfaces, so if you're working with live voice inputs, give Moonshine a try.