How accurate is Indonesian speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 7.1% word error rate on Indonesian FLEURS. That is better than French at 8.3% and Turkish at 8.4% in the same table, which makes Indonesian the strongest-scoring language in this group. Indonesian orthography helps here: it uses the plain 26-letter Latin alphabet with no diacritics and a highly regular sound-to-spelling mapping, so there is less for a model to get wrong on the page.
Indonesian is also the language where we have the clearest published evidence of how far a benchmark sits from daily use. A production engineering write-up measured whisper-small at 30.87% word error rate across a mixed 80.54-hour corpus, and 40.66% on clean but spontaneous informal speech. That is a smaller model on much harder material, so it is not the case that 7.1% becomes 40.66% - but the direction and the scale of the gap are documented, and pretending otherwise would be dishonest.
Treat the benchmark as a floor for careful, formal speech and expect materially worse on casual conversational Indonesian. Accuracy tracks model size far more than it tracks local versus cloud, so use the largest model you can run.
What makes Indonesian genuinely hard for speech models
Code-switching with English, and it is severe. Indonesian professional and urban speech is densely mixed with English. The measurements are stark: one study reported 2.43% character error on monolingual English, 4.10% on monolingual Indonesian, 37.57% on synthesised code-switched speech, and 91.76% on natural spontaneous Indonesian-English mixing with slight noise. The root cause is general rather than Indonesian-specific - Whisper is trained on monolingual data and handles one language per window - but Indonesian speech hits it constantly. If your dictation habitually mixes English terms into Indonesian sentences, this is the thing that will frustrate you.
The language decision is made once and applied to everything after it. Whisper detects the language from roughly the first thirty seconds and does not revisit it, so an English greeting at the start can cause everything following to be decoded under an English prior. For dictation this has a straightforward remedy, and it is advice rather than a feature: choose Indonesian explicitly instead of relying on automatic detection.
Formal and colloquial Indonesian are meaningfully different. Standard written Bahasa Indonesia and everyday Jakartan speech diverge in vocabulary, affixes and contractions. Benchmarks are built from the formal register. We make no claim about colloquial or slang coverage.
What we actually offer for Indonesian, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Indonesian we do not. There is no Indonesian tuning layer, no Indonesian cleanup pass and no Indonesian fine-tune. We do not handle code-switching, we do not claim colloquial register support, and we do not support Indonesia's regional languages - Indonesian here means Bahasa Indonesia and nothing else. Javanese, for instance, is a separate language with far worse published results.
What SnailText does give you is the best open Indonesian speech model running entirely on your own hardware, wired into a global hotkey that pastes at your cursor in any app. Indonesian is comparatively poorly served by good offline tools, so for many people the meaningful comparison is not against a better Indonesian product but against having nothing that works without a connection.
The audio is processed in RAM and never uploaded, so the privacy guarantee is architectural, and it works with no connection at all.
The single most useful thing you can do for accuracy: select Indonesian explicitly rather than leaving the language on automatic, and keep English terms to a minimum in the same sentence where you can.