How accurate is Vietnamese speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 10.3% word error rate on Vietnamese FLEURS. Unlike Chinese, Japanese, Korean or Thai, this is a real word error rate rather than a character one: Vietnamese is written in spaced Latin script, and its precomposed diacritics survive the standard evaluation normalisation, so the number is directly comparable with Czech, Swedish or Ukrainian.
That is clean read speech and describes a floor. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. The compact models are not enough for real Vietnamese dictation.
The more useful question for Vietnamese is not how big the number is but what makes it up, and there the published analysis is genuinely surprising.
What makes Vietnamese genuinely hard for speech models
English technical terms, not tones, dominate the errors. An error analysis of Vietnamese speech recognition on FLEURS broke substitutions down by type. The largest category by far was out-of-vocabulary foreign terms - typically English words that the system transcribes as Vietnamese syllables approximating the foreign pronunciation, so that feet comes out as phit. That accounted for 2612 of 4698 substitutions. Tonal confusion accounted for 175. If your Vietnamese is peppered with English product names and technical vocabulary, as most professional Vietnamese is, that is where your errors will come from.
Tone errors are real but rarer than expected. The same analysis documents a distinct tonal-confusion category where the reference and the output are segmentally identical and differ only in tone marking - ma becoming ma with a different tone mark, changing the word entirely. Vietnamese is tonal and the tone carries meaning, so these errors change sense rather than looking like typos, but they are a smaller share of the total than the reputation suggests.
Vowel quality and diacritic confusion is the second-largest class. Substitutions involving changes in vowel quality, vowel diacritics or vowel length while the consonants stay similar accounted for 566 errors in the same analysis - sau becoming sao, for instance. Vietnamese marks both tone and vowel quality with diacritics on nearly every syllable, so there is a lot of surface for this kind of error.
What we actually offer for Vietnamese, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Vietnamese we do not. There is no Vietnamese tuning layer, no Vietnamese cleanup pass and no Vietnamese fine-tune. We do not handle English technical terms specially, which given the error breakdown above is the honest headline limitation of this page.
What SnailText does give you is the best open Vietnamese speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app. The audio is processed in RAM and never uploaded, and it works with no connection at all.
Practical advice worth more than any feature claim: if you dictate mixed Vietnamese and English regularly, expect to fix the English terms, and consider typing the short technical ones rather than speaking them. That is a workflow suggestion, not a product capability, and we would rather give you the useful version.
Set the dictation language to Vietnamese explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, so starting a passage with an English sentence can affect everything after it.