SnailText
EN

Vietnamese speech to text

Vietnamese speech to text, running on your own machine

Dictate in Vietnamese into any app on Mac or Windows. The biggest risk is not tones - it is English technical terms. SnailText runs Whisper locally, so nothing is uploaded.

Download for Macand start dictating in any app
cursor.txt
Đang nói…
▍
Done

The short version

In the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 10.3% word error rate on Vietnamese FLEURS - a genuine word error rate, comparable with Czech and Swedish, because Vietnamese uses spaced Latin script with precomposed diacritics that survive standard normalisation. The interesting finding is what actually goes wrong. An error analysis of Vietnamese recognition found that foreign terms, typically English vocabulary rendered as approximating Vietnamese syllables, accounted for 2612 of 4698 substitution errors, while tone errors accounted for only 175. If you dictate Vietnamese with English technical terms mixed in, that is your real problem, not the tones. SnailText runs the model locally with no Vietnamese-specific tuning layer.

Vietnamese dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Vietnamese STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Plain VietnameseGood - around 10.3% on the read benchmarkGenerally good too
English terms mixed inThe dominant error source, and we say soSame problem; rarely disclosed
Tone diacriticsWhatever the open model producesUsually preserved, varies by vendor
Vietnamese-specific tuningNone - same open model as any local Whisper setupSome vendors tune per language, most run stock

How accurate is Vietnamese speech to text with local Whisper?

In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 10.3% word error rate on Vietnamese FLEURS. Unlike Chinese, Japanese, Korean or Thai, this is a real word error rate rather than a character one: Vietnamese is written in spaced Latin script, and its precomposed diacritics survive the standard evaluation normalisation, so the number is directly comparable with Czech, Swedish or Ukrainian.

That is clean read speech and describes a floor. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. The compact models are not enough for real Vietnamese dictation.

The more useful question for Vietnamese is not how big the number is but what makes it up, and there the published analysis is genuinely surprising.

What makes Vietnamese genuinely hard for speech models

English technical terms, not tones, dominate the errors. An error analysis of Vietnamese speech recognition on FLEURS broke substitutions down by type. The largest category by far was out-of-vocabulary foreign terms - typically English words that the system transcribes as Vietnamese syllables approximating the foreign pronunciation, so that feet comes out as phit. That accounted for 2612 of 4698 substitutions. Tonal confusion accounted for 175. If your Vietnamese is peppered with English product names and technical vocabulary, as most professional Vietnamese is, that is where your errors will come from.

Tone errors are real but rarer than expected. The same analysis documents a distinct tonal-confusion category where the reference and the output are segmentally identical and differ only in tone marking - ma becoming ma with a different tone mark, changing the word entirely. Vietnamese is tonal and the tone carries meaning, so these errors change sense rather than looking like typos, but they are a smaller share of the total than the reputation suggests.

Vowel quality and diacritic confusion is the second-largest class. Substitutions involving changes in vowel quality, vowel diacritics or vowel length while the consonants stay similar accounted for 566 errors in the same analysis - sau becoming sao, for instance. Vietnamese marks both tone and vowel quality with diacritics on nearly every syllable, so there is a lot of surface for this kind of error.

What we actually offer for Vietnamese, and what we do not

For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Vietnamese we do not. There is no Vietnamese tuning layer, no Vietnamese cleanup pass and no Vietnamese fine-tune. We do not handle English technical terms specially, which given the error breakdown above is the honest headline limitation of this page.

What SnailText does give you is the best open Vietnamese speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app. The audio is processed in RAM and never uploaded, and it works with no connection at all.

Practical advice worth more than any feature claim: if you dictate mixed Vietnamese and English regularly, expect to fix the English terms, and consider typing the short technical ones rather than speaking them. That is a workflow suggestion, not a product capability, and we would rather give you the useful version.

Set the dictation language to Vietnamese explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, so starting a passage with an English sentence can affect everything after it.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Vietnamese speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 10.3% word error rate on Vietnamese FLEURS. This is a genuine word error rate, comparable with Czech or Swedish, because Vietnamese uses spaced Latin script - unlike Chinese, Japanese, Korean or Thai, which are scored by character. It is clean read speech, so everyday dictation runs higher.

Are tones the main problem with Vietnamese dictation?

+

Surprisingly, no. A published error analysis of Vietnamese recognition found that foreign terms - mostly English words rendered as approximating Vietnamese syllables - accounted for 2612 of 4698 substitution errors, while tonal confusion accounted for only 175. Vowel and diacritic confusion accounted for another 566. If your Vietnamese includes English technical vocabulary, that is your dominant error source.

Does it handle Vietnamese mixed with English technical terms?

+

Not well, and this is the honest limitation of this page. It is the largest documented error category for Vietnamese recognition, we ship no Vietnamese tuning layer, and nothing on our side addresses it. If you dictate mixed speech regularly, expect to correct the English terms afterwards.

Do you tune SnailText specifically for Vietnamese?

+

No. There is no Vietnamese tuning layer, no Vietnamese cleanup pass and no Vietnamese fine-tune. For Vietnamese you get the same open Whisper model everyone runs, locally and privately, with the desktop plumbing done for you.

Will the tone marks come out right?

+

Usually, and Vietnamese diacritics survive standard processing better than those of several other scripts. But tonal confusion is a documented error category where the output differs from what you said only in the tone mark, which changes the word entirely rather than looking like a typo. We make no claim about our accuracy on it.

Does Vietnamese dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Vietnamese dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

·

Vietnamese speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Vietnamese, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app