How accurate is Thai speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 11.5% on Thai FLEURS. The paper's table is captioned as word error rate, but its appendix explains that for languages written without spaces between words - Chinese, Japanese, Thai, Lao and Burmese - spaces are inserted between individual characters before scoring, which effectively measures the character error rate instead. So 11.5% for Thai and 13.3% for Czech are not the same kind of measurement and should not be ranked against each other.
Independent work on Thai evaluation reinforces how badly this can go wrong. One analysis of evaluation pipelines found that standard text normalisation renders Thai unusable through excessive spacing, and the authors excluded Thai from word-error comparisons altogether on the grounds that character error rate is the metric reported for languages where the space is not a word delimiter.
Accuracy tracks model size far more than it tracks local versus cloud. It is also fair to note that specialised Thai fine-tunes of Whisper exist, which is a signal that the community considers the base model improvable for Thai rather than finished.
What makes Thai genuinely hard for speech models
Thai is written without spaces between words, and segmentation is not solved. Research on Thai speech recognition states the problem plainly: Thai ASR still faces challenges due to the lack of spaces in Thai sentences, and Thai admits several written forms for the same spoken word - numerals, the repetition marker, borrowed words - so transcripts have to be normalised to a single form before they can even be compared. The direct answer to the obvious question is that Whisper emits Thai text but does not provide reliable word boundaries. It is a speech model, not a Thai word segmenter.
What that means in practice. You get Thai text you can read and edit, not text with dependable word boundaries. If your workflow needs segmented Thai words - for indexing, for word counts, for alignment - you will need a dedicated Thai segmenter after us. SnailText produces timing per speech segment rather than per word, so it is not a tool for word-level alignment in any language, and Thai is where that matters most.
Tone is encoded by a combination of factors, not one mark. Thai tone comes from the consonant class, the vowel length and the tone marks together, rather than from a single diacritic. Thai vowel signs also belong to the Unicode mark category, which is why naive text normalisation strips them and corrupts the text.
What we actually offer for Thai, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Thai we do not. There is no Thai tuning layer, no Thai cleanup pass, no Thai fine-tune and no word segmenter. If your workflow depends on properly segmented Thai words, this is the wrong tool and we would rather say so here than have you find out later.
What SnailText does give you is the best open Thai speech model running entirely on your own hardware, with a global hotkey that pastes text at your cursor in any app, and audio that never leaves your machine. For dictating Thai into a document or a chat window, where you read and edit the text yourself, that works fine.
Expect to review the output. Between the segmentation behaviour and the error rate, Thai is one of the languages where proofreading is part of the workflow rather than an occasional check.
One practical tip: set the dictation language to Thai explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.