How accurate is Korean speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 14.3% character error rate on Korean FLEURS. For context from the same table and the same model: Italian 4.0%, Japanese 5.3%, Polish 5.4%, Russian 5.6%. Korean is roughly three times the error rate of those languages, and any page that puts it beside them without saying so is misleading you.
That figure is a character error rate, not a word error rate. The paper uses character-level scoring for Chinese, Japanese and Korean, and for Korean there is an additional reason: word spacing is applied inconsistently even by human transcribers, to the point that researchers introduced a space-normalised error rate specifically to stop spacing disagreements from corrupting Korean evaluations. So part of any published Korean number is spacing convention rather than misrecognition.
Model size matters more for Korean than for any other language on this site. On the same benchmark the compact model scores 36.1% and the small model 19.6%, against 14.3% for the largest. That is not a small gradient - it is the difference between unusable and workable. The free-tier compact models are genuinely not enough for serious Korean dictation, and we would rather say that than sell you a disappointment.
What makes Korean genuinely hard for speech models
Word spacing is inconsistent, even among people. Korean spacing rules exist but are widely varied in practice, and transcription corpora contain sequences of inconsistent spacing that produce invalid evaluations. This is why researchers proposed a space-normalised word error rate with flexible spacing rules for Korean. The practical consequence for you is that some of what looks like an error is a spacing convention the model chose differently from you.
One wrong syllable can change the meaning. Many Sino-Korean morphemes are single syllables, and phonologically similar syllables can map to different meanings or syntactic roles, so a single-character substitution can change what a sentence is asking. Research on Korean spoken question answering identifies exactly this: meaning flips hide under a low character error rate. A small number therefore understates the damage more in Korean than in a language where errors look like obvious typos.
Politeness is grammar, and it lives at the end of every sentence. Korean marks formality and deference in the verb ending. Getting that ending wrong is not a cosmetic error - it changes who the speaker is being to whom. We make no claim about handling speech levels or honorific forms, and you should check sentence endings in anything that goes to another person.
What we actually offer for Korean, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Korean we do not. There is no Korean tuning layer, no Korean cleanup pass and no Korean fine-tune. We do not normalise spacing, we do not handle Hanja, and we make no claim about honorific or speech-level correctness.
What SnailText gives you is the best open Korean speech model running entirely on your own hardware, with a global hotkey that pastes text at your cursor in any app, and audio that never leaves your machine. Given the accuracy above, the honest positioning is that this is a privacy-first way to run the same model everyone runs - not a claim that Korean recognition is solved.
If your Korean dictation needs to be production-accurate, use the largest model you can run and expect to proofread. If you are evaluating tools, the useful comparison is not our number against a competitor's marketing - it is whether a tool tells you the number at all.
One practical tip: set the dictation language to Korean explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.