How accurate is Japanese speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 5.3% character error rate on Japanese FLEURS and 9.1% on Japanese Common Voice 9. Japanese is not in the Multilingual LibriSpeech table, so FLEURS is the primary source, and Common Voice - crowd-recorded and noisier - gives the honest upper end.
The metric matters as much as the number. Those are character error rates, not word error rates. The paper scores Chinese, Japanese and Korean by character because word segmentation is not well defined in those writing systems: Japanese is written without spaces between words, so there is no word boundary for a model to get right or wrong. That means a Japanese figure of 5.3% and a Spanish figure of 3.0% are not the same kind of measurement, and lining them up as a ranking would be misleading.
Both numbers are clean read speech and describe a floor. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. SnailText's free tier covers the compact models, and the larger ones are on Pro.
What makes Japanese genuinely hard for speech models
Homophones resolved to the wrong kanji, producing output that sounds right and is wrong. Japanese has very high homophone density, and the characteristic ASR failure is a phonetic-similarity error: the model picks characters with the correct reading but the wrong meaning. Research on Japanese speech annotation documents Whisper-small emitting character combinations that read correctly aloud but are not real words. This is the most important thing to know about Japanese dictation, because the error is invisible if you check by reading the text back to yourself.
Script choice is genuinely ambiguous, and inflates every published error rate. The same word can legitimately be written in kanji, hiragana or katakana, and human transcribers disagree about which to use. A study of Japanese ASR evaluation found many cases where one reference used kanji and another used kana, and concluded that errors were more frequently due to orthographic variation than to misrecognition - which is why the authors proposed a lenient scoring scheme. Part of any Japanese error rate you read, including the ones on this page, is disagreement about spelling rather than about hearing.
No spaces, so no word boundaries. This follows from the writing system rather than from any defect in the model, and it is the cleanest explanation of why Japanese numbers are not comparable to Spanish or German ones. It is also why the metric is per-character.
What we actually offer for Japanese, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Japanese we do not. There is no Japanese tuning layer, no Japanese cleanup pass and no Japanese fine-tune. In particular we cannot offer a script preference: if you want a word in kanji rather than kana, there is no setting for that, and we do not claim to influence homophone selection.
What SnailText does give you is the best open Japanese speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app - email, docs, Slack, your editor. No Python to set up, no model plumbing, no account needed to start.
Given the homophone issue above, the honest recommendation for Japanese is to proofread names, technical terms and anything where a wrong character would change the meaning rather than look like a typo. A larger model reduces the rate but does not remove the failure mode.
One practical tip: set the dictation language to Japanese explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.