SnailText
EN

Japanese transcription

Japanese transcription, running on your own machine

Dictate in Japanese into any app on Mac or Windows. SnailText runs Whisper locally - the same open model everyone else runs - so your audio never leaves the machine. Press a hotkey, speak Japanese, and the text lands at your cursor.

Download for Macand start dictating in any app
cursor.txt
音声入力中…
▍
Done

The short version

Japanese is a strong Whisper language, but its accuracy number is measured differently from European languages and is not comparable to them. In the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 5.3% character error rate on Japanese FLEURS and 9.1% on Common Voice 9. That is character error rate, not word error rate: the paper scores Japanese by character because Japanese is written without spaces, so there are no word boundaries to be right or wrong about. Both figures are clean read speech, so real dictation runs above them. SnailText runs the same open model locally on Mac and Windows, with no cloud and no account to start. We do not ship a Japanese-specific tuning layer, and this page does not pretend otherwise.

Japanese dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Japanese STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Kanji homophone choiceWhatever the open model producesSame underlying difficulty for every tool
Script choice (kanji / hiragana / katakana)Not controllable - no script optionRarely controllable either
Japanese-specific tuningNone - same open model as any local Whisper setupSome vendors tune per language, most run stock
Account / costNo account to startAccount + per-minute or per-seat billing

How accurate is Japanese speech to text with local Whisper?

In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 5.3% character error rate on Japanese FLEURS and 9.1% on Japanese Common Voice 9. Japanese is not in the Multilingual LibriSpeech table, so FLEURS is the primary source, and Common Voice - crowd-recorded and noisier - gives the honest upper end.

The metric matters as much as the number. Those are character error rates, not word error rates. The paper scores Chinese, Japanese and Korean by character because word segmentation is not well defined in those writing systems: Japanese is written without spaces between words, so there is no word boundary for a model to get right or wrong. That means a Japanese figure of 5.3% and a Spanish figure of 3.0% are not the same kind of measurement, and lining them up as a ranking would be misleading.

Both numbers are clean read speech and describe a floor. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. SnailText's free tier covers the compact models, and the larger ones are on Pro.

What makes Japanese genuinely hard for speech models

Homophones resolved to the wrong kanji, producing output that sounds right and is wrong. Japanese has very high homophone density, and the characteristic ASR failure is a phonetic-similarity error: the model picks characters with the correct reading but the wrong meaning. Research on Japanese speech annotation documents Whisper-small emitting character combinations that read correctly aloud but are not real words. This is the most important thing to know about Japanese dictation, because the error is invisible if you check by reading the text back to yourself.

Script choice is genuinely ambiguous, and inflates every published error rate. The same word can legitimately be written in kanji, hiragana or katakana, and human transcribers disagree about which to use. A study of Japanese ASR evaluation found many cases where one reference used kanji and another used kana, and concluded that errors were more frequently due to orthographic variation than to misrecognition - which is why the authors proposed a lenient scoring scheme. Part of any Japanese error rate you read, including the ones on this page, is disagreement about spelling rather than about hearing.

No spaces, so no word boundaries. This follows from the writing system rather than from any defect in the model, and it is the cleanest explanation of why Japanese numbers are not comparable to Spanish or German ones. It is also why the metric is per-character.

What we actually offer for Japanese, and what we do not

For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Japanese we do not. There is no Japanese tuning layer, no Japanese cleanup pass and no Japanese fine-tune. In particular we cannot offer a script preference: if you want a word in kanji rather than kana, there is no setting for that, and we do not claim to influence homophone selection.

What SnailText does give you is the best open Japanese speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app - email, docs, Slack, your editor. No Python to set up, no model plumbing, no account needed to start.

Given the homophone issue above, the honest recommendation for Japanese is to proofread names, technical terms and anything where a wrong character would change the meaning rather than look like a typo. A larger model reduces the rate but does not remove the failure mode.

One practical tip: set the dictation language to Japanese explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Japanese speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 5.3% character error rate on Japanese FLEURS and 9.1% on Common Voice 9. Note that this is character error rate, not word error rate - the paper measures Japanese by character because Japanese is written without spaces, so word boundaries are undefined. That makes these numbers non-comparable with the word error rates quoted for European languages. Both are clean read speech, so everyday dictation runs higher.

Why is Japanese measured in character error rate instead of word error rate?

+

Because Japanese is written without spaces between words, so there is no agreed word boundary to score against. The Whisper paper uses character error rate for Chinese, Japanese and Korean for exactly this reason. It is a property of the writing system and the metric, not a defect in the model - but it does mean you cannot line a Japanese figure up against a Spanish one and call it a ranking.

Do you tune SnailText specifically for Japanese?

+

No. SnailText ships language-specific prompts for Spanish, German, French, Portuguese and Dutch, but there is no Japanese tuning layer, no Japanese cleanup pass and no Japanese fine-tune. For Japanese you get the same open Whisper model everyone runs, running locally and privately, with the desktop plumbing done for you.

Will it pick the right kanji?

+

Usually, but this is the honest weak spot of Japanese dictation on any tool. Japanese has very high homophone density, and the characteristic error is a character with the correct reading and the wrong meaning. Research on Japanese speech annotation has documented Whisper producing readings that are phonetically right but are not real words. The error is invisible when you read the text aloud, so proofread names and technical terms. We do not claim to influence homophone selection.

Can I choose whether a word comes out in kanji, hiragana or katakana?

+

No, there is no script setting. The same word can legitimately be written three ways in Japanese, and the model picks one. This ambiguity is real enough that researchers studying Japanese ASR evaluation found errors were more often caused by orthographic variation than by misrecognition, and proposed a more lenient scoring scheme to account for it.

Does Japanese dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Japanese dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

·

Japanese speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Japanese, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app