How accurate is Russian speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 5.6% word error rate on Russian FLEURS and 7.1% on Russian Common Voice 9. Russian is not in the Multilingual LibriSpeech table, so FLEURS is the primary source here; Common Voice is crowd-recorded and noisier, which is why it sits higher and why quoting both gives an honest range.
Both are read speech recorded in decent conditions, so they describe a floor rather than what you should expect in a normal room. Everyday dictation runs above both numbers. As with every language, accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way, and Russian reaches production quality on the larger models.
We do not publish an accuracy figure of our own. SnailText's Russian prompt is measured internally against a single reference recording, which is useful for our development and is not a benchmark, so it does not belong on this page as a claim.
What makes Russian genuinely hard for speech models
Unstressed vowels collapse, and stress is unpredictable and unwritten. Every unstressed Russian syllable reduces, and all but one vowel drift toward a neutral schwa - unstressed o is pronounced as an a, unstressed e centralises toward i. Russian stress can fall on any syllable and is not marked in ordinary writing. The result is that words differing only in stress, or only in an unstressed vowel, are close to acoustically identical, and the model has nothing but context to separate them.
Rich inflection produces a very high out-of-vocabulary rate. Published work on large-vocabulary Russian recognition puts it plainly: Russian is a highly inflectional language with rich morphology, which leads to high out-of-vocabulary word rates. One dictionary word yields dozens of surface forms through prefixes, suffixes and endings, and Russian word order is free enough that surrounding-word context helps less than it would in English.
English technical terms get transliterated into Cyrillic. Say Python or git push inside a Russian sentence and an untuned model tends to write it out in Cyrillic, because that is what its Russian training transcripts do. This is the specific problem SnailText's Russian prompt targets, and it is why that prompt is built from Latin technical tokens - Docker, Kubernetes, GitHub, TypeScript, Postgres - that bias the model toward emitting Latin script where a developer expects it.
Why local, and why Russian-tuned
Most dictation tools ship stock Whisper and inherit its Russian behaviour unchanged, including the Cyrillic-transliteration habit above. The tools that do tune per language are cloud services, which means your audio leaves your machine. SnailText is the other combination: the same best-in-class local Whisper model everyone runs, plus a Russian-tuned prompt and a Russian-aware cleanup pass, running entirely on your own hardware.
To be precise about what that is and is not - it is a prompt and a post-processing pass, not a fine-tuned model, and we make no claim that it lowers the word error rate, because we have no published benchmark of our own. What it does is bias the output toward the conventions a Russian-speaking developer actually wants, which is a different and more honest claim than a number.
That matters when what you dictate is not casual. Client notes, technical writing, internal company text - the audio is processed in RAM and never uploaded, so the privacy guarantee is the architecture rather than a policy you have to take on trust. It also works on a plane or with no wifi at all.
If you want the whole interface in Russian too, that is a separate thing from dictating in Russian; the dictation engine is the same either way.