SnailText
EN

Russian transcription

Russian transcription, running on your own machine

Dictate in Russian into any app on Mac or Windows. SnailText runs Whisper locally with a Russian-tuned prompt that keeps English technical terms in Latin script instead of transliterating them into Cyrillic. Nothing is uploaded.

Download for Macand start dictating in any app
cursor.txt
Говорю…
▍
Done

The short version

Russian is a strong Whisper language: in the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 5.6% word error rate on Russian FLEURS and 7.1% on Common Voice 9 - the same model on a cleaner and a noisier corpus. Both are read speech, so real dictation sits above them. Russian is also the one language on this site where SnailText adds a genuine per-language layer: a Russian-specific Whisper prompt built from Latin technical tokens, plus a Russian-aware post-processing pass. That targets the documented failure where an untuned model writes a spoken English term like Python as the Cyrillic transliteration instead. Everything runs locally on Mac and Windows, with no cloud and no account to start.

Russian dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Russian STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Latin tech terms in Cyrillic textRussian prompt biases toward Latin scriptStock Whisper transliterates them into Cyrillic
Russian-specific tuningRussian prompt + Russian post-processingMost run stock Whisper with no per-language layer
Dictate into any appGlobal hotkey, pastes at your cursorUsually a web app you copy out of
Account / costNo account to startAccount + per-minute or per-seat billing

How accurate is Russian speech to text with local Whisper?

In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 5.6% word error rate on Russian FLEURS and 7.1% on Russian Common Voice 9. Russian is not in the Multilingual LibriSpeech table, so FLEURS is the primary source here; Common Voice is crowd-recorded and noisier, which is why it sits higher and why quoting both gives an honest range.

Both are read speech recorded in decent conditions, so they describe a floor rather than what you should expect in a normal room. Everyday dictation runs above both numbers. As with every language, accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way, and Russian reaches production quality on the larger models.

We do not publish an accuracy figure of our own. SnailText's Russian prompt is measured internally against a single reference recording, which is useful for our development and is not a benchmark, so it does not belong on this page as a claim.

What makes Russian genuinely hard for speech models

Unstressed vowels collapse, and stress is unpredictable and unwritten. Every unstressed Russian syllable reduces, and all but one vowel drift toward a neutral schwa - unstressed o is pronounced as an a, unstressed e centralises toward i. Russian stress can fall on any syllable and is not marked in ordinary writing. The result is that words differing only in stress, or only in an unstressed vowel, are close to acoustically identical, and the model has nothing but context to separate them.

Rich inflection produces a very high out-of-vocabulary rate. Published work on large-vocabulary Russian recognition puts it plainly: Russian is a highly inflectional language with rich morphology, which leads to high out-of-vocabulary word rates. One dictionary word yields dozens of surface forms through prefixes, suffixes and endings, and Russian word order is free enough that surrounding-word context helps less than it would in English.

English technical terms get transliterated into Cyrillic. Say Python or git push inside a Russian sentence and an untuned model tends to write it out in Cyrillic, because that is what its Russian training transcripts do. This is the specific problem SnailText's Russian prompt targets, and it is why that prompt is built from Latin technical tokens - Docker, Kubernetes, GitHub, TypeScript, Postgres - that bias the model toward emitting Latin script where a developer expects it.

Why local, and why Russian-tuned

Most dictation tools ship stock Whisper and inherit its Russian behaviour unchanged, including the Cyrillic-transliteration habit above. The tools that do tune per language are cloud services, which means your audio leaves your machine. SnailText is the other combination: the same best-in-class local Whisper model everyone runs, plus a Russian-tuned prompt and a Russian-aware cleanup pass, running entirely on your own hardware.

To be precise about what that is and is not - it is a prompt and a post-processing pass, not a fine-tuned model, and we make no claim that it lowers the word error rate, because we have no published benchmark of our own. What it does is bias the output toward the conventions a Russian-speaking developer actually wants, which is a different and more honest claim than a number.

That matters when what you dictate is not casual. Client notes, technical writing, internal company text - the audio is processed in RAM and never uploaded, so the privacy guarantee is the architecture rather than a policy you have to take on trust. It also works on a plane or with no wifi at all.

If you want the whole interface in Russian too, that is a separate thing from dictating in Russian; the dictation engine is the same either way.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Russian speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 5.6% word error rate on Russian FLEURS and 7.1% on Russian Common Voice 9 - the same model on a clean read corpus and a noisier crowd-recorded one. Russian is not in the Multilingual LibriSpeech table, so FLEURS is the primary source. Both are read speech, so everyday dictation in a normal room runs above both. Accuracy otherwise depends mostly on model size.

Do you tune SnailText specifically for Russian?

+

Yes, and Russian is one of only a handful of languages where we do. SnailText adds a Russian-specific Whisper prompt built from Latin technical tokens, plus a Russian-aware post-processing pass. To be precise: that is a prompt and a cleanup pass, not a fine-tuned model, and we do not claim it lowers the error rate, because we have no published benchmark of our own. What it does is bias the output toward Russian conventions - in particular keeping English technical terms in Latin script.

Will it write Python as a Cyrillic word?

+

That is exactly the problem the Russian prompt targets. An untuned Whisper tends to transliterate English technical terms into Cyrillic when they appear inside a Russian sentence, because its Russian training transcripts do that. SnailText's Russian prompt is built from Latin technical tokens - Docker, Kubernetes, GitHub, TypeScript, Postgres - to bias the model toward keeping them in Latin script. It is a bias, not a guarantee, so check terms that matter.

Does Russian dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Russian dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

Does it handle the yo and ye distinction, or stressed homographs?

+

We make no claim about either, because we have not measured them. In ordinary Russian text the yo letter is routinely written as plain ye, so both spellings appear in training data and either can come out. In a small set of pairs that distinction carries meaning. Stress is not marked in normal Russian orthography at all, so homographs that differ only by stress are identical in text and the model has only context to go on.

Is it really different from just running Whisper myself?

+

For Russian, yes, more than for most languages. It is the same underlying Whisper model, but with a Russian-specific prompt and a Russian-aware post-processing pass on top, plus the desktop plumbing: a global hotkey, paste-at-cursor into any app, model management, and a free tier. You get the Russian-tuning layer already in place without wiring up Python.

·

Russian speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Russian, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app