How accurate is Finnish speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 9.7% word error rate on Finnish FLEURS. That places Finnish close to Norwegian at 9.5% and comfortably ahead of Czech at 13.3% and Danish at 13.8% in the same table. An independent Finnish team reported 10.20% on the same benchmark for their own large-v2 fine-tune, which corroborates the order of magnitude.
That is clean read speech and describes a floor; everyday dictation runs higher. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way.
There is a measurement caveat specific to agglutinative languages. Word error rate counts whole words, and a Finnish word form can carry as much information as an English phrase, so one wrong suffix inside a long word costs a full word error. Finnish numbers therefore look worse relative to English than the underlying accuracy warrants.
What makes Finnish genuinely hard for speech models
Agglutination produces a combinatorial explosion of word forms. Finnish builds words by concatenating suffixes onto a stem, and combined with compounding and inflection this yields millions of distinct yet genuinely frequent word forms. Classical dictionary-based language modelling does not work on Finnish for precisely this reason, which is why Finnish speech research pioneered morph-based models.
Out-of-vocabulary is effectively unavoidable with whole-word units. The measurements here are striking: Finnish research found that moving from word units to morph-like sub-word units reduced the out-of-vocabulary rate from 20% to zero and cut word error rate from 56% to 32%. Whisper uses byte-level sub-word tokenisation, so it is on the right side of that finding - but the result establishes Finnish morphology as a measured, recognised difficulty rather than a vague one.
Vowel and consonant length change the word. Finnish distinguishes short and long sounds phonemically: tuli is fire, tuuli is wind, tulli is customs. Length alone carries the meaning, so an error here produces a different real word rather than something that looks wrong. This is a fact about Finnish phonology; we have not measured how often Whisper gets it right and make no claim that we fix it.
What we actually offer for Finnish, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Finnish we do not. There is no Finnish tuning layer, no Finnish cleanup pass and no Finnish fine-tune. We do not claim to get suffix chains right and we do not claim to handle vowel length correctly, because we have not measured either.
What SnailText does give you is the best open Finnish speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app. There is a specific ergonomic argument for Finnish: long agglutinated and compound forms are slow and error-prone to type and perfectly natural to say. That benefit stands independently of any accuracy claim.
The audio is processed in RAM and never uploaded, which matters for the medical and research contexts where dictation is common and cloud processing is often not permitted. The privacy guarantee is the architecture, not a policy.
One practical tip: set the dictation language to Finnish explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.