SnailText
EN

Finnish speech to text

Finnish speech to text, running on your own machine

Dictate in Finnish into any app on Mac or Windows. Finnish words get very long, and speaking them beats typing them. SnailText runs Whisper locally, so nothing is uploaded.

Download for Macand start dictating in any app
cursor.txt
Puhutaan…
▍
Done

The short version

In the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 9.7% word error rate on Finnish FLEURS, which puts Finnish in the same band as Norwegian and ahead of Czech and Danish. That is clean read speech, so real dictation runs higher. The structural difficulty is agglutination: Finnish builds words by stacking suffixes, which together with compounding produces millions of distinct but genuinely common word forms, and classical vocabulary-based recognition breaks down on it. Measured work on Finnish showed that moving from whole-word units to sub-word units cut out-of-vocabulary rate from 20% to zero and word error from 56% to 32%. SnailText runs the same open model locally with no Finnish-specific tuning layer.

Finnish dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Finnish STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Long compound word formsFar quicker to speak than to typeSame benefit, if you accept the upload
Vowel and consonant lengthMeaning-bearing - worth proofreadingSame difficulty for every tool
Finnish-specific tuningNone - same open model as any local Whisper setupSome vendors tune per language, most run stock
Account / costNo account to startAccount + per-minute or per-seat billing

How accurate is Finnish speech to text with local Whisper?

In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 9.7% word error rate on Finnish FLEURS. That places Finnish close to Norwegian at 9.5% and comfortably ahead of Czech at 13.3% and Danish at 13.8% in the same table. An independent Finnish team reported 10.20% on the same benchmark for their own large-v2 fine-tune, which corroborates the order of magnitude.

That is clean read speech and describes a floor; everyday dictation runs higher. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way.

There is a measurement caveat specific to agglutinative languages. Word error rate counts whole words, and a Finnish word form can carry as much information as an English phrase, so one wrong suffix inside a long word costs a full word error. Finnish numbers therefore look worse relative to English than the underlying accuracy warrants.

What makes Finnish genuinely hard for speech models

Agglutination produces a combinatorial explosion of word forms. Finnish builds words by concatenating suffixes onto a stem, and combined with compounding and inflection this yields millions of distinct yet genuinely frequent word forms. Classical dictionary-based language modelling does not work on Finnish for precisely this reason, which is why Finnish speech research pioneered morph-based models.

Out-of-vocabulary is effectively unavoidable with whole-word units. The measurements here are striking: Finnish research found that moving from word units to morph-like sub-word units reduced the out-of-vocabulary rate from 20% to zero and cut word error rate from 56% to 32%. Whisper uses byte-level sub-word tokenisation, so it is on the right side of that finding - but the result establishes Finnish morphology as a measured, recognised difficulty rather than a vague one.

Vowel and consonant length change the word. Finnish distinguishes short and long sounds phonemically: tuli is fire, tuuli is wind, tulli is customs. Length alone carries the meaning, so an error here produces a different real word rather than something that looks wrong. This is a fact about Finnish phonology; we have not measured how often Whisper gets it right and make no claim that we fix it.

What we actually offer for Finnish, and what we do not

For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Finnish we do not. There is no Finnish tuning layer, no Finnish cleanup pass and no Finnish fine-tune. We do not claim to get suffix chains right and we do not claim to handle vowel length correctly, because we have not measured either.

What SnailText does give you is the best open Finnish speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app. There is a specific ergonomic argument for Finnish: long agglutinated and compound forms are slow and error-prone to type and perfectly natural to say. That benefit stands independently of any accuracy claim.

The audio is processed in RAM and never uploaded, which matters for the medical and research contexts where dictation is common and cloud processing is often not permitted. The privacy guarantee is the architecture, not a policy.

One practical tip: set the dictation language to Finnish explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Finnish speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 9.7% word error rate on Finnish FLEURS, close to Norwegian at 9.5% and well ahead of Czech at 13.3%. An independent Finnish team reported 10.20% on the same benchmark for their own fine-tune, which corroborates that range. It is clean read speech, so everyday dictation runs higher.

Why does Finnish score worse than English if the model handles it well?

+

Partly because of how the metric works. Word error rate counts whole words, and a single Finnish word form can carry as much meaning as an English phrase, so one wrong suffix inside a long word costs a full word error. Finnish numbers therefore look worse relative to English than the underlying accuracy justifies.

What do Finnish recognition errors usually look like?

+

Typically the stem is right and something in the suffix chain is wrong, which leaves a readable word with a changed grammatical role. The other characteristic error is vowel or consonant length - tuli, tuuli and tulli differ only in length and mean fire, wind and customs respectively. Both kinds of error produce real Finnish words, so proofreading means checking word endings rather than scanning for nonsense.

Do you tune SnailText specifically for Finnish?

+

No. There is no Finnish tuning layer, no Finnish cleanup pass and no Finnish fine-tune. For Finnish you get the same open Whisper model everyone runs, locally and privately, with the desktop plumbing done for you.

Is dictating faster than typing Finnish?

+

For long compound and agglutinated forms, considerably. Those are slow and error-prone to type and entirely natural to speak, which is a practical everyday benefit independent of any accuracy claim.

Does Finnish dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Finnish dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

·

Finnish speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Finnish, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app