How accurate is Swedish speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 8.5% word error rate on Swedish FLEURS. That puts Swedish in the same band as French at 8.3% and Turkish at 8.4% in the same table - solid, and comfortably better than Czech or Danish.
Swedish also gives the clearest demonstration on this site of why one number is never the whole story. Research on Swedish speech recognition measured the newer large-v3 model at 7.8% on FLEURS, 9.5% on Common Voice and 11.3% on the NST corpus. Same model, three corpora, and the answer moves by three and a half points depending purely on how the audio was recorded and who was speaking. Treat 8.5% as a floor for clean, neutral, well-recorded Swedish.
Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. SnailText's free tier covers the compact models, and the larger ones are on Pro.
What makes Swedish genuinely hard for speech models
Dialects are transcribed measurably worse, and this is documented. Research into Swedish speech models states it directly: dialects are generally harder to transcribe correctly because they occur less often in the training data, and this is especially noticeable for lower-resourced languages. Swedish spans a wide range, from Scanian in the south through to Norrland in the north, plus Finland Swedish - and the further a speaker sits from the neutral standard, the more the error rate climbs.
The training data biases toward web-typical speech. The same work notes that models trained on web material have effectively only learned to recognise the speakers commonly represented in that material, such as people in YouTube videos. In practice that means neutral broadcast-style Swedish does well and speech outside that profile does less well - which is a property of how the model was built, not something a setting can change.
Diacritics are separate letters, not decoration. Swedish a-ring, a-umlaut and o-umlaut are letters in their own right with their own alphabet positions, and losing one changes the word rather than producing a visible typo. We have no measurement of how often the model gets them right, so we describe this as what a correct Swedish transcript must do, not as something we fix.
What we actually offer for Swedish, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Swedish we do not. There is no Swedish tuning layer, no Swedish cleanup pass and no Swedish fine-tune. We also do not claim to handle compound words correctly or to manage English loanwords - Swedish speech research does not establish those as measured Whisper failure modes, so we are not going to assert them either way.
What SnailText does give you is the best open Swedish speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app - email, docs, Slack, your editor. No Python to set up, no model plumbing, no account needed to start.
The audio is processed in RAM and never uploaded, which matters more than usual for an audience working under European data-protection expectations: the guarantee is the architecture rather than a policy you have to trust. It also means it works with no connection at all.
One practical tip: set the dictation language to Swedish explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, which matters here because Swedish, Danish and Norwegian are close enough to be confused on short samples.