SnailText
EN

Swedish speech to text

Swedish speech to text, running on your own machine

Dictate in Swedish into any app on Mac or Windows. SnailText runs Whisper locally - the same open model everyone else runs - so your audio never leaves the machine.

Download for Macand start dictating in any app
cursor.txt
Talar…
▍
Done

The short version

Swedish sits comfortably among Whisper's well-supported languages. In the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 8.5% word error rate on Swedish FLEURS. Independent work on Swedish ASR puts the newer large-v3 at 7.8% on FLEURS, 9.5% on Common Voice and 11.3% on the NST corpus - one model, three corpora, three different answers, which is the clearest illustration on this site that a benchmark figure is a floor and not a promise. The documented weak point is dialect: speech further from the neutral media-style Swedish that dominates web training data is transcribed measurably worse. SnailText runs the model locally on Mac and Windows with no Swedish-specific tuning layer.

Swedish dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Swedish STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Neutral media-style SwedishGood - around 8.5% on the read benchmarkGenerally good too
Regional dialectsMeasurably worse, and we say soAlso worse; rarely disclosed
Swedish letters a-ring, a-umlaut, o-umlautWhatever the open model producesUsually preserved, varies by vendor
Swedish-specific tuningNone - same open model as any local Whisper setupSome vendors tune per language, most run stock

How accurate is Swedish speech to text with local Whisper?

In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 8.5% word error rate on Swedish FLEURS. That puts Swedish in the same band as French at 8.3% and Turkish at 8.4% in the same table - solid, and comfortably better than Czech or Danish.

Swedish also gives the clearest demonstration on this site of why one number is never the whole story. Research on Swedish speech recognition measured the newer large-v3 model at 7.8% on FLEURS, 9.5% on Common Voice and 11.3% on the NST corpus. Same model, three corpora, and the answer moves by three and a half points depending purely on how the audio was recorded and who was speaking. Treat 8.5% as a floor for clean, neutral, well-recorded Swedish.

Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. SnailText's free tier covers the compact models, and the larger ones are on Pro.

What makes Swedish genuinely hard for speech models

Dialects are transcribed measurably worse, and this is documented. Research into Swedish speech models states it directly: dialects are generally harder to transcribe correctly because they occur less often in the training data, and this is especially noticeable for lower-resourced languages. Swedish spans a wide range, from Scanian in the south through to Norrland in the north, plus Finland Swedish - and the further a speaker sits from the neutral standard, the more the error rate climbs.

The training data biases toward web-typical speech. The same work notes that models trained on web material have effectively only learned to recognise the speakers commonly represented in that material, such as people in YouTube videos. In practice that means neutral broadcast-style Swedish does well and speech outside that profile does less well - which is a property of how the model was built, not something a setting can change.

Diacritics are separate letters, not decoration. Swedish a-ring, a-umlaut and o-umlaut are letters in their own right with their own alphabet positions, and losing one changes the word rather than producing a visible typo. We have no measurement of how often the model gets them right, so we describe this as what a correct Swedish transcript must do, not as something we fix.

What we actually offer for Swedish, and what we do not

For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Swedish we do not. There is no Swedish tuning layer, no Swedish cleanup pass and no Swedish fine-tune. We also do not claim to handle compound words correctly or to manage English loanwords - Swedish speech research does not establish those as measured Whisper failure modes, so we are not going to assert them either way.

What SnailText does give you is the best open Swedish speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app - email, docs, Slack, your editor. No Python to set up, no model plumbing, no account needed to start.

The audio is processed in RAM and never uploaded, which matters more than usual for an audience working under European data-protection expectations: the guarantee is the architecture rather than a policy you have to trust. It also means it works with no connection at all.

One practical tip: set the dictation language to Swedish explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, which matters here because Swedish, Danish and Norwegian are close enough to be confused on short samples.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Swedish speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 8.5% word error rate on Swedish FLEURS, similar to French at 8.3%. Independent Swedish ASR research measured the newer large-v3 at 7.8% on FLEURS, 9.5% on Common Voice and 11.3% on the NST corpus - one model, three corpora, a three-and-a-half point spread. Treat the benchmark as a floor for clean neutral Swedish and expect real dictation to run higher.

Does it handle Swedish dialects?

+

Less well than neutral Swedish, and this is documented rather than speculative. Swedish ASR research states that dialects are harder to transcribe because they appear less often in training data, and that models trained on web material have mostly learned speakers typical of that material. Scanian, Norrland and Finland Swedish speech will see a higher error rate than broadcast-style Swedish. No setting changes this.

Do you tune SnailText specifically for Swedish?

+

No. SnailText ships language-specific prompts for Spanish, German, French, Portuguese and Dutch, but there is no Swedish tuning layer, no Swedish cleanup pass and no Swedish fine-tune. For Swedish you get the same open Whisper model everyone runs, running locally and privately, with the desktop plumbing done for you.

Does it get a-ring, a-umlaut and o-umlaut right?

+

Usually at the larger model sizes, but we make no claim about it because we have not measured it. These are separate letters in Swedish rather than accented variants, so losing one produces a different word rather than an obvious typo - worth a proofread.

Does Swedish dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Swedish dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

Is it really different from just running Whisper myself?

+

For Swedish, the model is exactly the same - there is no Swedish-specific layer on our side. What you get is everything around it: a global hotkey, paste-at-cursor into any application, model management, a free tier, and the guarantee that audio stays in RAM on your machine.

·

Swedish speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Swedish, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app