How accurate is Danish speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 13.8% word error rate on Danish FLEURS. Set beside its neighbours in the same table and on the same model - Norwegian 9.5%, Swedish 8.5%, German 4.5% - Danish is clearly the hardest of the Scandinavian group out of the box.
This is not a quirk of one benchmark. The Whisper paper establishes a strong relationship between how much training data a language had and how well the model performs on it, and Danish sits low on that axis. The existence of several dedicated Danish fine-tunes is the community's own verdict that the base model leaves room on Danish.
That figure is clean read speech and acts as a floor; real dictation runs higher. Accuracy tracks model size far more than local versus cloud, so use the largest model you can run.
What makes Danish genuinely hard for speech models
Stod, and the fact that nobody can define it acoustically. Danish stod is a phonological feature that distinguishes words, but research on Danish speech recognition documents that it is highly variable and can be clearly audible while not being visible on a spectrogram - and that there is still no generally accepted acoustic or phonetic definition of it. A feature that human listeners use to tell words apart but that resists acoustic definition is close to a worst case for a speech model.
Heavy reduction, and a large gap between writing and speech. Connected Danish is strongly reduced and its vowels weaken, and the same research finds that distinguishing long from short vowels makes a measurable contribution to recognition accuracy. This is the commonly cited reason Danish is harder than Swedish or Norwegian for speech recognition generally, not only for Whisper.
Danish and Norwegian Bokmal can be identical in writing. Work on Scandinavian language identification found sentences that were orthographically identical in Danish and Bokmal, distinguished only by a comma, which forced the researchers to adopt multi-label annotation because individual sentences could not be assigned to one language. We have not found a documented case of Whisper mistaking Danish for Norwegian from audio, and we are not going to claim one - but the textual closeness is a good reason to set the language explicitly rather than trusting automatic detection.
What we actually offer for Danish, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Danish we do not. There is no Danish tuning layer, no Danish cleanup pass and no Danish fine-tune. Given that dedicated Danish fine-tunes exist and perform better than base Whisper, the honest statement is that we ship the base model and you should calibrate your expectations to the numbers above.
What SnailText does give you is the best open Danish speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app. The audio is processed in RAM and never uploaded, so the privacy guarantee is architectural, and it works with no connection at all.
Set the dictation language to Danish explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, and given how close written Danish and Norwegian Bokmal can be, this is worth doing every time.
Danish also rewards speaking a little more deliberately than feels natural. Given how much connected Danish reduces, clearer articulation genuinely helps, which is advice rather than a feature claim.