SnailText
EN

Polish speech to text

Polish speech to text, running on your own machine

Dictate in Polish into any app on Mac or Windows. SnailText runs Whisper locally - the same open model everyone else runs - so your audio never leaves the machine. Press a hotkey, speak Polish, and the text lands at your cursor.

Download for Macand start dictating in any app
cursor.txt
Mówię…
▍
Done

The short version

Polish is a well-supported Whisper language, and unusually, its two published benchmarks agree. In the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 5.0% word error rate on Polish Multilingual LibriSpeech and 5.4% on Polish FLEURS. Both are clean read speech, so everyday dictation runs above them - a published study of Polish medical conversation put Whisper at 20.84% on that much harder material. SnailText runs the same open model locally on Mac and Windows, with no cloud and no account to start. We do not ship a Polish-specific tuning layer, and this page does not pretend otherwise.

Polish dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Polish STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Diacritics (ą ć ę ł ń ó ś ź ż)Whatever the open model producesUsually preserved, varies by vendor
Polish-specific tuningNone - same open model as any local Whisper setupSome vendors tune per language, most run stock
Hallucination guard on silenceAlways on, and language-agnosticVaries; often no filtering at all
Account / costNo account to startAccount + per-minute or per-seat billing

How accurate is Polish speech to text with local Whisper?

Polish is one of the languages where the published numbers line up cleanly. In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 5.0% word error rate on Polish Multilingual LibriSpeech and 5.4% on Polish FLEURS. Two different corpora agreeing within half a point is a good sign - it means the figure is about the language, not about which recordings someone picked.

Both of those are clean read speech, so treat them as a floor rather than an expectation. How far the floor sits from real life is documented for Polish specifically: a published comparison measured Whisper at 20.84% word error rate on Polish medical conversation, a deliberately hard domain with specialist vocabulary and overlapping speakers. That is roughly four times the benchmark, on the same language and the same model family.

Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. The compact models are not enough for real Polish dictation; quality becomes usable at the small model and best at the largest. SnailText's free tier covers the compact models, and the larger ones are on Pro.

What makes Polish genuinely hard for speech models

Seven cases, and a long tail of word forms. Polish inflects nouns, adjectives and numerals heavily, so a single dictionary word appears in the wild as dozens of surface forms. Research on Polish ASR benchmarking names this directly: complex morphology drives a high out-of-vocabulary rate, because no training set can contain every form of every word. The practical shape of the error is that the stem is recognised and the ending is not - and in Polish the ending is what carries the case, the number and often the meaning.

Three sibilant series that differ very little acoustically. Polish contrasts s/z, s-acute/z-acute and sz/z-dot, plus the matching affricates c, c-acute and cz. Polish speech-assessment research probes exactly these contrasts as the difficult set, using test words like szafa, szufelka, czapka and dziadek. That work is about screening children's pronunciation rather than about Whisper, so treat it as evidence that these are the hard contrasts in Polish phonology, not as a measurement of what our model gets wrong.

Whisper invents text when it hears no speech, and this was studied on Polish. A 2025 paper investigated Whisper hallucinations induced by non-speech audio and found the model emits fabricated transcript text when fed silence or noise. This is the one challenge on this page where SnailText has a real answer that is not a tuning claim: silence filtering, voice-coverage filtering and a hallucination blocklist run on every transcript. None of that is Polish-specific - it is global and language-agnostic - but the failure it prevents was documented on Polish audio.

What we actually offer for Polish, and what we do not

For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Polish we do not. There is no Polish tuning layer, no Polish cleanup pass and no Polish fine-tune, so this page makes no accuracy claim beyond what the open model does on its own. In particular we do not claim to fix dropped diacritics, and we do not claim to resolve the o-acute versus u or z-dot versus rz spellings, which are homophones in modern Polish and can only be inferred from vocabulary.

What SnailText does give you is the best open Polish speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app - email, docs, Slack, your editor. No Python to set up, no model plumbing, no account needed to start.

The hallucination guard is worth one more line because it is genuinely ours and genuinely useful. Whisper's habit of filling silence with invented sentences is well documented, and SnailText filters it at three levels before the text ever reaches your cursor. That protection applies to every language, Polish included.

One practical tip: set the dictation language to Polish explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, so an English word at the start can colour everything after it.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Polish speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 5.0% word error rate on Polish Multilingual LibriSpeech and 5.4% on Polish FLEURS. Two corpora agreeing that closely is unusual and means the number is reliable for clean read speech. Real dictation runs higher: a published study measured Whisper at 20.84% on Polish medical conversation, a much harder domain. Accuracy otherwise depends mostly on model size - the compact free-tier models are for casual use, and the larger Pro models are where Polish gets genuinely good.

Do you tune SnailText specifically for Polish?

+

No, and we would rather say so. SnailText ships language-specific prompts for Spanish, German, French, Portuguese and Dutch, but there is no Polish tuning layer, no Polish cleanup pass and no Polish fine-tune. For Polish you get the same open Whisper model everyone runs, with no per-language layer on top - just running locally, privately, with the desktop plumbing done for you.

Does Polish dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Polish dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

Does it get Polish diacritics right?

+

Usually, at the larger model sizes, but we make no claim about it because we have not measured it. Polish has nine diacritic letters and several of them form minimal pairs against their bare counterparts, so a dropped mark tends to produce a different real word rather than a visible typo. Worth a proofread. Note also that o-acute versus u, and z-dot versus rz, are homophones in modern Polish - the model has to infer those from vocabulary, since they sound identical.

Will it make things up if I pause or stay silent?

+

This is a real Whisper failure mode - a 2025 study documented the model fabricating transcript text when fed non-speech audio, using Polish recordings. SnailText filters it at three levels: silence trimming, voice-coverage checking per segment, and a blocklist of known hallucination phrases. All three are always on and apply to every language. It is not a Polish-specific feature, but it is a real one.

Is it really different from just running Whisper myself?

+

For Polish, the model is exactly the same - there is no Polish-specific layer on our side. What you get is everything around it: a global hotkey, paste-at-cursor into any application, model management, a free tier, the hallucination filtering described above, and the guarantee that audio stays in RAM on your machine. If you already have a local Whisper setup you are happy with, the honest answer is that the difference is convenience and privacy plumbing, not Polish accuracy.

·

Polish speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Polish, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app