How accurate is Italian speech to text with local Whisper?
Italian is the language where a single accuracy number is most misleading, and it is worth being precise about why. In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 4.0% word error rate on Italian FLEURS (short read sentences) and 13.8% on Italian Multilingual LibriSpeech (read audiobooks). Same model, same paper, same year - a 3.5x spread that comes entirely from which recordings the benchmark used.
So the honest framing is a range, not a headline: on clean read speech, published results put the largest model somewhere between roughly 4% and 14%. Both of those are read-speech numbers measured in good conditions. Everyday dictation in a normal room, with a normal microphone, sits above that range rather than inside it. Anyone quoting you one Italian figure is quoting a corpus, not a capability.
Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. The compact models are not enough for real Italian dictation; quality becomes usable at the small model and best at the largest. SnailText's free tier covers the compact models, and the larger ones are on Pro.
What makes Italian genuinely hard for speech models
Double consonants change the word, and barely change the sound. Italian distinguishes single from double consonants lexically: pala is a shovel, palla is a ball; casa is a house, cassa is a crate. The acoustic difference can be under 19 milliseconds, and in published measurements it does not always reach statistical significance. The consequence for dictation is specific and unpleasant: when a model guesses wrong, the output is still a real Italian word, so it does not look like an error. It reads as a typo you made.
Regional variation is unusually wide, and the training data under-represents it. A survey cataloguing 66 spoken Italian datasets notes that Italian "is marked by significant dialectal variation, yet publicly available large-scale corpora have remained comparatively underrepresented compared to major world languages". Worse for the point above, how speakers realise double consonants itself differs between northern and central-southern Italy - so the gemination problem compounds with where the speaker is from.
Accents are load-bearing in the writing, not decoration. A correct Italian transcript has to get the acute and grave accents right, and they are not interchangeable: perché takes an acute, città takes a grave, and writing perchè is simply wrong. Dropping the accent on è ("is") turns it into e ("and") - a meaning change no spellchecker flags. We have no measurement of how often Whisper gets these right, so we are describing what a correct transcript must do, not claiming we fix it.
What we actually offer for Italian, and what we do not
Being straight about this is more useful than a marketing line. For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Italian we do not. There is no Italian tuning layer, no Italian cleanup pass, and no Italian fine-tune, so this page makes no accuracy claim beyond what the open model does on its own.
What SnailText does give you for Italian is the part that is genuinely missing elsewhere: the best open Italian speech model, running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app - email, docs, Slack, your editor. No Python to set up, no model plumbing, no account needed to start.
That combination matters most when what you dictate is not casual. Client notes, medical or legal drafting, internal company text - the audio is processed in RAM and never uploaded, so the privacy guarantee is the architecture rather than a policy you have to take on trust. It also means it works on a plane, or with no wifi at all.
One practical tip that does help Italian: set the dictation language to Italian explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, so an English word at the start can colour everything after it.