SnailText
EN

Italian speech to text

Italian speech to text, running on your own machine

Dictate in Italian into any app on Mac or Windows. SnailText runs Whisper locally - the same open model everyone else runs - so your audio never leaves the machine. Press a hotkey, speak Italian, and the text lands at your cursor.

Download for Macand start dictating in any app
cursor.txt
Sto parlando…
▍
Done

The short version

Italian is the language where the benchmark decides the number. In the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 4.0% word error rate on Italian FLEURS (short read sentences) and 13.8% on Italian Multilingual LibriSpeech (read audiobooks). Those are two corpora, not two models, and both are clean read speech - so everyday dictation sits above both figures rather than between them. SnailText runs that same open model locally on Mac and Windows, with no cloud and no account to start, and pastes the result at your cursor in any app. We do not ship an Italian-specific tuning layer, and this page does not pretend otherwise: what you get is the best open Italian model, running privately, with the desktop plumbing already wired up.

Italian dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Italian STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Accented vowels (à è é ì ò ù)Whatever the open model producesUsually preserved, varies by vendor
Italian-specific tuningNone - same open model as any local Whisper setupSome vendors tune per language, most run stock
Dictate into any appGlobal hotkey, pastes at your cursorUsually a web app you copy out of
Account / costNo account to startAccount + per-minute or per-seat billing

How accurate is Italian speech to text with local Whisper?

Italian is the language where a single accuracy number is most misleading, and it is worth being precise about why. In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 4.0% word error rate on Italian FLEURS (short read sentences) and 13.8% on Italian Multilingual LibriSpeech (read audiobooks). Same model, same paper, same year - a 3.5x spread that comes entirely from which recordings the benchmark used.

So the honest framing is a range, not a headline: on clean read speech, published results put the largest model somewhere between roughly 4% and 14%. Both of those are read-speech numbers measured in good conditions. Everyday dictation in a normal room, with a normal microphone, sits above that range rather than inside it. Anyone quoting you one Italian figure is quoting a corpus, not a capability.

Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. The compact models are not enough for real Italian dictation; quality becomes usable at the small model and best at the largest. SnailText's free tier covers the compact models, and the larger ones are on Pro.

What makes Italian genuinely hard for speech models

Double consonants change the word, and barely change the sound. Italian distinguishes single from double consonants lexically: pala is a shovel, palla is a ball; casa is a house, cassa is a crate. The acoustic difference can be under 19 milliseconds, and in published measurements it does not always reach statistical significance. The consequence for dictation is specific and unpleasant: when a model guesses wrong, the output is still a real Italian word, so it does not look like an error. It reads as a typo you made.

Regional variation is unusually wide, and the training data under-represents it. A survey cataloguing 66 spoken Italian datasets notes that Italian "is marked by significant dialectal variation, yet publicly available large-scale corpora have remained comparatively underrepresented compared to major world languages". Worse for the point above, how speakers realise double consonants itself differs between northern and central-southern Italy - so the gemination problem compounds with where the speaker is from.

Accents are load-bearing in the writing, not decoration. A correct Italian transcript has to get the acute and grave accents right, and they are not interchangeable: perché takes an acute, città takes a grave, and writing perchè is simply wrong. Dropping the accent on è ("is") turns it into e ("and") - a meaning change no spellchecker flags. We have no measurement of how often Whisper gets these right, so we are describing what a correct transcript must do, not claiming we fix it.

What we actually offer for Italian, and what we do not

Being straight about this is more useful than a marketing line. For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Italian we do not. There is no Italian tuning layer, no Italian cleanup pass, and no Italian fine-tune, so this page makes no accuracy claim beyond what the open model does on its own.

What SnailText does give you for Italian is the part that is genuinely missing elsewhere: the best open Italian speech model, running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app - email, docs, Slack, your editor. No Python to set up, no model plumbing, no account needed to start.

That combination matters most when what you dictate is not casual. Client notes, medical or legal drafting, internal company text - the audio is processed in RAM and never uploaded, so the privacy guarantee is the architecture rather than a policy you have to take on trust. It also means it works on a plane, or with no wifi at all.

One practical tip that does help Italian: set the dictation language to Italian explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision, so an English word at the start can colour everything after it.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Italian speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 4.0% word error rate on Italian FLEURS and 13.8% on Italian Multilingual LibriSpeech. That is one model measured on two corpora, not an inconsistent model: FLEURS is short read sentences and MLS is read audiobooks. Both are clean read speech, so everyday dictation in a normal room runs above both numbers. Accuracy otherwise depends mostly on model size - the compact free-tier models are for casual use, and the larger Pro models are where Italian gets genuinely good.

Do you tune SnailText specifically for Italian?

+

No, and we would rather say so. SnailText ships language-specific prompts for Spanish, German, French, Portuguese and Dutch, but there is no Italian tuning layer, no Italian cleanup pass and no Italian fine-tune. For Italian you get the same open Whisper model everyone runs, with no per-language layer on top - just running locally, privately, with the desktop plumbing done for you.

Does Italian dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Italian dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

Will it confuse words like pala and palla?

+

Sometimes, and it is the honest weak spot of Italian dictation on any tool. Italian distinguishes single from double consonants lexically, and the acoustic difference can be under 19 milliseconds. When a model guesses wrong the result is still a real Italian word, so it reads like a typo rather than a recognition error - worth a proofread on words where the doubling carries the meaning. A larger model helps but does not eliminate it.

Does it handle Italian regional accents and dialects?

+

Standard Italian across the regions generally works, but we make no dialect claims. Neapolitan, Sicilian and Venetian are treated in much of the literature as distinct languages rather than accents, and we have no evidence about how the model handles them. Research also notes that large-scale Italian corpora under-represent the dialectal variation of the country, so accuracy is not uniform across regions.

Is it really different from just running Whisper myself?

+

For Italian, the model is exactly the same - there is no Italian-specific layer on our side. What you get is everything around it: a global hotkey, paste-at-cursor into any application, model download and management, a free tier, and the guarantee that audio stays in RAM on your machine. If you already have a local Whisper setup you are happy with, the honest answer is that the difference is convenience and privacy plumbing, not Italian accuracy.

·

Italian speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Italian, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app