How accurate is Turkish speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 8.4% word error rate on Turkish FLEURS. For context from the same table and the same model, French scores 8.3% and Dutch 6.7% - so Turkish sits comfortably among the well-supported languages, and the tone of this page can be confident about that.
There is a real caveat in how that number is calculated, and it works against Turkish. Word error rate counts whole words, and Turkish packs into one word what English spreads across several - tense, person, negation, case and possession all ride on suffixes attached to the stem. So a single wrong suffix, with the stem recognised perfectly, still counts as one full word error. In that sense 8.4% on Turkish represents a finer-grained accuracy than 8.3% on French, even though the numbers look identical.
As always, that is clean read speech and describes a floor. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. SnailText's free tier covers the compact models, and the larger ones are on Pro.
What makes Turkish genuinely hard for speech models
Agglutination makes the vocabulary effectively unbounded. Turkish builds words by stacking suffixes, and the morphology is productive enough that a single root generates an enormous number of valid forms. The Turkish speech recognition literature identifies this as the historical driver of high error rates: the out-of-vocabulary rate stays high because no training corpus can contain every form. The practical shape of the error is distinctive - the root comes out right and the suffix chain does not, so the sentence stays readable while the tense, the case or the negation quietly changes.
The metric punishes that pattern disproportionately. Because one Turkish word does the work of several English ones, an error that would cost a fraction of a word in English costs a whole one here. This is worth knowing when you compare published Turkish numbers against published English numbers; they are not measuring equivalent amounts of meaning.
Dotted and dotless i are separate letters. Turkish is one of the few languages where i and dotless i are genuinely different letters with different sounds and their own alphabet positions, and the same applies to their capitals. The well-known consequence in software is that naive case conversion written for English corrupts Turkish words. We found no measurement of how often Whisper confuses the two, so treat this as a risk in whatever processes the text afterwards rather than as a claim about the model - and note that we do not do anything special about it either.
What we actually offer for Turkish, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Turkish we do not. There is no Turkish tuning layer, no Turkish cleanup pass and no Turkish fine-tune. We do not claim to get suffix chains right, and we do not claim to handle dotted and dotless i casing correctly, because we have not measured either.
What SnailText does give you is the best open Turkish speech model running entirely on your own hardware, with a global hotkey that pastes text at your cursor in any app. For Turkish there is a specific ergonomic argument: long agglutinated words are slow and error-prone to type, and saying them is neither. That is a genuine daily benefit independent of any accuracy claim.
The audio is processed in RAM and never uploaded, so the privacy guarantee is the architecture rather than a policy you have to trust, and it works with no connection at all.
One practical tip: set the dictation language to Turkish explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision. Note also that Azerbaijani, despite the similarity, scores far worse in the same benchmark table - do not assume it comes along for free.