How accurate is Hungarian speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 17.0% word error rate on Hungarian FLEURS. The comparison that matters is with the same table and the same model: German 4.5%, French 8.3%, Finnish 9.7%. Hungarian is roughly four times German and about one and three quarter times Finnish, and that gap is the single most useful fact on this page.
It gets harder off the benchmark. Published measurements of Hungarian recognition on spontaneous speech - people talking normally rather than reading - run above 30%. Roughly one word in three needing attention is a different kind of workflow from occasional proofreading, and it is better to know that before you build a habit around it.
Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. For Hungarian, use the largest model you can run. To be direct about what that means commercially: the compact models are the free ones and they are not a realistic option for Hungarian, so this language effectively needs Pro. We would rather say that than imply it.
What makes Hungarian genuinely hard for speech models
Agglutination drives a high out-of-vocabulary rate. Hungarian builds words by stacking suffixes onto a stem, so a single dictionary word appears in text as a very large number of distinct forms. This is the best-established reason for high error rates in Hungarian speech recognition: no training corpus contains every form, so the model regularly meets word shapes it has not seen. The characteristic error is the same as in Finnish and Turkish - the stem is recognised and the ending is not, leaving a readable sentence whose grammar has shifted.
Hungarian is under-represented, and scale alone does not fix it. The Whisper paper establishes a strong relationship between the volume of training data for a language and the model's performance on it. Hungarian sits low on that axis, and the research on agglutinative languages finds that morphological complexity continues to hurt even large models rather than being absorbed by them.
Vowel length is phonemic and compulsory in writing, but blurred in speech. Hungarian marks long vowels with accents, and the distinction changes meaning. In natural speech the contrast is often much less clear than the orthography implies, which leaves the model inferring from context where the writing system demands certainty.
What we actually offer for Hungarian, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Hungarian we do not. There is no Hungarian tuning layer, no Hungarian cleanup pass and no Hungarian fine-tune. We make no claim about suffix chains and none about vowel length marking.
What SnailText gives you is the best open Hungarian speech model running entirely on your own hardware, with a global hotkey that pastes text at your cursor in any app, and audio that never leaves your machine. Given the accuracy above, the honest positioning is that this is a private, offline way to run the same model everyone else runs - not a claim that Hungarian recognition works well.
Where it is genuinely useful: first drafts you will edit anyway, notes for yourself, and long compound forms that are tedious to type. Where it is not: anything going out unreviewed. We would rather draw that line clearly than have you find it.
Set the dictation language to Hungarian explicitly rather than leaving it on automatic, and speak in complete, deliberate sentences. Both help more than any setting we could offer.