How accurate is Greek speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 12.5% word error rate on Greek FLEURS. This is the most reliably cross-checked figure on this site: research published at Interspeech 2024, building an 800-hour Greek podcast corpus, independently measured the same model at 12.95% on the same benchmark, and the newer large-v3 at 11.02%. Two independent sources agreeing within half a point is about as solid as a published number gets.
That is clean read speech and acts as a floor. The same peer-reviewed work classifies modern Greek as low-resourced for speech technology, which is the honest context for why the figure sits where it does rather than alongside German or Spanish.
Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. Use the largest model you can run.
What makes Greek genuinely hard for speech models
The accent mark is compulsory, not decorative. Modern Greek uses the monotonic orthography introduced in 1982, which replaced the older system of multiple accent and breathing marks with a single stress mark. That mark is orthographically required on multi-syllable words, so a transcript without it is not simply plain - it is misspelled Greek. The older polytonic system is still used in ecclesiastical, classical and some literary text, which means the correct output depends on context the model does not know about.
Letter forms change by position. Lowercase sigma has two forms: one used at the end of a word and another everywhere else. This is purely positional, so any error in word segmentation or a word wrongly joined to its neighbour produces an orthographically wrong form. Capitalisation has its own Greek-specific rules too, including dropping accents in all-caps text and placing the accent to the left of a capitalised initial letter rather than above it - which means correct Greek casing needs Greek-specific logic, not a generic uppercase function.
Hallucination on music, and English code-switching. The Greek podcast corpus research documents both: Whisper producing invented text over musical passages that silence detection had not removed, and Greek-English code-switching in phrases like follow me on Instagram breaking monolingual alignment. The first is a class of failure we do guard against - silence trimming, voice-coverage filtering and a hallucination blocklist run on every transcript, in every language. The second we do not handle.
What we actually offer for Greek, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Greek we do not. There is no Greek tuning layer, no Greek cleanup pass and no Greek fine-tune. We make no claim about restoring the stress mark, handling final sigma, or applying Greek capitalisation rules, and we do not support polytonic orthography - so this is not the tool for classical or ecclesiastical text.
The hallucination guard is worth stating precisely because it is real and it maps onto a documented Greek problem. Whisper's habit of inventing text over music and silence is filtered at three levels before anything reaches your cursor. That is global and language-agnostic rather than Greek-specific, but the failure it prevents was documented on Greek audio.
There is one audience this page is genuinely well suited to, and it is worth naming: people whose spoken Greek is more fluent than their written Greek, which describes a great deal of the Greek diaspora. Dictation lets you produce correct written Greek - accents, final sigma and all - from speech you are already comfortable with. We are not claiming the output is flawless; we are saying the gap it closes is a real one.
Set the dictation language to Greek explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.