SnailText
EN

Greek speech to text

Greek speech to text, running on your own machine

Dictate in Greek into any app on Mac or Windows. Useful if your spoken Greek runs ahead of your written Greek. SnailText runs Whisper locally, so nothing is uploaded.

Download for Macand start dictating in any app
cursor.txt
Μιλάω…
▍
Done

The short version

In the Whisper paper (Radford et al., 2022, arXiv:2212.04356), the large-v2 model scores 12.5% word error rate on Greek FLEURS. That figure is unusually well corroborated: work published at Interspeech 2024 on a Greek podcast corpus independently measured large-v2 at 12.95% on the same benchmark, so the two agree within normalisation differences. That is clean read speech, so real dictation runs higher. Greek is classified in peer-reviewed work as low-resourced for speech technology, and the same research documents two practical artefacts: Whisper hallucinating on music not removed by silence filtering, and Greek-English code-switching breaking alignment. SnailText runs the model locally with no Greek-specific tuning layer, though our hallucination filtering does address the first of those.

Greek dictation: local Whisper vs cloud STT

SnailText (local)Typical cloud Greek STT
Where audio goesStays on your device, in RAMUploaded to a server for every phrase
Works offlineYes, after the model downloads onceNo, needs a connection every time
Accent marks and final sigmaWhatever the open model producesUsually handled, varies by vendor
Hallucination on silence or musicFiltered - always on, language-agnosticVaries; often no filtering at all
Greek-specific tuningNone - same open model as any local Whisper setupSome vendors tune per language, most run stock
Account / costNo account to startAccount + per-minute or per-seat billing

How accurate is Greek speech to text with local Whisper?

In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 12.5% word error rate on Greek FLEURS. This is the most reliably cross-checked figure on this site: research published at Interspeech 2024, building an 800-hour Greek podcast corpus, independently measured the same model at 12.95% on the same benchmark, and the newer large-v3 at 11.02%. Two independent sources agreeing within half a point is about as solid as a published number gets.

That is clean read speech and acts as a floor. The same peer-reviewed work classifies modern Greek as low-resourced for speech technology, which is the honest context for why the figure sits where it does rather than alongside German or Spanish.

Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. Use the largest model you can run.

What makes Greek genuinely hard for speech models

The accent mark is compulsory, not decorative. Modern Greek uses the monotonic orthography introduced in 1982, which replaced the older system of multiple accent and breathing marks with a single stress mark. That mark is orthographically required on multi-syllable words, so a transcript without it is not simply plain - it is misspelled Greek. The older polytonic system is still used in ecclesiastical, classical and some literary text, which means the correct output depends on context the model does not know about.

Letter forms change by position. Lowercase sigma has two forms: one used at the end of a word and another everywhere else. This is purely positional, so any error in word segmentation or a word wrongly joined to its neighbour produces an orthographically wrong form. Capitalisation has its own Greek-specific rules too, including dropping accents in all-caps text and placing the accent to the left of a capitalised initial letter rather than above it - which means correct Greek casing needs Greek-specific logic, not a generic uppercase function.

Hallucination on music, and English code-switching. The Greek podcast corpus research documents both: Whisper producing invented text over musical passages that silence detection had not removed, and Greek-English code-switching in phrases like follow me on Instagram breaking monolingual alignment. The first is a class of failure we do guard against - silence trimming, voice-coverage filtering and a hallucination blocklist run on every transcript, in every language. The second we do not handle.

What we actually offer for Greek, and what we do not

For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Greek we do not. There is no Greek tuning layer, no Greek cleanup pass and no Greek fine-tune. We make no claim about restoring the stress mark, handling final sigma, or applying Greek capitalisation rules, and we do not support polytonic orthography - so this is not the tool for classical or ecclesiastical text.

The hallucination guard is worth stating precisely because it is real and it maps onto a documented Greek problem. Whisper's habit of inventing text over music and silence is filtered at three levels before anything reaches your cursor. That is global and language-agnostic rather than Greek-specific, but the failure it prevents was documented on Greek audio.

There is one audience this page is genuinely well suited to, and it is worth naming: people whose spoken Greek is more fluent than their written Greek, which describes a great deal of the Greek diaspora. Dictation lets you produce correct written Greek - accents, final sigma and all - from speech you are already comfortable with. We are not claiming the output is flawless; we are saying the gap it closes is a real one.

Set the dictation language to Greek explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.

Talk instead of typing

Download for Mac

and start dictating in any app

Frequently asked questions

How accurate is Greek speech to text?

+

In the Whisper paper (arXiv:2212.04356), the large-v2 model scores 12.5% word error rate on Greek FLEURS. Research published at Interspeech 2024 independently measured 12.95% for the same model on the same benchmark, and 11.02% for large-v3 - two sources agreeing within half a point, which makes this the best cross-checked figure on the site. It is clean read speech, so everyday dictation runs higher.

Does it get the Greek accent marks right?

+

Usually at larger model sizes, but we make no claim about it. Modern Greek uses the monotonic system, where the single stress mark is orthographically compulsory on multi-syllable words - a transcript without it is misspelled rather than merely unadorned. We add no Greek-specific post-processing, so this is down to the model.

Does it support polytonic Greek?

+

No. The model targets modern monotonic Greek. Polytonic orthography is still used for ecclesiastical, classical and some literary text, and we make no claim about producing it. If you work with classical or liturgical Greek, this is not the right tool.

Do you tune SnailText specifically for Greek?

+

No. There is no Greek tuning layer, no Greek cleanup pass and no Greek fine-tune. We make no claim about final sigma placement or Greek capitalisation rules, both of which need Greek-specific logic. For Greek you get the same open Whisper model everyone runs, locally and privately.

Will it invent text over music or silence?

+

That is a documented Whisper failure mode, and researchers building a Greek podcast corpus specifically recorded hitting it on musical passages. SnailText filters it at three levels: silence trimming, voice-coverage checking per segment, and a blocklist of known hallucination phrases. All three are always on and apply to every language, Greek included.

Does Greek dictation work offline?

+

Yes. SnailText runs the Whisper speech model on your own Mac or Windows machine, so Greek dictation works with no internet connection once the model has downloaded. The audio is processed in RAM and is never uploaded to any server.

·

Greek speech to text, on your own machine.

Free to start on Mac and Windows. Press Option+Space (Mac) / Ctrl+Space (Windows), speak Greek, and the text lands at your cursor in any app. No account, nothing uploaded, works offline.

Download for Macand start dictating in any app