How accurate is Romanian speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 14.4% word error rate on Romanian FLEURS. That is usable but clearly behind the big Western European languages, and it reflects the fact that research literature still describes Romanian as a low-resource language in the context of speech recognition.
That number is clean read speech and acts as a floor. Romanian ASR research also notes that dialect makes a large difference, with measurable divergence between standard Romanian and Moldovan speech, so the figure describes standard Romanian rather than the whole language area.
Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way. Use the largest model you can run.
What makes Romanian genuinely hard for speech models
Diacritics get dropped, and that is a real error rather than a cosmetic one. Romanian has five diacritic letters, and dropping them changes words. We do not add any Romanian post-processing to restore them, so whatever the model produces is what you get - worth a proofread, particularly on the two letters that carry commas below.
Two of those letters are routinely mis-encoded. The correct Romanian characters use a comma below, but visually similar Turkish letters with a cedilla have historically been substituted by software and fonts, and that substitution was the default for years in some environments. The result looks almost right and is wrong, which means it survives casual proofreading. This is a text-encoding hazard in whatever handles the text after us, and we make no claim to normalise it.
Silence detection can clip the end of a sentence, and this one is about our own pipeline. Romanian speech research documents that voice activity detection often truncates the final phonemes of a sentence when a speaker hesitates or a conversation turns over quickly, giving a clipped word instead of the full form. SnailText uses voice activity detection as part of its transcription pipeline, so this applies to us. We would rather name it than let you discover it: if you trail off or pause mid-sentence, check the last word.
What we actually offer for Romanian, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Romanian we do not. There is no Romanian tuning layer, no Romanian cleanup pass and no Romanian fine-tune. We do not restore dropped diacritics, we do not normalise the comma-versus-cedilla encoding issue, and we do not perform the morphological completion that research proposes for clipped sentence endings.
What SnailText does give you is the best open Romanian speech model running entirely on your own hardware, wired into a global hotkey that pastes text at your cursor in any app. The audio is processed in RAM and never uploaded, and it works with no connection at all.
Practical advice that follows from the section above: speak through to the end of your sentences rather than trailing off, and proofread diacritics. Both are usage habits rather than product features, and both will do more for your results than any setting.
Set the dictation language to Romanian explicitly rather than leaving it on automatic. Whisper decides the language from roughly the first thirty seconds and does not revisit that decision.