How accurate is Catalan speech to text with local Whisper?
In the Whisper paper itself (arXiv:2212.04356), the large-v2 model scores 7.3% word error rate on Catalan FLEURS and 14.1% on Catalan Common Voice 9. That is one model on two corpora, nearly doubling - FLEURS is read sentences recorded in good conditions, Common Voice is crowd-recorded by volunteers on whatever hardware they had. Quoting 7.3% alone would flatter the language; the honest range is 7 to 14% depending on how controlled the speech is.
For context, 7.3% puts Catalan level with French on the read benchmark and ahead of Dutch, which is a genuinely good result for a language with far fewer speakers. The reason is not an accident: Catalan has one of the largest validated-hours collections in Mozilla's Common Voice project, backed by sustained public language-technology funding, so there is unusually good open training data for its size.
Both figures describe controlled speech and act as a floor. Accuracy tracks model size far more than it tracks local versus cloud - it is the same open-source Whisper either way.
What makes Catalan genuinely hard for speech models
Mixed Catalan-Spanish speech is the hard limit, and it is measured. Work presented at Interspeech 2025 on Catalan-Spanish code-switching describes the mixing as commonplace in everyday conversation in Catalan-speaking areas and particularly challenging because the two languages are so similar. Their measurements of base Whisper on genuinely mixed speech land around 21 to 35% word error depending on the corpus, against 7.3% on clean monolingual Catalan. If your speech routinely switches between the two languages mid-sentence, no setting here will fix that.
Getting the language tag wrong is worse than it sounds. The same work reports that forcing an incorrect language tag pushes the error rate above 60%. That is the practical reason this page keeps telling you to select Catalan explicitly: the cost of the model operating under the wrong assumption is severe, not marginal.
Automatic language detection is not reliable enough to lean on. Whisper's language identification is roughly 64.5% accurate overall and is decided from about the first thirty seconds of audio. There is a documented pattern of under-represented Romance languages being collapsed into a larger neighbour - it has been measured for Galician, whose identification accuracy against Spanish is dramatically lower. We have not seen an equivalent published measurement for Catalan specifically, so we are not going to assert one; the defensible point is simply that automatic detection is a coin-flip you do not need to take.
What we actually offer for Catalan, and what we do not
For Spanish, German, French, Portuguese and Dutch, SnailText ships a language-specific prompt. For Catalan we do not. There is no Catalan tuning layer and no Catalan cleanup pass. It is worth being explicit about one thing here: we do not apply our Spanish tuning to Catalan and call it Catalan support. Catalan is its own language and is treated as its own selection.
In SnailText the dictation language is an explicit choice - Catalan is in the list, and you pick it. Given how badly a wrong language tag degrades output, that is the most useful practical property this page can offer, and it is a fact about how the app works rather than a claim about accuracy.
We do not offer Valencian or Balearic as separate options. Whisper folds Valencian into the same Catalan language code, so there is no separate model to select, and published results suggest Balearic speech fares measurably worse. Treat the page as covering standard Catalan.
The audio is processed in RAM and never uploaded, so the privacy guarantee is architectural, and dictation works with no connection at all.