How accurate is Arabic transcription with local Whisper?
For read Modern Standard Arabic there is a published figure worth trusting, because it comes from the model's own paper. In arXiv:2212.04356, large-v2 scores 16.0% word error rate on Arabic FLEURS. Read against the same table, that is roughly three and a half times German (4.5%) and about twice French (8.3%). It is drafting quality: good enough to capture a thought and fix it, not good enough to send unreviewed.
FLEURS is studio-read speech, so treat 16.0% as a floor rather than an expectation. Anything with background noise, a phone microphone or natural conversational pace will measure worse, and that is true of every language on this site - we just say it rather than quoting the benchmark as though it were a promise.
Accuracy tracks model size far more than it tracks local versus cloud, since it is the same open Whisper model either way. For Arabic, use the largest model your machine will run. The compact models are the free ones, and for Arabic they are not a realistic choice - so being direct about the commercial part, this language effectively needs Pro.
Which Arabic? The dialect question decides everything
Most pages selling Arabic transcription quote one accuracy number. That number is almost always Modern Standard Arabic, which is the language of news broadcasts and formal writing - and which very few people speak conversationally. If you dictate the way you talk, MSA figures do not describe your experience.
There is research specifically on this. Talafha, Waheed and Abdul-Mageed (2023) benchmarked Whisper across Arabic datasets and found that while zero-shot Whisper beats fully finetuned XLS-R models overall, its performance "deteriorates significantly in the zero-shot setting for five unseen dialects" - naming Algerian, Jordanian, Palestinian, Emirati and Yemeni. Zero-shot is exactly how SnailText runs it, and how every local Whisper setup runs it.
The practical reading: Egyptian and Levantine material fares better than Maghrebi, MSA fares better than any dialect, and a dictation habit built on formal register will hold up where one built on everyday speech will not. Anyone promising uniform Arabic accuracy across dialects is either finetuning per dialect or not measuring.
What we actually offer for Arabic, and what we do not
For Spanish, German, French, Portuguese, Dutch and Russian, SnailText ships a language-specific prompt that biases the model toward that language's spelling and punctuation. For Arabic we do not. There is no Arabic tuning layer, no Arabic cleanup pass, and no dialect fine-tune. We make no claim about diacritics, no claim about hamza placement, and no claim about dialect handling beyond what the research above reports.
What SnailText gives you is the strongest open Arabic speech model running entirely on your own hardware, reached with one hotkey in any application, with the audio processed in RAM and never uploaded. For Arabic that last part carries more weight than usual: a great deal of Arabic dictation is journalism, legal work, and conversation about matters people have concrete reasons not to route through someone else's server.
If you need production-grade dialect transcription, what you need is a dialect-finetuned model, and we would rather point you at that than take your money for something we have not built.
Right-to-left text, and where it actually breaks
Arabic is written right to left, and dictation tools rarely mention what that means in practice. SnailText transcribes speech to text and pastes it at your cursor; the direction of that text is then handled by whatever application received it. In an app with proper bidirectional support the result is correct. In one without, mixed Arabic and Latin content - a sentence with a product name or a URL in it - can display in a confusing order even though the underlying characters are right.
This is a text-field problem rather than a speech-recognition one, and it affects cloud tools identically. Worth knowing before you conclude the transcription is wrong: copy the same text into a different editor and check whether the order changes. If it does, the model heard you correctly and the display is at fault.