Best Arabic OCR Models in 2026: Misraj-DocOCR Benchmark Results
Fifteen OCR systems and vision-language models compared on Misraj-DocOCR, a public 400-page Arabic benchmark: word and character error rates, table-structure scores, and how to choose.
Short answer: on Misraj-DocOCR, a public benchmark of 400 expert-checked Arabic document pages, the lowest word error rate belongs to Baseer (0.25), followed by Gemini 2.5 Pro (0.37) and Azure AI Document Intelligence (0.44). Baseer also keeps table structure best (TEDS 66 against Gemini's 52), but it does not win every metric: Azure and Gemini make fewer character-level errors, and Gemini scores slightly higher on BLEU and ChrF.
Every number below comes from the Misraj-DocOCR benchmark and the Baseer paper (arXiv:2509.18174, published 17 September 2025), and the benchmark's Hugging Face card carries the same table. They are the latest published results on this benchmark, and we have not added scores for any model that was not in them.
Why Arabic OCR needs its own benchmark
Arabic is hard on OCR systems built for Latin script. Letters join and change shape by position, diacritics sit above and below the line, text runs right to left, and real documents mix fonts, columns, tables and footnotes. A model can read every word correctly and still return a page nobody can use, because the table came back as a run of loose numbers.
Before Misraj-DocOCR, the usual public test set was the PDF-to-Markdown subset of KITAB-Bench. When the Misraj team reviewed it, they found reference answers containing text that was not on the page, missing page numbers, and small print left out. They published a corrected version, then built a larger benchmark that checks structure as well as characters.
One thing to say plainly: Misraj built both the benchmark and Baseer. That is a fair reason to read the results with some scepticism. It is also why the benchmark is open. The 400 pages and their reference transcriptions are on Hugging Face under an Apache-2.0 licence, so anyone can rerun the comparison instead of taking our word for it.
Methodology
- Dataset. 400 page images, a mix of real and synthetic pages: books, reports, forms, scholarly pages and complex layouts. Human experts reviewed every page for both text and structure. The reference output is Markdown, with tables written as HTML.
- How models were run. Document-OCR systems used their own system prompts. General multimodal models got prompts the team tested beforehand.
- Normalisation. Before scoring, every output went through the same clean-up: Markdown tables converted to HTML, stray HTML outside tables removed, headers and horizontal rules standardised, and model-specific tags such as page-number and watermark markers stripped. That way a model is not penalised for a formatting choice that means the same thing.
- Date. The results were published on 17 September 2025 and reflect the model versions available then.
Six metrics are reported. Two count errors, so lower is better:
- WER (word error rate): the share of words substituted, deleted or inserted. It can go above 1.0 when a model adds a lot of text that is not on the page.
- CER (character error rate): the same idea, counted in characters.
Four measure similarity, so higher is better:
- BLEU: n-gram overlap with the reference text.
- ChrF: a character-level F-score that suits a morphologically rich language like Arabic.
- TEDS: tree edit distance similarity, which checks whether tables and nested structure came back intact.
- MARS: a layout-aware score that combines text and structure.
Results: Arabic OCR models on Misraj-DocOCR
| Model | WER ↓ | CER ↓ | BLEU ↑ | ChrF ↑ | TEDS ↑ | MARS ↑ |
|---|---|---|---|---|---|---|
| Baseer (Misraj) | 0.25 | 0.53 | 76.18 | 87.77 | 66 | 76.885 |
| Gemini 2.5 Pro | 0.37 | 0.31 | 77.92 | 89.55 | 52 | 70.775 |
| Azure AI Document Intelligence | 0.44 | 0.27 | 62.04 | 82.49 | 42 | 62.245 |
| Dots.ocr | 0.50 | 0.40 | 58.16 | 78.41 | 40 | 59.205 |
| Nanonets | 0.71 | 0.55 | 42.22 | 67.89 | 37 | 52.445 |
| Qari | 0.76 | 0.64 | 38.59 | 64.50 | 21 | 42.750 |
| Qwen2.5-VL-32B | 0.76 | 0.59 | 37.62 | 62.64 | 41 | 51.820 |
| GPT-5 | 0.86 | 0.62 | 40.67 | 61.6 | 48 | 54.8 |
| Qwen2.5-VL-3B-Instruct | 0.87 | 0.71 | 25.39 | 53.42 | 27 | 40.210 |
| Qwen2.5-VL-7B | 0.92 | 0.77 | 31.57 | 54.70 | 27 | 40.850 |
| Gemma3-12B | 0.96 | 0.80 | 19.75 | 44.53 | 33 | 38.765 |
| Gemma3-4B | 1.01 | 0.85 | 9.57 | 31.39 | 28 | 29.695 |
| GPT-4o-mini | 1.36 | 1.1 | 22.63 | 47.04 | 26 | 36.52 |
| AIN | 1.23 | 1.11 | 1.25 | 2.24 | 21 | 11.620 |
| Aya-vision | 1.41 | 1.07 | 2.91 | 9.81 | 26 | 17.905 |
Where Baseer wins
Word accuracy. Baseer's WER of 0.25 is the lowest in the table. Gemini 2.5 Pro comes second at 0.37, so Baseer makes about a third fewer word errors than the next system.
Tables and layout. Baseer's TEDS of 66 is the highest by a clear margin. Gemini scores 52, GPT-5 48 and Azure 42. Its MARS of 76.885 is also top. If your documents are financial statements, government forms or reports full of tables, these two columns matter more than any text metric.
Size. Baseer is a 3-billion-parameter model, smaller than the systems it is compared with. It is Qwen2.5-VL-3B-Instruct fine-tuned on 500,000 Arabic image-and-text pairs: 300,000 synthetic pages rendered from Markdown with varied fonts, layouts and simulated scanning damage, and 200,000 real pages from books, magazines, educational material and papers. The vision encoder stayed frozen and only the language decoder was trained. The fine-tuning is what moved the numbers: the base model scores a WER of 0.87 and a TEDS of 27 on the same benchmark.
Where Baseer does not win
Character error rate. Azure AI Document Intelligence has the lowest CER (0.27), followed by Gemini 2.5 Pro (0.31) and Dots.ocr (0.40). Baseer's CER is 0.53. If your acceptance test is character-level accuracy, those systems lead on this benchmark.
BLEU and ChrF. Gemini 2.5 Pro is first on both (77.92 and 89.55). Baseer is a close second (76.18 and 87.77).
The corrected KITAB-Bench set. The paper also tested open-source models on the corrected KITAB-Bench PDF-to-Markdown set. There, Dots.ocr has the best text scores, and Nanonets also posts a lower WER than Baseer. Baseer still leads on TEDS and MARS.
| Model | WER ↓ | CER ↓ | TEDS ↑ | MARS ↑ |
|---|---|---|---|---|
| Dots.ocr | 0.39 | 0.28 | 43 | 63.08 |
| Nanonets | 0.51 | 0.40 | 33 | 55.225 |
| Baseer (Misraj) | 0.61 | 0.40 | 56 | 68.13 |
| Qwen2.5-VL-3B | 0.70 | 0.57 | 31 | 48.89 |
| Qwen2.5-VL-7B | 0.76 | 0.63 | 24 | 43.225 |
The paper notes this set is small (it describes the evaluated subset as 30 samples), so each mistake moves the score a lot. On the 400 pages of Misraj-DocOCR the order reverses. Both results are worth knowing, and we would rather show you the one where Baseer is not first than leave it out.
The other models in the results
- GPT-5 has a WER of 0.86 but a TEDS of 48, the third-best table score. GPT-4o-mini scores 1.36 WER.
- Dots.ocr (0.50 WER) and Nanonets (0.71) are document-OCR specialists. They sit mid-table here and do better on the smaller KITAB-Bench set.
- Qari, an Arabic OCR model, has a WER of 0.76 and the joint-lowest TEDS (21).
- Gemma3-12B (0.96), Gemma3-4B (1.01), AIN (1.23) and Aya-vision (1.41) trail the field. The paper points to a sharp gap between the top systems and smaller or less specialised models on this benchmark.
If a model you are considering is missing, it is because it has no published score on Misraj-DocOCR, and we have not estimated one.
How to choose an Arabic OCR model
- Start from your documents. Take 50 to 100 pages that look like your real workload and score each candidate with the same metrics. The benchmark loads in a few lines with the Hugging Face
datasetslibrary, and you can use it as a template for your own test set. - Decide which metric matches the job. For reports, statements and forms where tables have to survive, weight TEDS and MARS. For plain running text, WER and CER tell you more.
- Check where the data goes. Gemini, GPT-5 and Azure are cloud services. Baseer can run on your own servers, which matters in Saudi government, banking and healthcare, where documents often cannot leave the organisation.
- Look at the output format. Baseer returns Markdown with tables as HTML, which goes straight into search indexes, retrieval pipelines and training sets without a separate layout step.
You can try Baseer on your own pages through the SaaS at baseerocr.com. For on-premises or sovereign deployment, read about Baseer OCR or contact our team.
FAQ
What is the best AI model for Arabic OCR in 2026?
On the published Misraj-DocOCR results, Baseer has the lowest word error rate (0.25) and the best table-structure score (TEDS 66). Gemini 2.5 Pro is the strongest general-purpose model in the comparison and leads on BLEU and ChrF, and Azure AI Document Intelligence has the lowest character error rate.
How accurate is Qwen2.5-VL at Arabic OCR?
On Misraj-DocOCR, Qwen2.5-VL scored a WER of 0.87 (3B), 0.92 (7B) and 0.76 (32B). Baseer, which is Qwen2.5-VL-3B fine-tuned for Arabic documents, scores 0.25.
Is the Misraj-DocOCR benchmark public?
Yes. All 400 pages and their expert-verified reference transcriptions are on Hugging Face under Apache-2.0, and the method is described in the paper on arXiv.
Why do some models have a WER above 1?
WER counts inserted words as errors. A model that repeats itself or writes text that is not on the page can make more errors than there are words in the reference, which pushes the rate past 1.0.
Can Baseer run on-premises?
Yes. Misraj deploys Baseer in the cloud, on-premises or in a sovereign environment inside the Kingdom. Talk to us about your setup.