KITAB PDF-to-Markdown (Reviewed)
A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic OCR evaluation.
What is in it
Sixty-two page-level samples, 68.6 MB in Parquet, with two fields per row: the page image, and human-verified Markdown for that page which preserves its structure. It is small on purpose. This is an evaluation set, not training data, and every one of the sixty-two pages has been read by a person against its image.
How it was built
It is a corrected version of the KITAB-Bench PDF-to-Markdown subset. Assessing that subset turned up three kinds of defect that make a fair comparison impossible: hallucinated ground truth, where the reference Markdown contained phrasing absent from the source page. One entry carried an English sentence beginning "You're right - let me write it exactly as it appears in the image"; missing page numbers in the references; and omitted small-font text plainly visible in the image. Each was repaired by hand: hallucinated phrases removed, omitted content restored, page markers added and verified, and minor formatting normalised so the task stays consistent across samples. The original task and schema are unchanged, so the set is a drop-in replacement for the subset it corrects.
Intended uses
Evaluating Arabic document OCR and document-understanding models on page-to-Markdown conversion. Its value is that the ground truth can be trusted: a model is no longer penalised for failing to reproduce text that was never on the page, or rewarded for skipping text the reference also skipped. At sixty-two pages it is a correctness check rather than a leaderboard. For a broader Arabic OCR measurement the Misraj-DocOCR benchmark covers 400 expert-reviewed pages across diverse layouts.
Licence and access
Apache 2.0 and openly downloadable from Hugging Face, with no access request. It builds on KITAB-Bench, whose own terms apply to the underlying material.
Related work
The Baseer paper introduces Misraj's Arabic document-OCR vision-language model and Misraj-DocOCR, the expert-verified benchmark released with it, on which Baseer reaches a word error rate of 0.25 against 0.37 for Gemini 2.5 Pro and 0.44 for Azure AI Document Intelligence. The upstream KITAB-Bench is the benchmark this subset corrects.
Cite this dataset
@misc{hennara2025baseer,
title = {Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR},
author = {Khalil Hennara and Muhammad Hreden and Mohamed Motasim Hamed and Ahmad Bastati and Zeina Aldallal and Sara Chrouf and Safwan AlModhayan},
year = {2025},
eprint = {2509.18174},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2509.18174}
}Dataset facts
- Records
- 62
- Format
- Parquet
- Licence
- apache-2.0
- Access
- Open download
- Released
- Built from
- KITAB-Bench
- Research track
- Arabic foundation models & open AI
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
