Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Datasets

KITAB PDF-to-Markdown (Reviewed)

A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic OCR evaluation.

62 pagesOpen downloadArabic foundation models & open AI

What is in it

Sixty-two page-level samples, 68.6 MB in Parquet, with two fields per row: the page image, and human-verified Markdown for that page which preserves its structure. It is small on purpose. This is an evaluation set, not training data, and every one of the sixty-two pages has been read by a person against its image.

How it was built

It is a corrected version of the KITAB-Bench PDF-to-Markdown subset. Assessing that subset turned up three kinds of defect that make a fair comparison impossible: hallucinated ground truth, where the reference Markdown contained phrasing absent from the source page. One entry carried an English sentence beginning "You're right - let me write it exactly as it appears in the image"; missing page numbers in the references; and omitted small-font text plainly visible in the image. Each was repaired by hand: hallucinated phrases removed, omitted content restored, page markers added and verified, and minor formatting normalised so the task stays consistent across samples. The original task and schema are unchanged, so the set is a drop-in replacement for the subset it corrects.

Intended uses

Evaluating Arabic document OCR and document-understanding models on page-to-Markdown conversion. Its value is that the ground truth can be trusted: a model is no longer penalised for failing to reproduce text that was never on the page, or rewarded for skipping text the reference also skipped. At sixty-two pages it is a correctness check rather than a leaderboard. For a broader Arabic OCR measurement the Misraj-DocOCR benchmark covers 400 expert-reviewed pages across diverse layouts.

Licence and access

Apache 2.0 and openly downloadable from Hugging Face, with no access request. It builds on KITAB-Bench, whose own terms apply to the underlying material.

Related work

The Baseer paper introduces Misraj's Arabic document-OCR vision-language model and Misraj-DocOCR, the expert-verified benchmark released with it, on which Baseer reaches a word error rate of 0.25 against 0.37 for Gemini 2.5 Pro and 0.44 for Azure AI Document Intelligence. The upstream KITAB-Bench is the benchmark this subset corrects.

Cite this dataset

@misc{hennara2025baseer,
  title         = {Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR},
  author        = {Khalil Hennara and Muhammad Hreden and Mohamed Motasim Hamed and Ahmad Bastati and Zeina Aldallal and Sara Chrouf and Safwan AlModhayan},
  year          = {2025},
  eprint        = {2509.18174},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2509.18174}
}

Dataset facts

Records
62
Format
Parquet
Licence
apache-2.0
Access
Open download
Released
Built from
KITAB-Bench
Research track
Arabic foundation models & open AI
Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.