Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Vision-Language

Baseer

Arabic document intelligence: scans to structured Markdown

About Baseer

Baseer is Misraj's vision-language model for Arabic documents. It reads a scanned page, a photo or a PDF and writes it out as structured Markdown, with headings, reading order and tables kept, and it recorded a word error rate of 0.25 on Misraj-DocOCR, the best result among the open-source and commercial systems its paper compares.

Arabic is hard for OCR. The script is cursive, letters change shape with their position, diacritics sit above and below the line, fonts vary widely and the page runs right to left. General multimodal models handle high-resource languages well but lose accuracy on Arabic, and classic OCR returns flat strings that throw the layout away. Baseer was built to close both gaps at once.

How it works

Baseer is a 3-billion-parameter model fine-tuned from Qwen2.5-VL-3B-Instruct.

  • Decoder-only fine-tuning. The vision encoder is frozen and only the language decoder is trained. In the paper's own comparison this beat full fine-tuning, and it keeps the base model's general visual features.
  • Training data. 500,000 page and text pairs. 300,000 are synthetic pages rendered in varied fonts and layouts, and 150,000 of those went through one to three of 29 image transformations, such as motion blur, folding, yellowing and low light. The other 200,000 are real pages from books, magazines, educational material and research papers.
  • Structured output. Tables come back as HTML inside the Markdown, and page numbers, watermarks and images are tagged so they stay out of the running text.

Capabilities

Benchmark results

Measured on Misraj-DocOCR, a public benchmark of 400 expert-verified Arabic pages drawn from books, reports, forms, scholarly pages and complex layouts. Figures are from the Baseer paper (arXiv:2509.18174) and the benchmark's leaderboard.

  • Word error rate: 0.25, against 0.37 for Gemini 2.5 Pro, 0.44 for Azure AI Document Intelligence, 0.50 for Dots.ocr and 0.86 for GPT-5. The base model, Qwen2.5-VL-3B-Instruct, scores 0.87.
  • Table structure (TEDS): 66, against 52 for Gemini 2.5 Pro and 42 for Azure.
  • Layout fidelity (MARS): 76.9, the highest on the board.
  • Where others lead: Azure has the lowest character error rate, 0.27 against Baseer's 0.53, and Gemini 2.5 Pro has the highest BLEU and chrF scores.

Deployment

Baseer runs through the Misraj cloud API or on-premises. It is the engine behind Baseer OCR (try it at baseerocr.com) and Baseer Extract, and it is a built-in node in Seamless API. Developers can call it through Kawn Console, the kawn.ai Python SDK, or BaseerReader in the LlamaIndex integration.

Who it is for

  • Archives, libraries and publishers digitizing Arabic books, records and periodicals.
  • Teams feeding Arabic PDFs into search, RAG or training pipelines that need the structure, not just the text.
  • Organizations that must keep their documents inside their own environment.

FAQ

Is Baseer the same as Baseer OCR? Baseer is the model. Baseer OCR is the product built on it that turns documents into Markdown.

Does it read handwriting? Baseer was built for printed documents. A version adapted to historical handwritten manuscripts, Baseer Nakba, took first place in the NAKBA-NLP 2026 handwriting recognition task with a word error rate of 0.2440.

Can I check the benchmark myself? Yes. Misraj-DocOCR is public on Hugging Face at huggingface.co/datasets/Misraj/Misraj-DocOCR under the Apache 2.0 licence.

At a glance

Type
Vision-Language
Parameters
3B
License
Commercial
Deployment
Cloud APIOn-premises
Research track
Arabic foundation models & open AI

Benchmarks

DatasetMetricResult
Misraj-DocOCRArabic document OCR had no expert-verified benchmark that measured structure, not just characters. Misraj-DocOCR is the evaluation Baseer was measured against, and it is public, so anyone can reproduce the comparison rather than take the number on trust.
Word Error Rate
0.25State of the art

Variants

Baseer Nakba

Handwriting-specialised variant of Baseer

Baseer Nakba is an open-source vision-language model for Arabic handwritten text recognition, optimized for historical manuscripts. It accurately transcribes highly cursive handwritten Arabic from degraded archival documents. It achieved 1st Place at the AR-MS NAKBA-NLP 2026 Arabic Manuscript Understanding Shared Task, establishing a new state of the art on the Nakba OCR benchmark. Open weights are published on Hugging Face.

Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.