Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Kawn Lab

Datasets

Arabic corpora built and released by Kawn Lab for pre-training, evaluation and multimodal work.

Every corpus, released or not

2T tokens

Arabic Pre-training Corpus

A 2-trillion-token proprietary Arabic corpus, the largest clean dataset built for Arabic LLM pre-training.

35B+ tokensOpen Source

Open Arabic Corpus

35B+ curated Arabic tokens released publicly, growing regularly, for researchers to build on.

Open Source

Wasm - Structured Arabic Multimodal Corpus

An open pipeline for building structured Arabic multimodal corpora from Common Crawl.

Open Source

MSDD - Misraj Structured Data Dump

A large-scale Arabic multimodal dataset built with the Wasm pipeline from Common Crawl.

4,758,338 rowsOpen Source

MUDD - Misraj Unstructured Data Dump

A large-scale Arabic plain-text dataset of 4.76M rows, translated from SlimPajama-627B for pre-training.

Open Source

Sadeed Tashkeela

A large, high-quality Arabic diacritized corpus for training and evaluating diacritization models.

100M captionsOpen Source

Arabic Image Captioning (100M)

A large-scale dataset of 100 million Arabic image captions, generated with the Mutarjim translation model.

Open Source

KITAB PDF-to-Markdown (Reviewed)

A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic OCR evaluation.

Get started

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.