Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Kawn Lab

Datasets

Arabic corpora for pretraining, evaluation and multimodal work. Built by us. Shared with everyone.

Every corpus, released or not

2T tokens

Arabic Pre-training Corpus

A 2-trillion-token proprietary Arabic corpus, the largest clean dataset built for Arabic LLM pre-training.

35B+ tokensOpen Source

Open Arabic Corpus

35B+ curated Arabic tokens released publicly, growing regularly, for researchers to build on.

Open Source

Wasm - Structured Arabic Multimodal Corpus

An open pipeline for building structured Arabic multimodal corpora from Common Crawl.

23M documentsOpen Source

MSDD - Misraj Structured Data Dump

A large-scale Arabic multimodal dataset built with the Wasm pipeline from Common Crawl.

4,758,338 rowsOpen Source

MUDD - Misraj Unstructured Data Dump

A large-scale Arabic plain-text dataset of 4.76M rows, translated from SlimPajama-627B for pre-training.

1,045,183 rowsOpen Source

Sadeed Tashkeela

A large, high-quality Arabic diacritized corpus for training and evaluating diacritization models.

100M captionsOpen Source

Arabic Image Captioning (100M)

A large-scale dataset of 100 million Arabic image captions, generated with the Mutarjim translation model.

62 pagesOpen Source

KITAB PDF-to-Markdown (Reviewed)

A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic OCR evaluation.

Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.