Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Kawn Lab

Open Source

Models, tools and corpora Misraj publishes openly, and where to find them.

Where we publish

Two accounts hold everything on this page: the Hugging Face organisation for weights, datasets and benchmark sets, and the GitHub organisation for code.

github.com

GitHub

Open datasets

Corpora released so others can train and evaluate on them. Where a set has no public repository yet it is listed without a link, rather than left out.

35B+ tokens

Open Arabic Corpus

35B+ curated Arabic tokens released publicly, growing regularly, for researchers to build on.

Wasm - Structured Arabic Multimodal Corpus

An open pipeline for building structured Arabic multimodal corpora from Common Crawl.

MSDD - Misraj Structured Data Dump

A large-scale Arabic multimodal dataset built with the Wasm pipeline from Common Crawl.

4,758,338 rows

MUDD - Misraj Unstructured Data Dump

A large-scale Arabic plain-text dataset of 4.76M rows, translated from SlimPajama-627B for pre-training.

Sadeed Tashkeela

A large, high-quality Arabic diacritized corpus for training and evaluating diacritization models.

100M captions

Arabic Image Captioning (100M)

A large-scale dataset of 100 million Arabic image captions, generated with the Mutarjim translation model.

KITAB PDF-to-Markdown (Reviewed)

A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic OCR evaluation.

Open benchmarks

Evaluation sets published so a result can be reproduced rather than taken on trust. Each is a dataset on Hugging Face.

Misraj-DocOCR

A curated, expert-verified Arabic OCR benchmark of 400 pages, with ground truth for text and layout fidelity.

SadeedDiac-25

SadeedDiac-25 unifies Modern Standard and Classical Arabic diacritization in one expert-reviewed benchmark.

Tarjama-25

A bidirectional Arabic-English translation benchmark of 5,000 expert-reviewed sentence pairs across five domains.

Code repositories

Every repository in the Misraj GitHub organisation, in catalogue order. Licences are shown only where the host reports one it can identify, and a repository that belongs to a model links to that model's page.

PythonMITkawn.ai

kawn.ai Python SDK

The official Python client for the kawn.ai Models API: embedding and Baseer OCR services, sync and async.

PythonMITllama-index-kawn

LlamaIndex Kawn integration

LlamaIndex wrappers for the Kawn SDK: KawnEmbedding for embeddings, and BaseerReader for OCR.

PLpgSQL

QuranHub API

A REST API over the Holy Quran: editions, translations, tafsir, audio recitations and search.

Jupyter Notebook

Sadeed

The code repository published alongside the Sadeed diacritization model and its paper.

Python

Mutarjim evaluation

The evaluation repository published alongside the Mutarjim translation model and its paper.

Python

Nakba pipeline

Reproduces the Misraj model's result in the Nakba-NLP 2026 Arabic handwriting recognition task.

Jupyter Notebook

Kuwain Arabic cleaner

The Arabic data-cleaning repository published alongside the Kuwain 1.5B model and its paper.

Python

Wasm

The pipeline repository published alongside the Wasm paper on Arabic multimodal corpora.

Get started

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.