Open Source
Models, tools and corpora Misraj publishes openly, and where to find them.
Where we publish
Two accounts hold everything on this page: the Hugging Face organisation for weights, datasets and benchmark sets, and the GitHub organisation for code.
GitHub
Open datasets
Corpora released so others can train and evaluate on them. Where a set has no public repository yet it is listed without a link, rather than left out.
Open Arabic Corpus
35B+ curated Arabic tokens released publicly, growing regularly, for researchers to build on.
Wasm - Structured Arabic Multimodal Corpus
An open pipeline for building structured Arabic multimodal corpora from Common Crawl.
MSDD - Misraj Structured Data Dump
A large-scale Arabic multimodal dataset built with the Wasm pipeline from Common Crawl.
MUDD - Misraj Unstructured Data Dump
A large-scale Arabic plain-text dataset of 4.76M rows, translated from SlimPajama-627B for pre-training.
Sadeed Tashkeela
A large, high-quality Arabic diacritized corpus for training and evaluating diacritization models.
Arabic Image Captioning (100M)
A large-scale dataset of 100 million Arabic image captions, generated with the Mutarjim translation model.
KITAB PDF-to-Markdown (Reviewed)
A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic OCR evaluation.
Open benchmarks
Evaluation sets published so a result can be reproduced rather than taken on trust. Each is a dataset on Hugging Face.
Misraj-DocOCR
A curated, expert-verified Arabic OCR benchmark of 400 pages, with ground truth for text and layout fidelity.
SadeedDiac-25
SadeedDiac-25 unifies Modern Standard and Classical Arabic diacritization in one expert-reviewed benchmark.
Tarjama-25
A bidirectional Arabic-English translation benchmark of 5,000 expert-reviewed sentence pairs across five domains.
Code repositories
Every repository in the Misraj GitHub organisation, in catalogue order. Licences are shown only where the host reports one it can identify, and a repository that belongs to a model links to that model's page.
kawn.ai Python SDK
The official Python client for the kawn.ai Models API: embedding and Baseer OCR services, sync and async.
LlamaIndex Kawn integration
LlamaIndex wrappers for the Kawn SDK: KawnEmbedding for embeddings, and BaseerReader for OCR.
QuranHub API
A REST API over the Holy Quran: editions, translations, tafsir, audio recitations and search.
Sadeed
The code repository published alongside the Sadeed diacritization model and its paper.
Mutarjim evaluation
The evaluation repository published alongside the Mutarjim translation model and its paper.
Nakba pipeline
Reproduces the Misraj model's result in the Nakba-NLP 2026 Arabic handwriting recognition task.
Kuwain Arabic cleaner
The Arabic data-cleaning repository published alongside the Kuwain 1.5B model and its paper.
Wasm
The pipeline repository published alongside the Wasm paper on Arabic multimodal corpora.
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.
