MUDD - Misraj Unstructured Data Dump
A large-scale Arabic plain-text dataset of 4.76M rows, translated from SlimPajama-627B for pre-training.
What is in it
4,758,338 rows of Arabic plain text, in 64 Parquet shards. Unlike MSDD it carries no images and no document structure: this is the translation arm of the pre-training data, Arabic rendered from English source material that had no Arabic counterpart worth crawling. The English side is SlimPajama-627B, the deduplicated open pre-training corpus.
How it was built
Translation here is a coverage decision, not a volume exercise. The stage begins with a gap map (which domains are thinly represented in native Arabic web data) and selects high-value source material across technology, science, healthcare, law and culture. Each segment keeps its source identifier, original language, domain, licence and retrieval date. The translation is done by Mutarjim, a 1.5-billion-parameter decoder-only model built on Kuwain 1.5B, continued-pretrained on roughly ten billion tokens of Arabic-English pairs and then fine-tuned on about six million high-quality parallel pairs, weighted two to one in favour of Arabic-origin samples so the Arabic generator learns Arabic rhythm rather than imitating English structure. Output is accepted only after it clears language verification, a length-proportion check that catches truncation and padding, deduplication, and terminology validation.
Intended uses
Pre-training and continued pre-training where the goal is to add knowledge Arabic does not already carry: the domains a crawl of the Arabic web will not fill. It is machine-translated text and should be labelled as such in any mixture: the quality gates are there to hold down translation noise, not to make the output indistinguishable from originally-Arabic writing, and a model trained on too high a share of it will acquire a translated style.
Licence and access
Listed publicly on Hugging Face with access granted on request. The dataset card states no SPDX licence of its own, so the terms of the underlying SlimPajama-627B material and of the request itself are what govern use; the licence field on this page is deliberately empty rather than guessed.
Related work
The Mutarjim paper documents the translation model, its two-phase training and the Tarjama-25 benchmark released with it; the Kuwain paper documents the base model Mutarjim was built on. MSDD is the sibling dump from the web pipeline.
Cite this dataset
@misc{hennara2025mutarjim,
title = {Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model},
author = {Khalil Hennara and Muhammad Hreden and Mohamed Motaism Hamed and Zeina Aldallal and Sara Chrouf and Safwan AlModhayan},
year = {2025},
eprint = {2505.17894},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2505.17894}
}Dataset facts
- Records
- 4,758,338
- Format
- Parquet
- Access
- Public listing, access on request
- Released
- Built from
- SlimPajama-627B
- Research track
- Arabic foundation models & open AI
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
