MSDD - Misraj Structured Data Dump
A large-scale Arabic multimodal dataset built with the Wasm pipeline from Common Crawl.
What is in it
Twenty-three million Arabic documents drawn from several Common Crawl snapshots, 63.4 GB in 361 Parquet shards. Each sample is Markdown text with the image marker left inline, an ordered list of the image URLs, and the list of their captions, so the position of every picture within the article survives, which is the whole point of the format. The images themselves are not included; the dump carries their links.
How it was built
MSDD is the published output of the Wasm pipeline, and inherits its decisions end to end: index-level filtering before download, two-pass HTML cleaning, conversion to structured Markdown, quality filters applied at both the text-block and document level and recalibrated for Arabic, perplexity scoring by a KenLM model trained on Arabic, and node-level near-duplicate removal at an 80 per cent Needleman-Wunsch threshold. Each crawl snapshot was processed independently, which made batches restartable and kept the time provenance of every document traceable.
Intended uses
Pre-training Arabic language models on text only, or Arabic multimodal models on interleaved image-text documents, from the same dump. It is a sample of one pipeline's output and is not a slice of Misraj's full two-trillion-token repository, so it should be described as web-derived Arabic data and not as a representative cross-section of Arabic knowledge: web text skews towards news, short forms and material written for search engines.
Licence and access
Apache 2.0, listed publicly on Hugging Face with access granted on request. The listing, the card and the file manifest are readable by anyone, and the Parquet shards are handed out after an access request. Because image links are carried rather than image files, anyone training on the multimodal form fetches the pictures themselves under the originating sites' terms.
Related work
The Wasm paper documents the pipeline that produced this dump and the comparison against other Arabic corpora. MUDD is its sibling from the translation pipeline. Baseer's training data drew on a sample of this same web output.
Cite this dataset
@misc{hennara2025wasm,
title = {Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora},
author = {Khalil Hennara and Ahmad Bastati and Muhammad Hreden and Mohamed Motasim Hamed and Zeina Aldallal and Sara Chrouf and Safwan AlModhayan},
year = {2025},
eprint = {2511.07080},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2511.07080}
}Dataset facts
- Size
- 23M documents
- Format
- Parquet
- Licence
- apache-2.0
- Access
- Public listing, access on request
- Released
- Built from
- Common Crawl
- Research track
- Arabic foundation models & open AI
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
