Wasm - Structured Arabic Multimodal Corpus
An open pipeline for building structured Arabic multimodal corpora from Common Crawl.
What is in it
Wasm is the pipeline, not a fixed archive: an open processing chain that turns Common Crawl into structured Arabic multimodal documents, published with a representative dump of its own output. Each document comes out as Markdown that keeps the distinction between headings, paragraphs, ordered and unordered lists and tables, with images left in the position they held on the page rather than stripped into separate image-text pairs, and with metadata for the source, the Common Crawl snapshot the page came from, its language and its domain.
How it was built
Wasm starts from the OBELICS framework and is adapted for Arabic. Because Arabic is only about 0.6 per cent of Common Crawl, the pipeline filters the crawl index before downloading anything, keeping the pages whose metadata indicates Arabic and whose fetch succeeded. The retrieved HTML is cleaned in two passes, converted to Markdown, then filtered twice, once per text block and once per whole document, so an advertisement or a navigation menu cannot condemn the article around it. Quality rules were recalibrated for Arabic: the weight on repeated-word ratios was reduced and the stopword, punctuation and common-word filters removed, because rhetorical repetition, few function words and inconsistent punctuation are normal in Arabic and are not evidence of a bad page. A KenLM model trained on Arabic across dialects and topics scores each unit's perplexity, and the Needleman-Wunsch algorithm removes near-duplicate nodes above an 80 per cent similarity threshold. The repeated node goes, the rest of the document stays.
Intended uses
Building Arabic pre-training corpora, in either of two shapes: text-only, or interleaved multimodal, from the same output. The paper's argument for the second is that multimodal models trained on natural documents where images and text interleave beat models trained on image-text pairs across a wide range of benchmarks, and that Arabic had no corpus preserving document structure to train that way on. The pipeline is also meant to be read and argued with: it ships with a comparative analysis of its filtering choices against those of the major existing corpora.
Licence and access
The pipeline is open source on GitHub under Apache 2.0, and a representative dataset dump is released with it. Image links are preserved rather than downloaded during extraction, so a user of the dump fetches images themselves and is bound by whatever terms the originating sites set.
Related work
MSDD is the published dump of this pipeline's output. The Wasm paper is on arXiv and was accepted at LaTell 2026. Baseer's training data drew on a sample of Wasm's output, and the same pipeline is the web stage of Misraj's two-trillion-token Arabic pre-training corpus.
Cite this dataset
@misc{hennara2025wasm,
title = {Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora},
author = {Khalil Hennara and Ahmad Bastati and Muhammad Hreden and Mohamed Motasim Hamed and Zeina Aldallal and Sara Chrouf and Safwan AlModhayan},
year = {2025},
eprint = {2511.07080},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2511.07080}
}Dataset facts
- Format
- Markdown
- Licence
- apache-2.0
- Access
- Open download
- Built from
- Common Crawl
- Research track
- Arabic foundation models & open AI
- Links
- Read the paper
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
