Open Arabic Corpus
35B+ curated Arabic tokens released publicly, growing regularly, for researchers to build on.
What is in it
More than 35 billion curated Arabic tokens, released publicly and added to regularly. It is the released counterpart to Misraj's internal two-trillion-token repository: the same curation standards, published so that a laboratory without a crawling and cleaning stack of its own can still pre-train or continue pre-training on Arabic text that has been through one. Every record keeps the same provenance fields the internal repository uses: a persistent document identifier, the source type, the domain, quality indicators, the extraction method and model version, the original URL and the crawl timestamp, so a team that later decides one source was weak can exclude it and re-measure rather than having to discard the whole mixture.
How it was built
The tokens come through the same three-stage machinery Misraj uses for its own pre-training data (the Wasm pipeline for Arabic web content, Mutarjim for knowledge translated into Arabic, and Baseer for text recovered from documents) and through the same quality gates. Filtering is applied at both the text-block and whole-document level so that an advertisement or a repeated navigation menu does not cause an otherwise valuable article to be thrown away; quality rules are calibrated for Arabic rather than borrowed from English pipelines, where lexical repetition, few function words and sparse punctuation are treated as defects they are not in Arabic.
Intended uses
Pre-training and continued pre-training of Arabic language models, and any downstream work that needs volume of clean Arabic text: tokeniser training, domain adaptation, retrieval corpora. It is a general corpus, not a task set: it carries no labels, no instruction pairs and no evaluation split, so a result measured on it is a language-modelling result and nothing more. Keeping the provenance fields is what makes it usable for more than one of those at once: a retrieval corpus wants the URL and the crawl date, a domain adaptation run wants the domain field, and a tokeniser wants none of them.
Licence and access
Publicly released and growing. Individual dumps are published through Misraj's Hugging Face organisation as they are prepared, which is why this page carries no single download link: the set is a running release rather than one frozen archive.
Related work
MSDD and MUDD are the two named dumps that have been published so far, the first from the web pipeline and the second from the translation pipeline. The two-trillion-token Arabic pre-training corpus is the internal repository this is drawn from, and is not released.
Dataset facts
- Size
- 35B+ tokens
- Access
- Open download
- Research track
- Arabic foundation models & open AI
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
