Arabic Pre-training Corpus
A 2-trillion-token proprietary Arabic corpus, the largest clean dataset built for Arabic LLM pre-training.
What is in it
About two trillion Arabic tokens, held as an internal repository rather than a public release. Three sources feed it: Arabic web pages retrieved and cleaned with the Wasm pipeline, high-value knowledge translated into Arabic with Mutarjim, and text recovered from books, periodicals and academic PDFs with Baseer. Every record carries the same fields: a persistent document identifier, the source type, the structured text and its visual elements, language and translation direction, domain, quality indicators, the extraction method and model version, the original URL and crawl timestamp, and licensing information.
How it was built
Each source has its own pipeline, because a web page, a translated passage and a scanned book fail in different ways and one tool across all three would have hidden which stage introduced an error. Wasm decomposes the DOM, scores every text unit with a KenLM language model trained on Arabic across dialects and topics, and removes near-duplicate nodes above an 80 per cent similarity threshold rather than discarding the whole page. Mutarjim translates only where a domain is thin in native Arabic, and every pair is checked for target language, length proportion, duplication and terminology. Baseer treats each page as a structured image and writes Markdown, with tables in HTML so merged cells survive. The three streams are then normalised and deduplicated against one another, so the same article cannot enter once from the web, again from a PDF and a third time as a republished translation.
Intended uses
Pre-training and continued pre-training of Arabic-aware language and vision-language models. The mixture is weighted by the effect each layer should have rather than by raw volume: original Arabic stays the foundation, translation fills domain gaps without imposing a translated style on the model, and books and research raise knowledge density and context length. Balance is monitored across source, domain, time period, Modern Standard Arabic against dialects, document length, share of translated text and OCR quality, and an independent validation set is kept for each segment so bias or degradation shows up per source rather than disappearing into an average.
Licence and access
This corpus is proprietary. It is not released, there is no public download, and there is no licence to quote. What is public is the machinery and one sample of its output: the Wasm pipeline on GitHub, the Arabic data-cleaning repository published alongside Kuwain 1.5B, and MSDD, a published dump of the web stage. MSDD is a sample of one pipeline, not a slice of this repository.
Related work
The three pipelines behind this corpus each have their own published record: the Wasm paper for the web stage, the Mutarjim paper for the translation stage, and the Baseer paper for document extraction. The Open Arabic Corpus is the separately curated set Misraj does release. MSDD is the public dump of Wasm's output.
Dataset facts
- Size
- 2T tokens
- Access
- Not released
- Built from
- Common CrawlTranslated English sourcesBooks and research papers
- Research track
- Arabic foundation models & open AI
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
