Arabic Pre-training Corpus
A 2-trillion-token proprietary Arabic corpus, the largest clean dataset built for Arabic LLM pre-training.
We understand deeply. We build from the source. And we run with confidence.
Arabic corpora built and released by Kawn Lab for pre-training, evaluation and multimodal work.
A 2-trillion-token proprietary Arabic corpus, the largest clean dataset built for Arabic LLM pre-training.
35B+ curated Arabic tokens released publicly, growing regularly, for researchers to build on.
An open pipeline for building structured Arabic multimodal corpora from Common Crawl.
A large-scale Arabic multimodal dataset built with the Wasm pipeline from Common Crawl.
A large-scale Arabic plain-text dataset of 4.76M rows, translated from SlimPajama-627B for pre-training.
A large, high-quality Arabic diacritized corpus for training and evaluating diacritization models.
A large-scale dataset of 100 million Arabic image captions, generated with the Mutarjim translation model.
A reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic OCR evaluation.
Tell us the problem. We'll tell you honestly whether AI is the right answer.