Datasets
Arabic Pre-training Corpus
2T tokensArabic foundation models & open AI
A 2-trillion-token proprietary Arabic corpus, the largest clean dataset built for Arabic LLM pre-training.
Get started
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.
