Benchmarks
Tarjama-25
Beats GPT-4o mini: En→Ar
State of the artArabic foundation models & open AI
A bidirectional Arabic-English translation benchmark of 5,000 expert-reviewed sentence pairs across five domains.
Why it matters
Most public Arabic-English test sets are English-source, short and narrow in domain, which flatters models that translate out of English and hides how they behave in the other direction. Tarjama-25 was released publicly so Arabic-English translation results can be compared on the same terms.
Get started
Let's talk about what you're trying to build.
Tell us the problem. We'll tell you honestly whether AI is the right answer.
