Benchmarks
Tarjama-25
Beats GPT-4o mini: En→Ar
State of the artArabic foundation models & open AI
A bidirectional Arabic-English translation benchmark of 5,000 expert-reviewed sentence pairs across five domains.
Why it matters
Most public Arabic-English test sets are English-source, short and narrow in domain, which flatters models that translate out of English and hides how they behave in the other direction. Tarjama-25 was released publicly so Arabic-English translation results can be compared on the same terms.
Get started
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
