Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Language

Mutarjim

Bidirectional Arabic–English translation built on Kuwain, for general and specialized domains.

1.5BOpen weights

About Mutarjim

Mutarjim is a 1.5-billion-parameter language model built for one job: translating between Arabic and English in both directions. On Tarjama-25, Misraj's public translation benchmark, it is the strongest system tested for English-to-Arabic translation, ahead of GPT-4o mini and of open models up to 32 billion parameters.

Most Arabic-English translation today comes either from large multilingual models, which are costly to run and often weaker on Arabic-specific text, or from systems measured on test sets of short, English-source sentences. Mutarjim takes the other route: a small model trained only on Arabic-English data and evaluated on text originally written in both languages. Its paper reports that it rivals models up to 20 times its size while needing far less compute to train and run.

How it works

Mutarjim starts from Kuwain-1.5B, Misraj's Arabic-English small language model, and adds two training stages:

  • Translation pre-training. About 10 billion tokens of Arabic-English parallel text, drawn from the OPUS collection and internally curated sets. Two special tokens mark each language, and the order of every sentence pair is shuffled so neither direction is favoured.
  • Fine-tuning. About 6 million filtered sentence pairs, around 3 billion tokens over two epochs, with Arabic-source examples weighted two to one over English-source ones. The model is trained only on producing the target sentence.

The pre-training stage matters: on WMT24++ it lifted the English-to-Arabic model's COMET score from 61.91 to 75.46.

Capabilities

Benchmark results

All figures are from the Mutarjim paper (arXiv:2505.17894).

  • Tarjama-25, English to Arabic: COMET 83.41, chrF++ 68.67, BLEU 43.71. GPT-4o mini scores 83.36, 66.36 and 38.52, and the strongest open model in the comparison, Cohere 32B, scores 82.09 COMET.
  • Tarjama-25, Arabic to English: COMET 82.63 and BLEU 55.28, against 83.67 and 54.24 for GPT-4o mini.
  • WMT24++ and IWSLT 2017: on these established test sets, GPT-4o mini and several larger open models score higher than Mutarjim, mostly from Arabic to English. We report that alongside the wins.

Tarjama-25 holds 5,000 sentence pairs, each 50 to 100 words long, corrected by professional translators and then reviewed by domain experts. Half of the source texts were written in Arabic and half in English, across general, medical, legal, technical and other fields. The benchmark is public on Hugging Face and the evaluation code is on GitHub.

Deployment

Mutarjim is available through a cloud API or on-premises, inside your own environment. At 1.5B parameters it fits resource-constrained settings where a model many times its size would not. The paper also trained one-direction variants: where only one direction is needed they score higher, for example 82.89 COMET from Arabic to English on WMT24++ against 79.73 for the bidirectional model.

Who it is for

  • Organizations translating Arabic-source material, not only English into Arabic, across legal, healthcare, scientific, technical, cultural and religious content.
  • Teams that need translation to run inside their own infrastructure.
  • Researchers comparing Arabic-English systems on a shared, balanced benchmark.

FAQ

Does Mutarjim beat GPT-4o mini? On Tarjama-25 from English to Arabic, yes, on COMET, chrF++ and BLEU. From Arabic to English it is close on COMET and ahead on BLEU. On WMT24++ and IWSLT 2017, GPT-4o mini scores higher.

What does "built on Kuwain" mean? Mutarjim starts from the weights of Kuwain-1.5B and adds the two translation training stages described above.

Can it run on-premises? Yes. On-premises deployment is supported alongside the cloud API.

Where can I check the results? The paper is arXiv:2505.17894. Tarjama-25 is published at huggingface.co/datasets/Misraj/Tarjama-25 and the evaluation toolkit at github.com/misraj-ai/Mutarjim-evaluation.

At a glance

Type
Language
Parameters
1.5B
License
Open weights
Deployment
Cloud APIOn-premises
Research track
Arabic foundation models & open AI

Benchmarks

DatasetMetricResult
Tarjama-25Most public Arabic-English test sets are English-source, short and narrow in domain, which flatters models that translate out of English and hides how they behave in the other direction. Tarjama-25 was released publicly so Arabic-English translation results can be compared on the same terms.
Beats GPT-4o mini
En→ArState of the art
Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.