Lahjawi
A family of Arabic dialect translation models covering dialect-to-dialect and dialect-to-MSA translation.
About Lahjawi
Lahjawi is a family of Arabic dialect translation models from Misraj. Lahjawi-D2D translates directly between 15 Arabic dialects, the first model built for that task according to its paper, and Lahjawi-D2MSA converts any of those dialects into Modern Standard Arabic.
Arabs speak dialects day to day, and the dialects differ in vocabulary, sentence structure, verb conjugation, plural forms and idioms such as folk proverbs. Most dialect translation research converts one dialect, or several, into MSA. Direct dialect-to-dialect translation had one earlier foundational work, the PADIC corpus, and no dedicated model.
How it works
Both models are fine-tuned from Kuwain-1.5B, Misraj's Arabic small language model.
- Training data. Seven open datasets (MADAR, PADIC, NADI 2023, Dial2MSA, Arabic STS, the UFAL North Levantine corpus and MDPCA), normalized for characters, numerals, punctuation and spacing.
- Two datasets built from them. A dialect-to-MSA set of 197,042 samples, and a dialect-to-dialect set of 266,871 samples built from every combination of dialect pairs, 210 in all.
- Translation as question answering. Each sample is wrapped in a prompt template: one for any dialect into MSA, and one that names the source and target dialects. Training uses next-token prediction with the system prompt masked.
One finding shaped the design: four separate single-dialect models did worse overall than one model trained on all four dialects together, which scored 13.55 BLEU against 12.13 on the NADI 2024 test set.
Capabilities
Benchmark results
Figures are from the Lahjawi paper, published in the proceedings of WACL-4 at COLING 2025.
- Dialect to MSA, MADAR test set: overall BLEU 9.62 across 15 dialects, from 11.52 for Jordanian and 10.81 for Saudi down to 6.47 for Tunisian.
- Dialect to dialect, MADAR test set: overall BLEU 9.88.
- Human evaluation: 58% accuracy and 78% fluency, scored by people on 50 sentences for the most widely spoken dialects, including Syrian, Jordanian, Palestinian, Tunisian, Egyptian, Saudi and Moroccan.
- NADI 2024 dialect-to-MSA task: 13.30 BLEU for Lahjawi-D2MSA. Teams using much larger models, such as Jais-13B and Llama 3 8B, or bigger augmented datasets scored higher, and the top system reached 20.44. We report that result as it stands.
Quality follows the training data. Levantine dialects, which dominate it, translate best, and Maghrebi dialects, Tunisian above all, are hardest. Translation is not always symmetric either: Qatari to Iraqi scores 17.21 BLEU, Iraqi to Qatari 6.02.
Deployment
Lahjawi is available through the Misraj cloud API or on-premises.
Who it is for
- Services that need to reach people across the Arab world in their own dialect.
- Teams that need dialect text normalized into MSA before search, analysis or further translation.
- Researchers working on Arabic dialects and cross-dialect translation.
FAQ
Which dialects does it cover? Fifteen: Saudi, Omani, Qatari, Iraqi, Jordanian, Lebanese, Palestinian, Syrian, Egyptian, Sudanese, Yemeni, Algerian, Libyan, Moroccan and Tunisian.
Does it handle Saudi and Gulf dialects? Yes. Saudi, Qatari and Omani are among the 15, and Saudi scored 10.81 BLEU into MSA, above the overall average.
What are its limits? Dialect data is small and uneven, many reference translations are paraphrases rather than literal, and a small model can produce inaccurate output. The paper names all three.
At a glance
- Type
- Language
- License
- Open weights
- Deployment
- Cloud APIOn-premises
- Research track
- Arabic foundation models & open AI
Papers
COLING Workshops (2025)
Lahjawi: Arabic Cross-Dialect Translator
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
