Kuwain
A 1.5B-parameter Arabic SLM built by language injection, the base layer for our downstream models.
About Kuwain
Kuwain is a 1.5-billion-parameter Arabic-English small language model that Misraj built by injecting Arabic into TinyLlama, a compact model trained mainly on English. The method raised Arabic performance by an average of 8% across benchmarks, kept the original English ability intact, and cut training cost by 70%.
Kuwain is the base layer for several of Kawn Lab's downstream models, including Mutarjim for Arabic-English translation and Sadeed for diacritization. Lahjawi, the cross-dialect translator, is also a fine-tuned version of Kuwain.
The problem it solves
Most open language models are built for English. The tokenizers in models such as Llama 2 and Mistral carry only 28 Arabic tokens, one per letter, which makes Arabic slow and costly to process. The usual fixes have their own price: training a bilingual model from scratch takes enormous data and compute, and continuing to train an English model on Arabic tends to erode what it already knew. Kuwain was built to add Arabic without paying either cost.
Capabilities
How language injection works
- Keep the original model. TinyLlama's existing 1.1 billion parameters stay frozen, so its English knowledge is not overwritten.
- Add new layers. Eight new layers are spread across the model rather than stacked together, which grows it by about 30%. Only the new layers and the final layer are trained; the paper found that stacking the new layers or freezing the last one made training unstable.
- Extend the vocabulary. A 26,000-token Arabic tokenizer is merged into the original one, for 54,000 tokens in total.
- Train on a lean mix. 110 billion tokens from public sources: 90 billion in Arabic, including some dialect data, and 20 billion in English. About 20% English was enough to hold English performance, against the 50% used in comparable work. Training ran on 8 A100 GPUs for 3 epochs.
The Arabic cleaning script used to prepare the data is released on GitHub as misraj-ai/Kuwain-Arabic-cleaner.
Benchmark results
All figures are from the Kuwain paper (arXiv:2504.15120).
- English preserved: an average of 53.28 across seven English benchmarks, against 52.99 for the original TinyLlama. With less than 20% English data the average fell to 49.56.
- Against conventional training: the same data without the new layers (Kuwain-Naive) reached similar Arabic scores, 42.17 against Kuwain's 42.27, but its English average dropped to 46.85.
- Arabic gained: an average improvement of 8% across Arabic benchmarks over the base model.
- Tokenizer efficiency: an expansion ratio of 2.30 with 26,000 Arabic tokens, against 2.51 for AraBERT with 54,000 and 2.19 for Jais with 44,000.
Deployment
Kuwain is available through a cloud API or on-premises. At 1.5B parameters it suits environments where compute is limited.
Who it is for
- Teams that need one small model to handle both Arabic and English.
- Developers building task-specific Arabic models on a compact base, the way Mutarjim, Sadeed and Lahjawi were built.
- Researchers extending an English-centric model to a new language without retraining it from scratch.
FAQ
Does adding Arabic hurt Kuwain's English? No. Its English average is 53.28, against 52.99 for the model it started from.
Why not train an Arabic model from scratch? It takes far more data and compute. Kuwain's approach cut training cost by 70%.
Which Misraj models are built on Kuwain? Mutarjim for translation, Sadeed for diacritization and Lahjawi for cross-dialect translation.
Where is the paper? "Kuwain 1.5B: An Arabic SLM via Language Injection", arXiv:2504.15120, also presented at LaTell 2026.
At a glance
- Type
- Language
- Parameters
- 1.5B
- License
- Open weights
- Deployment
- Cloud APIOn-premises
- Research track
- Arabic foundation models & open AI
- Links
- CodeRead the paper
Papers
arXiv (2025), LaTell 2026
Kuwain 1.5B: An Arabic SLM via Language Injection
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
