Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Publications

Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation

Proposes a multi-teacher knowledge distillation framework that specializes text embedding models in Arabic, and introduces Kawn-Embed-Light, the best-performing Arabic embedding model under one billion parameters.

LaTell 2026Arabic foundation models & open AI

About this paper

Multilingual embedding models include Arabic in their training data but treat it as one language among many. The particulars of Arabic, its morphology, the variation in how it is written, the distance between formal and everyday usage, are averaged away in a space shaped mostly by other languages.

This paper presents a multi-teacher knowledge distillation framework for specializing a text embedding model in Arabic. Instead of distilling from a single teacher, the student learns from several teacher models at once and is trained on Arabic text, so the resulting embedding space is shaped by Arabic rather than adapted to it.

The result is Kawn-Embed-Light, a 300M-parameter Arabic embedding model built for semantic search, retrieval, clustering and RAG. Among Arabic embedding models under one billion parameters it reaches state-of-the-art performance, while staying small enough to run at production speed.

Models in this paper

Get started

Let's talk about what you're trying to build.

Tell us the problem. We'll tell you honestly whether AI is the right answer.