Kawn Embed Light
Lightweight general-purpose Arabic embeddings for semantic search, retrieval, clustering, and RAG.
About Kawn Embed Light
Kawn Embed Light is a 300-million-parameter Arabic text embedding model for semantic search, retrieval, clustering and RAG. Among Arabic embedding models under one billion parameters it reaches state-of-the-art performance, while staying small enough to run at production speed.
Multilingual embedding models include Arabic in their training data but treat it as one language among many. Its morphology, the variation in how it is written and the distance between formal and everyday usage get averaged away in a vector space shaped mostly by other languages.
How it works
Kawn Embed Light comes from a multi-teacher knowledge distillation framework described in Misraj's paper Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation (LaTell 2026). Instead of learning from a single teacher, the student model learns from several teacher models at once and is trained on Arabic text, so the embedding space is shaped by Arabic rather than adapted to it. It is the general-purpose member of the Kawn Embed family, tuned for business and general Arabic content.
Capabilities
Deployment
Kawn Embed Light is available through the Misraj cloud API, including Kawn Console, or on-premises. Developers can call Kawn embeddings through the kawn.ai Python SDK, with sync and async clients, or through KawnEmbedding in the LlamaIndex integration. It is also the general-purpose embedding model inside Seamless Enterprise.
Who it is for
- Teams building Arabic semantic search or RAG that need speed as well as quality.
- Classification and clustering of general and business Arabic text.
- Organizations that want their embeddings to run inside their own environment.
FAQ
How does it differ from Kawn Embed Islamic and Kawn Embed Medical? Light is the general-purpose model. The other two are domain models, for Islamic literature and for healthcare text.
Where are the benchmark numbers? The published abstract states the result, state of the art among Arabic embedding models under one billion parameters, but gives no per-benchmark scores, so we do not list figures here.
At a glance
- Type
- Embedding
- Parameters
- 300M
- License
- Open weights
- Deployment
- Cloud APIOn-premises
- Research track
- Arabic foundation models & open AI
Let's talk about what you're trying to build.
Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.
