Build on Arabic
models we trained ourselves
Most of what an Arabic AI project needs already exists in our catalogue. What is left over is the part that is specific to you: your documents, your dialects, your regulator, your infrastructure. That is the work this page describes.
How an engagement starts
Most projects are one of these, or two of them in sequence.
Domain adaptation of a catalogue model
Start from a trained Arabic model, not from zero.
Kawn Embed ships in three forms rather than one because a general Arabic retrieval model was not accurate enough on jurisprudential or clinical vocabulary. The same route is open to your domain: we take the model in the catalogue closest to your task and specialise it on your material, so you inherit the Arabic the base model already learned instead of paying to rebuild it.
An evaluation set for your task
The measure that decides whether it worked.
Kawn Lab has published 3 benchmarks that did not exist until it built them: Misraj-DocOCR for Arabic document OCR, SadeedDiac-25 for diacritization, and Tarjama-25 for Arabic to English translation. Each was built because the available evaluations were too narrow to settle the question. We do the same inside an engagement: an expert-reviewed test set drawn from your own material and agreed before any model is touched, so what arrives at the end is a number you can check rather than a demonstration.
A system composed from the catalogue
Several models, one workflow, your systems at both ends.
Seamless Enterprise, Workforces and Sada are each built by putting several of our models behind a single workflow rather than by training one model to do everything. An engagement can produce the same thing for a process we do not already ship: document intelligence feeding retrieval, speech feeding extraction, extraction feeding an approval step, wired into the systems the work already runs in.
Deployment inside your boundary
The same models, running where your data has to stay.
Models in the catalogue are offered on-premises as well as through the cloud API, because the sectors we work in cannot always send documents to a third party. An engagement of this kind is an infrastructure project rather than a modelling one: sizing, installation, integration and handover inside your own environment, against whatever your regulator requires.
How the work runs
The same sequence Kawn Lab follows on its own models: agree the measure, build the evaluation, adapt what already exists, deploy where the constraint says, and hand the measure over.
Agree what would count as working
The first output is not a prototype. It is a written statement of the task, the material it runs on, and the measure that will decide it, because a project with no agreed measure ends in an argument about screenshots.
Build the test set before the model
The evaluation material is assembled and reviewed first, out of your own documents, recordings or records. Building it afterwards means grading the work with a ruler the work has already seen.
Adapt what already exists
Wherever the catalogue can carry the task, the work starts from a trained model rather than a blank one. Sadeed and Mutarjim are both built on Kuwain, our own compact Arabic language model, which is why each is small enough to run in production.
Deploy where the constraint says
Cloud, on-premises or inside your own boundary is a decision taken from your data residency and regulatory position, not from our convenience. The Trust Center sets out what each option involves.
Hand over the evaluation with the system
You keep the test set and the results alongside the system itself. That is what lets your own team re-run the measure after your data changes, and what stops the number ageing quietly.
What we commit to
Stated plainly here, with the detail in the Trust Center.
Your data stays where you put it
On-premises deployment exists so that documents, recordings and records never have to leave the environment they are governed in. Where a project does run in the cloud, data residency is settled and recorded up front rather than discovered later.
You own what the engagement produces
Data ownership is one of the things our architecture is built around, alongside Arabic at the core and fit to the sector. The material you bring stays yours, and so do the evaluation set and the results built from it.
Compliance is designed in, not retrofitted
Sovereign deployment options and regulatory alignment are part of the architecture rather than something added once a security review flags a gap. In government, finance, legal and healthcare, where the constraint is not optional, it is part of the design brief from the first conversation.
We publish what we can prove
10 research papers, 3 benchmarks we built ourselves and open model weights are the reason a claim on this site can be checked. Inside an engagement the same rule holds: a result we cannot measure is not reported as one, and no certification is claimed here that the Trust Center does not list.
Tell us what you are trying to build
The more concretely you can describe the task and the material it runs on, the more useful the first conversation is. We come back with what in the catalogue is closest to it, and what would have to be built.
