Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Build on Arabic models we trained ourselves

Most of what an Arabic AI project needs already exists in our catalogue. What is left over is the part that is specific to you: your documents, your dialects, your regulator, your infrastructure. That is the work this page describes.

How an engagement starts

Most projects are one of these, or two of them in sequence.

  • Domain adaptation of a catalogue model

    Start from a trained Arabic model, not from zero.

    Kawn Embed ships in three forms rather than one because a general Arabic retrieval model was not accurate enough on jurisprudential or clinical vocabulary. The same route is open to your domain: we take the model in the catalogue closest to your task and specialise it on your material, so you inherit the Arabic the base model already learned instead of paying to rebuild it.

  • An evaluation set for your task

    The measure that decides whether it worked.

    Kawn Lab has published 3 benchmarks that did not exist until it built them: Misraj-DocOCR for Arabic document OCR, SadeedDiac-25 for diacritization, and Tarjama-25 for Arabic to English translation. Each was built because the available evaluations were too narrow to settle the question. We do the same inside an engagement: an expert-reviewed test set drawn from your own material and agreed before any model is touched, so what arrives at the end is a number you can check rather than a demonstration.

  • A system composed from the catalogue

    Several models, one workflow, your systems at both ends.

    Seamless Enterprise, Workforces and Sada are each built by putting several of our models behind a single workflow rather than by training one model to do everything. An engagement can produce the same thing for a process we do not already ship: document intelligence feeding retrieval, speech feeding extraction, extraction feeding an approval step, wired into the systems the work already runs in.

  • Deployment inside your boundary

    The same models, running where your data has to stay.

    Models in the catalogue are offered on-premises as well as through the cloud API, because the sectors we work in cannot always send documents to a third party. An engagement of this kind is an infrastructure project rather than a modelling one: sizing, installation, integration and handover inside your own environment, against whatever your regulator requires.

How the work runs

The same sequence Kawn Lab follows on its own models: agree the measure, build the evaluation, adapt what already exists, deploy where the constraint says, and hand the measure over.

  • Agree what would count as working

    The first output is not a prototype. It is a written statement of the task, the material it runs on, and the measure that will decide it, because a project with no agreed measure ends in an argument about screenshots.

  • Build the test set before the model

    The evaluation material is assembled and reviewed first, out of your own documents, recordings or records. Building it afterwards means grading the work with a ruler the work has already seen.

  • Adapt what already exists

    Wherever the catalogue can carry the task, the work starts from a trained model rather than a blank one. Sadeed and Mutarjim are both built on Kuwain, our own compact Arabic language model, which is why each is small enough to run in production.

  • Deploy where the constraint says

    Cloud, on-premises or inside your own boundary is a decision taken from your data residency and regulatory position, not from our convenience. The Trust Center sets out what each option involves.

  • Hand over the evaluation with the system

    You keep the test set and the results alongside the system itself. That is what lets your own team re-run the measure after your data changes, and what stops the number ageing quietly.

What we commit to

Stated plainly here, with the detail in the Trust Center.

  • Your data stays where you put it

    On-premises deployment exists so that documents, recordings and records never have to leave the environment they are governed in. Where a project does run in the cloud, data residency is settled and recorded up front rather than discovered later.

  • You own what the engagement produces

    Data ownership is one of the things our architecture is built around, alongside Arabic at the core and fit to the sector. The material you bring stays yours, and so do the evaluation set and the results built from it.

  • Compliance is designed in, not retrofitted

    Sovereign deployment options and regulatory alignment are part of the architecture rather than something added once a security review flags a gap. In government, finance, legal and healthcare, where the constraint is not optional, it is part of the design brief from the first conversation.

  • We publish what we can prove

    10 research papers, 3 benchmarks we built ourselves and open model weights are the reason a claim on this site can be checked. Inside an engagement the same rule holds: a result we cannot measure is not reported as one, and no certification is claimed here that the Trust Center does not list.

Tell us what you are trying to build

The more concretely you can describe the task and the material it runs on, the more useful the first conversation is. We come back with what in the catalogue is closest to it, and what would have to be built.