Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Specialized NLP

Sadeed

Arabic diacritization built on Kuwain, restoring full Tashkeel for MSA and classical Arabic.

1.5BOpen weights

About Sadeed

Sadeed is a 1.5-billion-parameter model that restores diacritics, or tashkeel, to Arabic text. It is fine-tuned from Kuwain-1.5B, Misraj's Arabic small language model, and on the Fadel benchmark for Classical Arabic its paper reports a state-of-the-art word error rate.

Most modern Arabic is written without diacritics, so the same letters can spell different words: qalb (heart), qalaba (to turn) and qullab (volatile) look identical on the page. Speech synthesis, machine translation, morphological analysis and part-of-speech tagging all depend on resolving that ambiguity, and the right vowel often depends on the rest of the sentence, even on its punctuation. Labeled data is scarce and mostly Classical Arabic, and some widely used test sets overlap with common training sets, which inflates reported results.

How it works

  • Clean training data. Built from the Tashkeela corpus and the Arabic Treebank (ATB-3), cleaned with Kuwain's own pipeline, normalized for consistent diacritization, split into chunks of 50 to 60 words that keep sentence context, and filtered for fully diacritized text. The result is 1,042,698 examples, about 53 million words, with overlap against the Fadel test set cut to 0.4% of examples.
  • Diacritization as question answering. Each example is wrapped in a fixed prompt template, and the model generates the diacritized version of its input.
  • Hallucination repair. The output is aligned back to the input with the Needleman-Wunsch algorithm. Added words are removed, missing ones are restored, and altered words are returned without diacritics.
  • Modest compute. Three epochs on 8 A100 GPUs.

Capabilities

Benchmark results

All figures are from the Sadeed paper (arXiv:2504.21635). Lower is better.

  • Fadel test set (Classical Arabic): a word error rate of 1.83 without case endings, against 1.96 for SUKOUN, the strongest earlier system in the comparison. With case endings Sadeed scores 4.49, behind SUKOUN's 3.34; on the phonologically corrected Fadel set the team released, it scores 2.94.
  • WikiNews (Modern Standard Arabic): competitive, not first. Sadeed's word error rate without case endings is 8.44, against 2.9 for the feature-rich BiLSTM of Darwish et al., which was trained on data from the same distribution.
  • SadeedDiac-25: the best of the open models tested, with a word error rate of 9.92 without case endings against 40.25 for Aya-23-8B, the next open model. Claude 3.7 Sonnet leads on every metric (2.31). With case endings GPT-4 and Gemini Flash 2.0 are also ahead, at 5.27 and 7.99 against Sadeed's 13.74; without them, Sadeed's 9.92 beats GPT-4's 10.93. About 7.19 points of Sadeed's 9.92 come from hallucinated words.

SadeedDiac-25 is Misraj's public benchmark of 1,200 paragraphs. Half is Modern Standard Arabic (454 paragraphs curated from web articles plus 146 from WikiNews, 40 to 50 words each) and half is Classical Arabic (600 paragraphs from the Fadel test set). Two experts reviewed the text, then checked each other's corrections.

Deployment

Sadeed is available through the Misraj cloud API, including Kawn Console, or on-premises. The evaluation code is on GitHub at github.com/misraj-ai/Sadeed, and the cleaned training corpus, Sadeed Tashkeela, is listed on Hugging Face with access on request.

Who it is for

  • Teams building Arabic text-to-speech, where an unvowelled word is ambiguous to pronounce.
  • Machine translation and search pipelines that need to tell apart words spelled the same way.
  • Publishers and language-learning tools that present fully vowelled text, especially Classical Arabic.

FAQ

Is Sadeed stronger on Classical or Modern Standard Arabic? Classical. Its training data is mostly Classical Arabic, and the paper reports weaker results on text that is mostly MSA. The paper says the team is expanding its MSA training data.

Why does it sometimes leave a word without diacritics? When the model changes a word, the alignment step puts the original word back without diacritics rather than keep a wrong one.

What are its known limits? Hallucination, especially on non-Arabic words, and the small amount of openly available diacritized MSA data.

At a glance

Type
Specialized NLP
Parameters
1.5B
License
Open weights
Deployment
Cloud APIOn-premises
Research track
Arabic foundation models & open AI

Benchmarks

DatasetMetricResult
SadeedDiac-25Diacritization results were being reported on narrow, inconsistent test sets, which made models hard to compare honestly. A single benchmark covering both registers and several genres is what lets one diacritization result be read against another.
Coverage
MSA + Classical Arabic
Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.