Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Datasets

Sadeed Tashkeela

A large, high-quality Arabic diacritized corpus for training and evaluating diacritization models.

1,045,183 rowsPublic listing, access on requestArabic foundation models & open AI

What is in it

1,042,698 training rows and 2,485 test rows of diacritized Arabic, 1.43 GB in Parquet. Each row carries three fields: the source filename, the undiacritized input, and the fully diacritized output. The test split is held separate in the published files, so a fine-tune and its own held-out evaluation come out of the same download.

How it was built

The corpus is derived from Tashkeela, the standard large Arabic diacritized collection, put through the cleaning and normalisation pipeline built for the Sadeed paper. That work exists because diacritization data in the wild is inconsistent: the same word appears with partial, full and absent diacritics across sources, and a model fine-tuned on the mixture learns the inconsistency. The pipeline normalises the diacritic forms, removes the entries that cannot be reconciled, and keeps the input and output strictly aligned so the task stays a character-faithful restoration rather than a rewrite.

Intended uses

Fine-tuning and evaluating Arabic diacritization models: the use it was built for. Sadeed itself is a fine-tune of Kuwain 1.5B on exactly this kind of curated diacritized data. Downstream, accurate diacritization matters most to Arabic text-to-speech, to machine translation where an unvowelled word is ambiguous, and to language-learning tools. For publishable evaluation numbers the SadeedDiac-25 benchmark is the fairer instrument: it was built precisely because existing diacritization benchmarks were too narrow in genre and complexity.

Licence and access

Listed publicly on Hugging Face with access granted on request. The card declares no SPDX licence of its own, so use is governed by the terms of the underlying Tashkeela material and of the access request; the licence field on this page is left empty rather than guessed.

Related work

The Sadeed paper documents both the cleaning pipeline behind this corpus and the model trained on it, and introduces SadeedDiac-25, the benchmark released to evaluate diacritization across genres and complexity levels. Kuwain 1.5B is the base model Sadeed was adapted from.

Cite this dataset

@misc{aldallal2025sadeed,
  title         = {Sadeed: Advancing Arabic Diacritization Through Small Language Model},
  author        = {Zeina Aldallal and Sara Chrouf and Khalil Hennara and Mohamed Motaism Hamed and Muhammad Hreden and Safwan AlModhayan},
  year          = {2025},
  eprint        = {2504.21635},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2504.21635}
}

Dataset facts

Records
1,045,183
Format
Parquet
Access
Public listing, access on request
Released
Built from
Tashkeela
Research track
Arabic foundation models & open AI
Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.