Skip to content
MISRAJAI
What we believe

We understand deeply. We build from the source. And we run with confidence.

Datasets

Arabic Image Captioning (100M)

A large-scale dataset of 100 million Arabic image captions, generated with the Mutarjim translation model.

100M captionsPublic listing, access on requestArabic foundation models & open AI

What is in it

93,613,962 rows in the published train split, 73.7 GB across 148 Parquet shards. Each row holds three things: the image URL, the original caption as it was written, and its Arabic translation. Keeping the source caption beside the Arabic is deliberate: it makes every row auditable, and it lets the set be used as a parallel caption corpus as well as an Arabic one. The images are not redistributed; the rows carry their links.

How it was built

The Arabic side was generated with Mutarjim, Misraj's 1.5-billion-parameter Arabic-English translation model. Mutarjim was chosen for this because its fine-tuning data weighted Arabic-origin samples two to one over English-origin ones, which is the direction that matters here: the whole job is generating Arabic, the harder direction, and on the Tarjama-25 benchmark Mutarjim leads in exactly that direction with a COMET of 83.41 against GPT-4o mini's 83.36, chrF++ of 68.67 against 66.36, and BLEU of 43.71 against 38.52.

Intended uses

Pre-training and fine-tuning Arabic vision-language models: captioning, image-text retrieval, and multimodal alignment at a scale Arabic has not previously had. Because the captions are machine-translated, the set is training material rather than an evaluation benchmark: a caption model's score should be reported on a human-written Arabic set, and this corpus labelled as translated in any data statement.

Licence and access

Listed publicly on Hugging Face with access granted on request. The card declares no SPDX licence of its own, so the terms of the underlying image and caption sources and of the access request govern use, and the licence field here is deliberately empty. Since only image URLs are carried, anyone training on it fetches the pictures themselves under the originating sites' terms.

Related work

The Mutarjim paper documents the translation model that produced the Arabic side, its two-phase training, and Tarjama-25, the 5,000-pair expert-reviewed benchmark released with it. Kuwain 1.5B is the base model Mutarjim was built on.

Cite this dataset

@misc{hennara2025mutarjim,
  title         = {Mutarjim: Advancing Bidirectional Arabic-English Translation with a Small Language Model},
  author        = {Khalil Hennara and Muhammad Hreden and Mohamed Motaism Hamed and Zeina Aldallal and Sara Chrouf and Safwan AlModhayan},
  year          = {2025},
  eprint        = {2505.17894},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2505.17894}
}

Dataset facts

Size
100M captions
Records
93,613,962
Format
Parquet
Access
Public listing, access on request
Released
Research track
Arabic foundation models & open AI
Get started

Let's talk about what you're trying to build.

Every good solution starts with a clear question. Ask us, and our team is with you from question to solution.