MBZUAI builds first cultural benchmark for Arabic AI
AI understands dialects but struggles to speak them, Emirati Arabic among hardest
#UAE #LLMS - Abu Dhabi-based Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) has developed ArabCulture-Dialogue, the first benchmark testing AI models’ cultural reasoning across Modern Standard Arabic and 13 national dialects. Benchmark results revealed that leading global models correctly identify culturally appropriate responses in the mid-90 percent range. However, performance drops sharply when AI models are asked to generate dialect themselves, succeeding only about half the time. The new benchmark research was presented at the 63rd Annual Meeting of the Association for Computational Linguistics.
SO WHAT? - Arabic has more than 400 million speakers, but most AI models are trained almost exclusively on Modern Standard Arabic and much less on regional Arabic language dialects. The MBZUAI study shows models can recognise the right cultural answer in colloquial Arabic, but are largely unable to produce it in the dialect a real person would use. The gap matters directly for global and Arab AI developers building models that are intended to be capable in Arabic understanding, reasoning and writing.
KEY POINTS:
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) researchers have built ArabCulture-Dialogue, the first benchmark for testing Arabic cultural reasoning in multi-turn conversations. The benchmark works across Modern Standard Arabic (MSA) and 13 national dialects, built with 26 native Arabic speakers across 13 countries.
The dataset covers 12 everyday topics, from weddings and food to parenting, agriculture, arts and games, spanning 54 fine-grained subtopics and more than 340,000 words.
Models were tested on three tasks: selecting the culturally appropriate reply from a set of options, translating between MSA and a specific dialect, and continuing a conversation in a named dialect.
The strongest models scored in the mid-90s at picking correct culturally appropriate answers, even when conversations shifted from MSA into dialect, but performance dropped sharply on dialect production tasks.
Models produced the correct dialect for the target country in only about half of cases, with North African and Emirati dialogues among the most difficult overall.
The cultural knowledge already exists within the models but often needs only a small nudge, noting that specifying the country and region associated with a conversation improved accuracy.
The researchers evaluated a range of Arabic-centric models including Jais, ALLaM and SILMA, multilingual models such as Gemma and Qwen3, and proprietary models including GPT-5 and Gemini 2.5 Pro, finding smaller open-weight Arabic-centric models struggled most, in some cases nearing random guessing on dialect tasks.
The findings may carry a cautionary message for anyone developing Arabic-language AI. Professor Fajri Koto, Assistant Professor in the Department of Natural Language Processing, notes that “a system that supports Arabic isn’t necessarily one that understands Arabic across its dialects and the cultures embedded within them.”
The research team includes MBZUAI - Muhammad Dehan Al Kautsar∗1 Saeed Almheiri∗1 Momina Ahsan, Bilal Elbouardi,, Sarfraz Ahmad, Amr Keleg, Omar El Herraoui1 Kareem Elzeky, Abed Alhakim Freihat, Mohamed Anwar, Zhuohan Xie1 Junhong Liang, Preslav Nakov, Fajri Koto; IBM Research - Younes Samih; and American University in the Emirates - Mohammad Rustom Al Nasar.
[Written and edited with the assistance of AI]
Source: MBZUAI
LINKS
ArabCulture-Dialogue data (Hugging Face)
Read more about MBZUAI AI model research:
MBZUAI officially launches K2 Think V2 with mobile apps (Middle East AI News)
MBZUAI, Inception launch enhanced Nanda Hindi LLM (Middle East AI News)
MBZUAI releases K2 V2 open-source reasoning model (Middle East AI News)
New K2 Think model debuts on OnDemand platform (Middle East AI News)
K2 Think 32B rivals reasoning models 20 times its size (Middle East AI News)


