UAE: AI Learns to Speak the Dialects of the Arab World

From describing customs and traditions to expressing itself in the dialects spoken by local communities, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) is advancing efforts to develop AI systems with a deeper understanding of Arab culture and its linguistic diversity.

In a pioneering research achievement, researchers at the university have developed the first benchmark designed to measure AI models’ ability to understand and engage with Arab culture across 13 local dialects. The study has revealed a notable gap between models’ ability to understand what Arabs say and their ability to speak as people do in everyday life.

For example, when asked to describe a UAE wedding, an AI model can often provide accurate details about customs, traditions and clothing, as well as appropriate expressions used to congratulate the families of the bride and groom. However, when asked to phrase those congratulations in the way people in the UAE would naturally say them in everyday conversation, the model’s fluency, naturalness and cultural awareness can decline significantly.

This discrepancy is at the heart of a new research study by MBZUAI, which found that large language models (LLMs) may possess extensive knowledge of Arab culture but continue to face challenges in expressing that knowledge through the precise local dialects and voices of the communities they represent.

The study’s lead researchers, Professor Fajri Koto, Assistant Professor in the Department of Natural Language Processing, and Mohamed Dehane, a researcher in the same department, said closing this gap would pave the way for the next stage in the development of Arabic-speaking AI, enabling systems not only to understand the language but also to grasp the nuances of how it is used across different Arab societies.

Dehane explained that Arabic connects more than 400 million speakers worldwide, but one of the major challenges in developing Arabic AI is that models are trained almost exclusively on Modern Standard Arabic. As a result, they may appear fluent and capable in conventional tests without having been sufficiently evaluated in natural conversations involving different Arabic dialects and precise local cultural details.

To address this gap, the researchers developed the ArabCulture-Dialogue benchmark, which they described as the first benchmark for testing cultural reasoning in Arabic through multi-turn dialogues in Modern Standard Arabic and 13 local dialects from different Arab countries.

To ensure the benchmark was accurate and representative of real-world language use, the team recruited 26 native Arabic speakers, with two speakers from each of 13 Arab countries. Each participant had spent at least 10 years in their respective country.

Participants drew on situations rooted in their own cultural environments and developed them into short, interactive dialogues, with each dialogue then written in the participant’s local dialect.

The dialogue dataset covered 12 everyday topics, ranging from weddings and food to childcare, agriculture, arts and games.

Professor Koto said the researchers evaluated the models through three main tasks: selecting the most culturally appropriate response from several options; translating between Modern Standard Arabic and a specific dialect in both directions; and continuing a dialogue in the required dialect.

The results showed that the strongest models were highly capable of identifying culturally appropriate responses, achieving scores of close to 95 per cent even when conversations shifted from Modern Standard Arabic to a specific dialect. However, their performance declined sharply when the task shifted from understanding a dialect to producing it, whether translating a sentence into Emirati Arabic or continuing a dialogue in the dialect.

The findings highlight a subtle but important gap in AI development: the challenge is not always understanding meaning, but preserving the identity and cultural specificity of a particular dialect.

The study found that models performed better when dealing with customs shared across the Arab world, while they faced greater difficulty with practices specific to individual countries. North African dialect dialogues were among the most challenging overall, while Emirati Arabic also ranked among the dialects that posed the greatest difficulties for the models.

The models succeeded in producing the dialect corresponding to the target country in only around half of the cases. Some specialised Arabic models and open-weight models achieved the desired result only in a limited number of instances.

By contrast, the models performed considerably better when translating from dialects into Modern Standard Arabic than in the opposite direction. Returning text to a standardised form is easier, while giving it an authentic local voice remains the greater challenge.

Dehane explained that the apparent paradox lies in the fact that cultural knowledge is already present within the models but can sometimes require only minimal guidance to be accessed effectively. Specifying the country and region associated with a dialogue improved the accuracy of the results.

This finding suggests that advancing Arabic AI requires not only larger quantities of data, but also more diverse and precise datasets that reflect the social and cultural contexts in which language is actually used.

The findings hold particular strategic significance for the UAE, which has made artificial intelligence a national priority, from the UAE National Strategy for Artificial Intelligence 2031 to the development of local AI models, including Jais, a bilingual large language model supporting Arabic and English, which MBZUAI contributed to developing.

The study highlights the university’s leading role in advancing Arabic AI research from general language support to a more sophisticated understanding of the linguistic and cultural diversity of the Arab world.

The next stage is therefore not simply for AI to be able to speak Arabic, but to know which Arabic it is speaking, in what context and with whom, while expressing local culture without losing its distinctive character.

Professor Koto said the study’s broader conclusion serves as a warning to those developing Arabic-language AI models: a system that supports Arabic is not necessarily a system that understands Arabic in all its dialects, along with the cultures and identities embedded within them.

The findings open the door to a new phase of research, including expanding the range of dialects covered by the benchmark and developing more precise methods for evaluating models’ ability to understand local cultural contexts and produce language authentically.

The research team published its paper, titled “Cultural Benchmarking of Large Language Models in Modern Standard Arabic and Arabic Dialect Dialogues”, which was presented at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2026). The study provides a foundation for future research aimed at developing AI models that are better equipped to understand Arab linguistic and cultural diversity while supporting efforts to preserve local cultures and identities.

Facebook
WhatsApp
Al Jundi

Please use portrait mode to get the best view.