الإمارات.. ذكاء اصطناعي يتعلم التحدث بلهجات العرب - وكالة وام

15 أغسطس 2026·وكالة أنباء الإمارات
الإمارات.. ذكاء اصطناعي يتعلم التحدث بلهجات العرب - وكالة وام

أبوظبي في 16 أغسطس/ وام/ من وصف العادات والتقاليد إلى القدرة على التعبير بلهجة أهلها، تضطلع جامعة محمد بن زايد للذكاء الاصطناعي بتطوير ذكاء اصطناعي أكثر فهماً للثقافة العربية وتنوعها اللغوي.

وفي إنجاز بحثي رائد طور باحثون في الجامعة أول معيار مرجعي لقياس قدرة نماذج الذكاء الاصطناعي على فهم الثقافة العربية والتفاعل معها عبر 13 لهجة محلية، كاشفين عن فجوة لافتة بين قدرة النماذج على فهم ما يقوله العرب، وقدرتها على التحدث كما يتحدثون في حياتهم اليومية.

وعندما يطلب من أحد نماذج الذكاء الاصطناعي وصف حفل زفاف إماراتي، يستطيع غالباً تقديم تفاصيل دقيقة عن العادات والتقاليد والملابس، وصولاً إلى عبارات التهنئة المناسبة لمباركة أهل العروسين، لكن ما إن يُطلب منه صياغة العبارة بالطريقة التي يقولها أهل الإمارات في حياتهم اليومية، حتى تتراجع تلك الطلاقة والانسيابية والإلمام الثقافي الذي أظهره النموذج في البداية.

وهذه المفارقة كانت محور دراسة بحثية جديدة من جامعة محمد بن زايد للذكاء الاصطناعي، كشفت أن النماذج اللغوية الكبيرة قد تمتلك معرفة واسعة بالثقافة العربية، لكنها لا تزال تواجه تحدياً في التعبير عنها بصوت أهلها ولهجاتهم المحلية الدقيقة.

ويرى الباحثان الرئيسيان في الدراسة، البروفيسور فجري كوتو، الأستاذ المساعد في قسم معالجة اللغة الطبيعية، ومحمد ديهان، الباحث في القسم ذاته، في تصريحات لوكالة أنباء الإمارات "وام" أن سد هذه الفجوة يمهد للمرحلة المقبلة من تطور الذكاء الاصطناعي الناطق بالعربية، بحيث لا يكتفي بفهم اللغة، وإنما يستوعب أيضاً خصوصيات استخدامها في المجتمعات العربية المختلفة.

وقال محمد ديهان إن اللغة العربية توحد أكثر من 400 مليون متحدث حول العالم، إلا أن أحد أبرز التحديات في تطوير الذكاء الاصطناعي يتمثل في أن النماذج تدرب بصورة شبه حصرية على العربية الفصحى الحديثة، ولذلك قد تبدو طليقة وقادرة في الاختبارات، من دون أن تكون قد اختبرت بصورة كافية في محادثات طبيعية باللهجات المختلفة في الدول العربية، مع الالتزام بدقة تفاصيل الثقافية المحلية.

ولمعالجة هذه الثغرة، طور الباحثون معيار ArabCulture-Dialogue، الذي وصفوه بأنه أول معيار مرجعي لاختبار الاستدلال الثقافي باللغة العربية ضمن حوارات متعددة الجولات، بالعربية الفصحى الحديثة و13 لهجة محلية من دول عربية مختلفة.

ولضمان دقة المعيار وارتباطه بالاستخدام الحقيقي للغة، استعان الفريق بـ26 ناطقاً أصلياً بالعربية، بواقع متحدثين اثنين من كل دولة من 13 دولة عربية، ممن أمضى كل واحد منهم عشرة أعوام على الأقل في بلده.

وانطلق المشاركون من مواقف مستمدة من بيئاتهم الثقافية، ووسعوها إلى حوارات قصيرة متبادلة، ثم صاغ كل منهم الحوار بلهجته المحلية.

وغطت مجموعة الحوارات 12 موضوعاً من الحياة اليومية، بدءاً من حفلات الزفاف والطعام، وصولاً إلى تربية الأطفال والزراعة والفنون والألعاب.

وقال البروفيسور فجري كوتو إن الباحثين اختبروا النماذج من خلال ثلاث مهام رئيسية، شملت اختيار الرد الأنسب ثقافياً من بين عدة خيارات، والترجمة بين العربية الفصحى الحديثة ولهجة محددة في الاتجاهين، إلى جانب مواصلة الحوار باللهجة المطلوبة.

وأظهرت النتائج أن النماذج الأقوى تمتلك قدرة كبيرة على تمييز الرد المناسب ثقافياً، إذ سجلت نتائج قاربت 95 % حتى عند انتقال الحوار من العربية الفصحى إلى اللهجة المحددة، غير أن أداءها انخفض بشكل حاد عندما انتقلت المهمة من فهم اللهجة إلى إنتاجها، سواء عند ترجمة جملة إلى اللهجة الإماراتية أو عند مواصلة حوار بها.

وتكشف هذه النتائج عن فجوة دقيقة في تطوير الذكاء الاصطناعي؛ فالمشكلة لا تكمن دائماً في فهم المعنى، وإنما في القدرة على الحفاظ على هوية اللهجة وخصوصيتها الثقافية.

وتشير الدراسة إلى أن النماذج كانت أفضل في التعامل مع العادات المشتركة على امتداد العالم العربي، بينما واجهت صعوبة أكبر في فهم العادات الخاصة بكل دولة، وكانت حوارات لهجات شمال أفريقيا من بين الأصعب إجمالاً، كما جاءت الحوارات الإماراتية بين أكثر اللهجات تحدياً للنماذج.

ولم تنجح النماذج في إنتاج اللهجة المطابقة للدولة المستهدفة إلا في نحو نصف الحالات، فيما لم تحقق بعض النماذج العربية المتخصصة ومفتوحة الأوزان هذه النتيجة إلا في حالات معدودة.

وفي المقابل، كانت النماذج أفضل بكثير في الترجمة من اللهجات إلى العربية الفصحى الحديثة مقارنة بالاتجاه المعاكس؛ فالعودة بالنص إلى اللغة المعيارية أسهل، أما منحه نبرة محلية أصيلة، فهو التحدي الأكبر.

وأوضح محمد ديهان أن المفارقة تتمثل في أن المعرفة الثقافية موجودة بالفعل داخل النماذج، لكنها تحتاج أحياناً إلى قدر بسيط من التوجيه، موضحاً أن تحديد الدولة والمنطقة اللتين ينتمي إليهما الحوار أدى إلى تحسن دقة النتائج.

وتشير هذه النتيجة إلى أن تطوير الذكاء الاصطناعي العربي لا يتطلب فقط كميات أكبر من البيانات، وإنما يحتاج أيضاً إلى بيانات أكثر تنوعاً ودقة، تعكس السياقات الاجتماعية والثقافية التي تستخدم فيها اللغة فعلياً.

وتكتسب هذه النتائج أهمية إستراتيجية خاصة في دولة الإمارات، التي جعلت الذكاء الاصطناعي أولوية وطنية، بدءاً من إستراتيجية الإمارات الوطنية للذكاء الاصطناعي 2031، وصولاً إلى تطوير نماذج ذكاء اصطناعي محلية، من بينها "جيس"، النموذج اللغوي الكبير ثنائي اللغة الذي يدعم العربية والإنجليزية، وأسهمت جامعة محمد بن زايد للذكاء الاصطناعي في تطويره.

ويبرز من خلال هذه الدراسة الدور الريادي للجامعة في الانتقال بأبحاث الذكاء الاصطناعي العربي من مرحلة دعم اللغة العربية بصورة عامة إلى مرحلة أكثر تقدماً، تتمثل في فهم التنوع اللغوي والثقافي داخل العالم العربي.

فالمرحلة المقبلة لا تتمثل في أن يكون الذكاء الاصطناعي قادراً على التحدث بالعربية فحسب، وإنما أن يعرف أي عربية يتحدث، وفي أي سياق، ومع من، وأن يتمكن من التعبير عن الثقافة المحلية دون أن يفقد خصوصيتها.

وقال البروفيسور فجري كوتو إن الخلاصة الأوسع للدراسة تحمل رسالة تنبيه إلى كل من يعمل على تطوير نماذج ذكاء اصطناعي باللغة العربية، مفادها أن النظام الذي يدعم اللغة العربية ليس بالضرورة نظاماً يفهم العربية بمختلف لهجاتها، وما تحمله هذه اللهجات في طياتها من ثقافات وهويات.

وتفتح نتائج الدراسة الباب أمام مرحلة جديدة من البحث العلمي، يمكن خلالها توسيع نطاق اللهجات المشمولة بالمعيار، وتطوير أساليب أكثر دقة لتقييم قدرة النماذج على فهم السياقات الثقافية المحلية وإنتاج اللغة بصورة أصيلة.

ونشر الفريق ورقته البحثية المعنونة "التقييم المعياري الثقافي للنماذج اللغوية الكبيرة في حوارات العربية الفصحى واللهجات العربية"، والتي عُرضت في الاجتماع السنوي الثالث والستين لجمعية اللغويات الحاسوبية "ACL 2026"، لتشكل أساساً لأبحاث لاحقة تسهم في تطوير نماذج ذكاء اصطناعي أكثر قدرة على فهم التنوع اللغوي والثقافي العربي، وتعزز في الوقت ذاته جهود صون الثقافات والهويات المحلية.

AIArabicLevantine
الخبر متوفّر باللغات التالية:English
المصدر الأصلي للخبر
وكالة أنباء الإمارات

UAE: Artificial Intelligence Learns to Speak Arabic Dialects - WAM

August 15, 2026·WAM
UAE: Artificial Intelligence Learns to Speak Arabic Dialects - WAM

Abu Dhabi, August 16 (WAM) - From describing customs and traditions to being able to express themselves in the local dialect, the Mohamed Bin Zayed University of Artificial Intelligence is developing an artificial intelligence that better understands Arabic culture and its linguistic diversity.

In a pioneering research achievement, researchers at the university have developed the first reference standard for measuring AI models' ability to understand and interact with Arab culture through 13 local dialects, revealing a striking gap between the models' ability to comprehend what Arabs say and their ability to speak as they do in daily life.

When asked to describe an Emirati wedding, one of the AI models can often provide detailed information about customs, traditions, clothing, and even appropriate phrases for blessing the bride and groom. However, when asked to phrase the sentence in the way that Emiratis would say it in their daily lives, the fluency, ease, and cultural knowledge that the model initially demonstrated diminishes.

This paradox is the focus of a new study from the Mohamed Bin Zayed University of Artificial Intelligence, which reveals that large language models may have extensive knowledge of Arab culture but still face challenges in expressing themselves in the precise local dialects of their people.

The two main researchers of the study, Professor Fجري كوتو, an assistant professor in the Department of Natural Language Processing, and Mohamed Dehan, a researcher in the same department, stated to the Emirates News Agency (WAM) that bridging this gap paves the way for the next stage of development of Arabic-speaking artificial intelligence. This will not only enable it to understand the language but also absorb its specific use within different Arab societies.

Mohamed Dehan said that Arabic unifies more than 400 million speakers worldwide, but one of the main challenges in developing AI is that models are almost exclusively trained on modern Standard Arabic. Therefore, they may appear fluent and capable in tests without being sufficiently tested in natural conversations using different dialects in Arab countries while adhering to the details of local culture.

To address this gap, the researchers developed the ArabCulture-Dialogue standard, which they described as the first reference standard for testing cultural reasoning in Arabic within multi-round dialogues in modern Standard Arabic and 13 local dialects from different Arab countries.

To ensure the accuracy of the standard and its relevance to real language use, the team relied on 26 native Arabic speakers, with two speakers from each of 13 Arab countries who had spent at least ten years in their country.

The participants started from situations derived from their cultural environments and expanded them into short reciprocal dialogues. Then each one formulated the dialogue in his/her local dialect.

The dialogue set covered 12 topics from daily life, ranging from weddings and food to child-rearing, agriculture, arts, and games.

Professor Fجري كوتو said that researchers tested the models through three main tasks, including selecting the most culturally appropriate response from several options, translating between modern Standard Arabic and a specific dialect in both directions, and continuing the dialogue in the required dialect.

The results showed that stronger models have great ability to distinguish the culturally appropriate response, recording scores of up to 95% even when the dialogue shifted from Standard Arabic to a specific dialect. However, their performance dropped sharply when the task changed from understanding the dialect to producing it, whether translating a sentence into the Emirati dialect or continuing a conversation in it.

These results reveal a subtle gap in artificial intelligence development; the problem is not always about understanding meaning, but also about maintaining the identity and cultural specificity of the dialect.

The study indicates that models were better at dealing with common customs across the Arab world, while they faced greater difficulty in understanding country-specific customs. North African dialect dialogues were among the most challenging overall, and Emirati dialogues were among the most challenging for the models.

Models only succeeded in producing the target state's dialect in about half of cases, while some specialized Arabic open-weight models did not achieve this result even in a few cases.

On the other hand, models were much better at translating from dialects to modern standard Arabic than vice versa; it is easier to return text to the standard language, but giving it an authentic local tone is a greater challenge.

Mohammed Dehan explained that the paradox is that cultural knowledge already exists within the models, but sometimes needs a little guidance. He said that specifying the country and region from which the dialogue originates has improved the accuracy of results.

This result indicates that developing Arabic artificial intelligence does not just require larger amounts of data, but also more diverse and accurate data that reflects the social and cultural contexts in which language is actually used.

These findings are particularly significant for the United Arab Emirates, which has made artificial intelligence a national priority. This includes the UAE National Artificial Intelligence Strategy 2031 and the development of local AI models such as "Jais", a bilingual large language model that supports Arabic and English. The Mohammed Bin Zayed University for Artificial Intelligence contributed to its development.

This study highlights the university's pioneering role in advancing Arabic artificial intelligence research from a general support level for the Arabic language to a more advanced stage, which is understanding linguistic and cultural diversity within the Arab world.

The next phase is not just about AI being able to speak Arabic, but also knowing what kind of Arabic it is speaking, in what context, with whom, and being able to express local culture without losing its specificity.

Professor Fagrī Kūṭō said that the broader conclusion of the study carries a message of caution for those working on developing Arabic language AI models. The message is that a system supporting Arabic does not necessarily mean it understands Arabic in all its dialects and the cultural and identity dimensions they carry.

The results of the study open the door to a new phase of scientific research, during which the scope of the included dialects can be expanded, and more accurate methods developed for evaluating models' ability to understand local cultural contexts and produce language authentically.

The team published its research paper entitled "Cultural Standard Evaluation of Large Language Models in Arabic Literary and Dialectal Conversations," presented at the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2026). This forms a foundation for future research aimed at developing AI models that are more capable of understanding Arab linguistic and cultural diversity, while also enhancing efforts to preserve local cultures and identities.

AIArabicLevantine
This article is also available in:العربية
Original source article
WAM