#SaudiArabia #Arabic - Arabic AI Atlas, an open-source map and repository of Arabic AI models, datasets and tools has been launched as a free-to-use resource for AI developers. Created by Riyadh-based AI professional Hesham Haroon, it lists more than 1,000 models and tools and more than 1,000 datasets, covering LLMs, speech recognition, text-to-speech, OCR and embeddings. The repository is also a Claude Code plugin, with an offline server that lets AI agents recommend models. The Atlas aims to provide developers with one reliable place to look for Arabic-capable open-source code and data sets.
SO WHAT? - Open-source Arabic AI code and data is scattered across Hugging Face, GitHub, other community sites and research archives, presenting a challenge for anyone who needs a clear picture of what exists. The Atlas aims to fix that, and aggregates data for open-source Arabic models, tools and data sets. This is an important new resource, because much of the work on regional dialects comes from individuals and communities with low visibility (compared to research released with corporate or government backing).
KEY POINTS:
Riyadh-based Hesham Haroon has released the Arabic AI Atlas of the open-source Arabic AI ecosystem, consisting of an interactive map and repository. The hub helps answer questions developers now chase across Google, Hugging Face, GitHub and research papers, such as “which open Arabic text-to-speech model runs on a mobile?”
The new Atlas lists more than 1,000 models and tools and more than 1,000 datasets from developers all over the world. Categories cover LLMs, speech recognition, text-to-speech, OCR, embeddings, tools, benchmarks and organisations. Each entry records country, organisation, licence, size, Hugging Face downloads and the project’s last update.
An interactive map places entries by country, allowing users to filter by dialect or licence and share a link to a filtered view. A grid view of the data shows which countries work in each category and the most downloaded entries in each field.
The repository doubles as a Claude Code plugin, with five skills and an MCP (Model Context Protocol) server that reads the Atlas offline. Its recommend tool answers queries from the data, and any MCP client can connect, not just Claude.
YAML files (at Github file used to configure automated workflows, repository settings, and issue templates) in the data folder are the only hand-edited source. Continuous integration checks every pull request against a schema. Download counts are pulled down from Hugging Face daily to update and regenerate the map, README, atlas.json and llms.txt for LLMs.
Saudi Arabia leads the Atlas country listings with 355 entries, ahead of the UAE on 184, Egypt on 175, Qatar on 128 and Morocco on 71. By dialect, Modern Standard Arabic has 441 entries, mixed 310, Maghrebi 307 and Gulf 170.
The new Atlas already highlights certain patterns and trends. For example, most models for Egyptian, Moroccan Darija, Tunisian, Algerian, Saudi, Jordanian, Kuwaiti, Bahraini, Yemeni and Libyan dialects have been developed from individuals and communities. Industry and national AI programmes often overlook them, and previously growth in this work goes largely untracked.
The Atlas also lists 15 gaps in Arabic AI development. These include open Kuwaiti or Qatari text-to-speech, a Libyan Arabic corpus, Palestinian speech recognition and an open Arabic handwriting benchmark. Mauritania has no listed entry of any kind.
Code for the Arabic AI Atlas is MIT licensed, data is CC BY 4.0, and contributors can add entries by pull request.
ZOOM OUT - Arabic AI has lagged behind English, and data is the main reason. LLMs typically need datasets of tens of billions of tokens, and quality Arabic text is scarce, especially for colloquial speech. Training data for Arabic dialects is much less available than data for standard Arabic (fusha). Egypt, Qatar, Saudi Arabia and the UAE have invested heavily in acquiring and digitising Arabic content, but most of these projects remain work in progress. Meanwhile, Arabic model developers are also spread thinly across the region, with few shared community resources.. So, tools like the Arabic AI Atlas help by bringing models, tools and datasets together, making it easier for the development community to identify, evaluate and use them.
[Written and edited with the assistance of AI]
Source: Hesham Haroon, MEAIN
LINKS
Arabic AI Atlas (Github)
Arabic AI Atlas repo (Github)
Read more recent news about Arabic language AI:
NVIDIA halves Saudi dialect errors with SDAIA dataset (Middle East AI News)
Morocco, Mistral partnership releases open-source Darija AI (Middle East AI News)
TokenAI open-sources Horus Taleeq Arabic language model (Middle East AI News)
HUMAIN unveils 428B parameter Arabic MiniMax model (Middle East AI News)
HUMAIN invests in Arabic-first technology firm Arabic AI (Middle East AI News)





