#SaudiArabia #AImodels - NVIDIA has used the Saudi Audio Dataset for Arabic (SADA) to fine-tune its Nemotron 3.5 ASR speech recognition model for Najdi and Hijazi Saudi Arabian dialects. According to NVIDIA’s developer blog, training on SADA audio data cut the word error rate on Saudi Arabic dialects from 55 percent to 30 percent, roughly halving transcription mistakes, while the error rate across SADA’s full multi-dialect set fell from 59 percent to 36 percent. The dataset was built by Saudi Data and Artificial Intelligence Authority (SDAIA) in collaboration with the Saudi Broadcasting Authority,
SO WHAT? - Most Arabic-capable models are trained principally using Modern Standard Arabic (MSA), the formal written form of Arabic. However, most spoken Arabic is dialectal, not formal Arabic and training data for Arabic dialects is scarce. SDAIA built SADA specifically to close the gap in Arabic training data, and NVIDIA’s results show the approach works. According to the AI chipmaker, dialect accuracy roughly doubled without degrading performance on English or Modern Standard Arabic (which actually improved slightly alongside it).
KEY POINTS:
NVIDIA has fine-tuned its open weight Nemotron 3.5 ASR model, which supports around 40 languages and dialects, using the Saudi Audio Dataset for Arabic (SADA) dataset to specifically target Najdi and Hijazi Saudi Arabic dialects and Khaleeji (Gulf) Arabic.
Developed by Saudi Data and Artificial Intelligence Authority (SDAIA), SADA comprises around 667 hours of transcribed audio across more than 10 Saudi dialects. The dataset includes over 600 hours from 57 Saudi Broadcasting Authority television programmes, and more than 125,000 categorised audio clips, released publicly on Kaggle.
The initial pretrained NVIDIA model initially misidentified about 55 of every 100 words in Saudi dialects. After fine-tuning, the word error rate dropped to roughly 30 words, a reduction of 25 percentage points.
Across the full SADA test set covering all dialects, word error rate fell from 58.84% to 35.61%, while character-level error rate dropped from 35.40% to 15.97%.
Fine-tuning for the audio model was completed in around 4.5 hours over 12,000 training steps using two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs.
The approach combined three techniques: training narrowly on the two target dialects rather than all eleven in the dataset. The approach also mixed in a small “replay” stream of previously learned English and Arabic data to prevent the model forgetting those skills, and grouping similar-length audio clips into batches for more efficient training.
Performance on English and Modern Standard Arabic (MSA) improved alongside the dialect gains rather than being traded off. FLEURS English word error rate fell from 11.04% to 10.42% and FLEURS Arabic from 12.67% to 11.41%.
Updating all 24 layers of the model’s encoder outperformed partial updates, reducing word error rate to 29.96%, compared with 32.32% when only the top eight layers were updated.
The enhanced model supports real-time applications with latency starting at 80 milliseconds, suited to voice assistants, conversational agents, live subtitling and transcription of call centres, media archives and meetings.
NVIDIA has published the full fine-tuning workflow and tools, allowing developers to replicate the same dialect-adaptation methodology for other languages and dialects.
ZOOM OUT - SADA exists because there are few resources for Arabic speech data, and even fewer for dialectal Arabic. Before its release, available Arabic dialect datasets ran to 20 hours of Egyptian, 60 of Levantine and just 15 of Gulf Arabic. Meanwhile, most published Arabic corpora focused on Modern Standard Arabic rather than the dialects that people speak day-to-day. SDAIA's National Center for Artificial Intelligence built SADA in collaboration with the Saudi Broadcasting Authority to close the gap. The dataset draws on roughly 667 hours of audio from more than 57 TV shows and covers Saudi Arabia's three major dialect families, Najdi, Hijazi and Khaleeji, plus several sub-dialects. The dataset is split into an 418-hour training set and 10-hour validation and test sets. Each set is transcribed and annotated by speaker age, gender, dialect and recording environment, and released under a Creative Commons licence (CC BY-NC-SA) for public use.
[Written and edited with the assistance of AI]
Source: SPA, SDAIA, NVIDIA
LINKS
NVIDIA Technical Blog (NVIDIA Developer)
Read more about Arabic speech AI:
Cohere launches high accuracy Arabic transcription model (Middle East AI News)
CNTXT AI acquires Actualize to expand Arabic voice agents (Middle East AI News)
UAE-based CNTXT claims most accurate Arabic AI voice (Middle East AI News)
Saudi AI firm launches Arabic TTS leaderboard (Middle East AI News)
Arabic AI speech & Intella’s $12.5 million funding (Middle East AI News)


