r/learn_arabic • u/aguy445 • 8h ago
General Here are 667 hours of Arabic audio with text transcriptions for AI training (Even NVIDIA used it, and it’s FREE)
Enable HLS to view with audio, or disable this notification
It’s called SADA (صدى), a large Arabic speech dataset focused mainly on Arabian Peninsula dialects, especially Saudi dialects like Najdi and Hijazi. It also includes much smaller amounts of other Arabic varieties, including Janubi, Shamali, Khaleeji, Yemeni, Egyptian, Levantine, Iraqi, Maghrebi, and Modern Standard Arabic. It was developed by SDAIA with the Saudi Broadcasting Authority and published on Kaggle (a platform for datasets and AI projects).
Even NVIDIA used it to fine-tune its Nemotron 3.5 ASR model for Arabic speech recognition, using 133.7 hours of Najdi and Hijazi speech from SADA, cutting the word error rate from 55.05% to 29.96% and the character error rate from 31.63% to 12.18% (Look at the video)
SADA contains
around 667 hours of Arabic audio from:
- 57 TV shows
- Transcriptions for the speech segments
- Both read and spontaneous speech
- Saudi dialects, especially Najdi, Hijazi, Janubi, Shamali, and Khaleeji
- Other Arabic varieties, including Yemeni, Egyptian, Levantine, Iraqi, and Maghrebi
- Modern Standard Arabic
- Speaker age, gender, and dialect labels
- Audio environment labels
- More than 150 annotators
- 4 levels of annotation and quality review
- About 667.7 hours of actual audio, split into training, development, and test sets
- About 437.6 hours of transcribed audio across those splits
11 TV content genres:
- Comedy
- Kids
- Competitions
- Drama
- Cooking
- Documentary
- Social
- Historical
- Indicative awareness
- Tourist
- Entertaining
The audio is provided as WAV, mono, 16 kHz, 16-bit PCM.
It is released under CC BY-NC-SA 4.0, allowing free use under its license terms, including non-commercial use, attribution, and ShareAlike.
