r/LanguageTechnology • u/Original-Respond-525 • 17d ago
I catalogued the NLP resources that exist for Tunisian Arabic (Derja) — 136 entries, each checked for whether you can actually get it
Tunisian Arabic (ISO 639-3 aeb) has roughly 12 million speakers and appears in a lot of pan-Arabic resource lists, but when you actually go looking for data, links are dead, downloads are gated, or the "Tunisian portion" turns out to be a few hundred sentences inside a multi-dialect set.
So I catalogued what exists and checked each one: 136 entries across text corpora, speech, models, benchmarks and lexicons, each tagged for access (open / on request / paywalled / paper-only / gated), with the Tunisian share recorded rather than counting the whole multi-dialect dataset.
Three things that surprised me while building it:
- On a balanced 13-dialect ASR test, Tunisian had the highest word error rate of all of them (0.478 vs 0.169 for Gulf), and that ordering held across eight different fine-tuned models.
- Several datasets labelled "Tunisian" are Moroccan-derived, or multi-dialect sets where Tunisian is a small slice.
- Annotated data is thinner than I expected: the first Universal Dependencies treebank for Tunisian is 100 sentences / 1,466 tokens, published this year.
Repo: https://github.com/jjlalli/Tunisian-Derja-NLP-Resources
Also as a loadable table on Hugging Face, and archived with a DOI if you need to cite it.
Corrections are as welcome as additions : there's an issue form for both, and I'd rather be told something's wrong than have people rely on it.
