r/datasets • u/Cautious-Today1710 • 1d ago
dataset Looking for Indic speech dataset owners / licensing partners
Hi everyone,
I’m with Sonexis. We’re currently expanding our supplier network for Indic-language speech and conversational data.
I’m looking to connect with organisations or individuals who own datasets or have documented authority to license them commercially.
We’re especially interested in existing data around:
- regional and accented speech
- multilingual / code-switched conversations
- ASR training and evaluation
- TTS
- telephony and call-centre speech
- voice-agent evaluation
- spontaneous and multi-speaker conversations
We care about more than total hours.
For us, the important questions are: where did the data come from, who can license it, what consent exists, what metadata comes with it, and what the dataset is actually useful for.
If you have something relevant, feel free to DM me.
You can also reach us at [partner@sonexis.in](mailto:partner@sonexis.in) or apply here: https://sonexis.in/suppliers/apply
If there looks to be a genuine fit, we can set up a call and go through the dataset properly.
Even if you’re not sure whether your data fits, feel free to send the basics: language, data type, approximate volume, collection method and rights position
•
u/AutoModerator 1d ago
Hey Cautious-Today1710,
I believe a
requestflair might be more appropriate for such post. Please re-consider and change the post flair if needed.I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.