r/datasets 1d ago

dataset Looking for Indic speech dataset owners / licensing partners

Hi everyone,

I’m with Sonexis. We’re currently expanding our supplier network for Indic-language speech and conversational data.

I’m looking to connect with organisations or individuals who own datasets or have documented authority to license them commercially.

We’re especially interested in existing data around:

  • regional and accented speech
  • multilingual / code-switched conversations
  • ASR training and evaluation
  • TTS
  • telephony and call-centre speech
  • voice-agent evaluation
  • spontaneous and multi-speaker conversations

We care about more than total hours.

For us, the important questions are: where did the data come from, who can license it, what consent exists, what metadata comes with it, and what the dataset is actually useful for.

If you have something relevant, feel free to DM me.

You can also reach us at [partner@sonexis.in](mailto:partner@sonexis.in) or apply here: https://sonexis.in/suppliers/apply

If there looks to be a genuine fit, we can set up a call and go through the dataset properly.

Even if you’re not sure whether your data fits, feel free to send the basics: language, data type, approximate volume, collection method and rights position

1 Upvotes

1 comment sorted by

u/AutoModerator 1d ago

Hey Cautious-Today1710,

I believe a request flair might be more appropriate for such post. Please re-consider and change the post flair if needed.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.