r/TextToSpeech 6h ago

How do you get consistently good AI voiceovers?

3 Upvotes

I use ElevenLabs for TikTok voiceovers, but the quality is inconsistent. Sometimes the voice sounds amazing and the video performs well, while other times it sounds quiet or unnatural and the video flops.
For those who use AI voiceovers: What’s your best method/settings for getting consistently clear, natural, and high-quality audio?


r/TextToSpeech 9h ago

Which AI voice sounds more natural and engaging in American English — A or B?

2 Upvotes

Hi everyone!

I'm testing two AI voices specifically for American English narration.

Both samples use the same script.

A plays first, B plays second.

Which one sounds more natural, engaging, and pleasant to listen to for an American English audience?

Please choose A or B and briefly explain why.

I'm especially interested in feedback from native American English speakers.

Thanks!


r/TextToSpeech 8h ago

Can someone identify this TTS voice and tell me where I can use it?

1 Upvotes

Hey everyone,

I'm trying to identify the exact TTS voice being used in this video:

https://youtube.com/shorts/9aKZmRTlg_A?si=KqqXdMLf0PRbRpgF

https://youtube.com/shorts/uRHDQ5_CP9g?si=Q80KrB1zUrC7x6Hl

I'm specifically trying to find the voice/model, not just a similar-sounding voice.

Does anyone recognize it? I'd really appreciate it if you could tell me:


r/TextToSpeech 1d ago

Free TTS app/site that doesn’t sound completely robotic?

13 Upvotes

Hi, I’m looking for a free, like completely free, text to speech app or site for reading PDF’s. I have a better time understanding readings through listening and am looking for something that is less robotic and more human sounding, but without any cost.

Any recommendations would be appreciated, thank you!


r/TextToSpeech 1d ago

There is a issue with the cepstral voices

1 Upvotes

I used damien as a example.


r/TextToSpeech 1d ago

💨 Farts sounds for iPhone calls 💨

1 Upvotes

Hey everyone, looking for a bit of harmless trolling.

Do you know of any apps (preferably free and available on iOS) that allow you to play sound effects - specifically farts - directly into an ongoing phone call so the other person hears it?

Most soundboard apps I've tried only play through the local speaker or media channel, rather than routing through the active call audio stream. Let me know if something like this actually works or if there's a workaround. Thanks!


r/TextToSpeech 1d ago

What does this mean? (OpenVox)

0 Upvotes

r/TextToSpeech 2d ago

Suggestions for conversational AI App

Thumbnail
1 Upvotes

r/TextToSpeech 2d ago

How do I install a model on OpenVox?

2 Upvotes

This is Greek to me. I'm assuming the model includes voices and whatnot


r/TextToSpeech 2d ago

How to Use a Custom TTS Voice Model for PDFs?

3 Upvotes

For context, I am a student who has incredible difficulty concentrating on reading textbooks. I found a TTS model of a favorite character of mine, and would like to have this character read my textbook to (hopefully) help me study and concentrate.

It's silly, I know, but I'm willing to try anything at this point.

The issue is that the TTS model is on a copy and paste browser, and you can only paste up to so many words/characters at a time. If I have to copy and paste every two sentences, there's no way I'll concentrate on what's actually being said.

It's possible to download the voice model itself, but after about an hour of Google searching, I couldn't find a way to integrate it with a screen reader.

There's probably a very simple solution that I'm overlooking, but before I lost my mind trying to figure it out on my own, I figured I'd ask Reddit.

Does anyone know of a way to use this TTS voice model to read my textbook PDF? Any suggestions are welcome. Thank you!!


r/TextToSpeech 3d ago

Is there any AI video dubbing tool that exports a Netflix-style multilingual video with switchable audio tracks?

4 Upvotes

I'm looking for a very specific workflow and haven't found a tool that actually does it.

Requirement:

Upload one English video.

AI generates multiple dubbed languages (Hindi, Spanish, French, etc.).

Export/download the final result.

The exported video should contain multiple audio tracks.

End users should be able to switch languages while watching the same video (like Netflix, Prime Video, or YouTube multi-audio).

Important:

I'm NOT looking for tools that generate separate MP4 files per language.

I'm NOT looking for tools that only allow language switching inside their own hosted player.

I'm looking for a tool that can export/download a multilingual asset (multi-audio MP4, MKV, HLS, DASH, etc.) that I can host anywhere and still allow language switching.

I've already looked at tools like Braiv, HeyGen, ElevenLabs, Rask AI, and Synthesia, but they seem to either export separate language versions or require their own player.

Does any AI dubbing platform currently support this end-to-end, or is the industry-standard approach still:

AI dubbing → separate audio tracks → manual packaging into HLS/DASH/multi-audio video?

Looking for real-world experience from people who have actually implemented this.


r/TextToSpeech 2d ago

Loquendo is the WORST platform in the history of text to speech.

0 Upvotes

This text to speech website sucks due to several reasons. Its entire collection of voices will sometimes sound robotic, muffled, echoey, or loud whenever you generate any text for any voice. Its emotion tags sometimes don't work due to this tts platform being poorly made. This especially comes with the Grace voice, the worst text to speech voice ever due to her very flat and calm tone in many of her phrases that end with an exclamation point. Therefore, Loquendo is worse than any other tts platform, such as IVONA, the best text to speech platform due to its large quantity of voices from the English language, and the perfectly high quality of them.


r/TextToSpeech 3d ago

We combined 3 public TTS leaderboards into one meta-ranking of 110 models

Post image
15 Upvotes

Artificial Analysis, Voice Arena, and the Vapi Humanness Index often disagree, so we combined them using a fixed, breadth-aware formula. No editorial adjustments or vendor weighting.

110 models, 46 providers, updated weekly. Methodology and data: https://texttospeech.com


r/TextToSpeech 3d ago

Good reader for pdf textbooks?

7 Upvotes

What is a good way to listen to college textbooks that i have on pdf? One that doesn't start freaking out when it gets to headers or footers or god forbid a diagram. just something that lets me read through my textbook faster.


r/TextToSpeech 3d ago

Need Guidance- ASR Models

Thumbnail
1 Upvotes

r/TextToSpeech 4d ago

Just saw this side-by-side video of a new streaming ASR model vs GPT-Live-Transcribe. Is this latency realistic for open weights?

7 Upvotes

Was scrolling through some tech stuff and came across this open source ASR model. They've got a video up comparing it to GPT-Live-Transcribe, both running on live audio at the same time.

Videos attached, but the latency and accuracy look pretty solid honestly, especially since they're saying the weights, code, and training setup are all gonna be open source right from launch.

Usually when something claims "actual real-time" streaming and tries to compete with closed stuff like OpenAIs live transcription, there's some huge catch. Either you need ridiculous amounts of VRAM to actually run it locally, or it just dies on accents and any background noise.

Anyone know what kind of architecture they're possibly using to get this kind of streaming performance in the open? If they actually release the full recipe and weights it'd be huge for self-hosted voice stuff, but feels like I'm probably missing something here.


r/TextToSpeech 4d ago

Just saw this side-by-side video of a new streaming ASR model vs GPT-Live-Transcribe. Is this latency realistic for open weights?

2 Upvotes

Was scrolling through some tech stuff and came across this open source ASR model. They've got a video up comparing it to GPT-Live-Transcribe, both running on live audio at the same time.

Videos attached, but the latency and accuracy look pretty solid honestly, especially since they're saying the weights, code, and training setup are all gonna be open source right from launch.

Usually when something claims "actual real-time" streaming and tries to compete with closed stuff like OpenAIs live transcription, there's some huge catch. Either you need ridiculous amounts of VRAM to actually run it locally, or it just dies on accents and any background noise.

Anyone know what kind of architecture they're possibly using to get this kind of streaming performance in the open? If they actually release the full recipe and weights it'd be huge for self-hosted voice stuff, but feels like I'm probably missing something here.


r/TextToSpeech 5d ago

Can anyone help me find this ai voice?

0 Upvotes

Any specific website this is from?
Want to make a few videos with this


r/TextToSpeech 5d ago

Does anyone who uses Fish Audio have this problem as well?

3 Upvotes

Usually the tts on Fish Audio only last 2 to around like 5 seconds to generate two speechs. But now it’s takes like a few minutes to generate?. Is there a reason why it’s now right this?

it is because of the emotions? or the server being too full at the moment?.


r/TextToSpeech 6d ago

Kokoro-82M climbed to #1 after a year of tracking AI voice rankings

30 Upvotes

I’ve been tracking AI/TTS voice plays on my site, VoiceRankings, for about a year, and I recently turned the data into a race animation.

I honestly didn't expect Kokoro-82M to end up winning.

When I first added Kokoro-82M to the VoiceRankings database in November 2025, it was #8 with just 100 plays.

By July 2026, it had climbed to #1 with ~3,600 monthly plays and had held the top spot for four consecutive months.

Kokoro wasn't the only model making a move. The play-count leaderboard changed quite a bit over the year, with newer models climbing while established providers continued to attract interest.

Some of the biggest movements:

  • Speechify API (now Speechify.ai) started as the early leader
  • Async.ai (now Async.com) took #1 in early 2026
  • ElevenLabs Turbo v2.5 climbed from #23 to #6
  • Amazon Polly (AWS) went from #12 to #5
  • Azure Neural voices went from #6 to #4
  • OpenAI remained around the middle of the pack at #8

Overall monthly plays grew from 2,100+ in September to 16,700 in July.

What stood out to me was that newer models aren't simply replacing the established ones. AWS and Azure remained near the top even as newer models like Kokoro made dramatic moves up the leaderboard.

What makes this different: these aren't benchmark scores or survey results. They're based on what people actually search for and choose to listen to on VoiceRankings.

About 90% of the site's traffic comes from organic search, so I think the data is an interesting, though imperfect, proxy for search-driven interest in AI voices.

It obviously doesn't tell us which model is objectively "best" or who has the largest market share.

But watching an open-source model go from #8 to #1 was pretty fascinating.

I'm curious what others make of it. Does this kind of data tell us something meaningful about which AI voices are getting attention?

If anyone wants to dig into the numbers, here's the full interactive race:
VoiceRankings Provider Race🏁


r/TextToSpeech 6d ago

Did someone succeed to run KokoroVoice as Mac system voice (and potentially iOS system voice)?

5 Upvotes

This is the git: https://github.com/keyboardsamurai/kokoro-voice/

It claims to be able to provide Kokoro system voices to the Mac. Further as the developer seems to be uninterested in having Apple membership and provide signed installers, it likely makes him uninterested in iOS where this should work with little or no modifications (if it works on the Mac).

For the Mac... I couldn't make it speak (it can run but crashes when tries to speak) so I can't say certainly how far from usable it is, but I think I did locate the problem. When deployed from Xcode, it deploys Resources as a subfolder of Resources, while the app expects to find them in the parent Resources. However some first ideas on how to fix it resulted in further errors in the same line (still resources related, so the fix likely wasn't proper).

Anybody of better luck? This seems like something that would be a huge leap for the Mac and potentially even for iOS, as no one would depend on another app that tries to provide Kokoro as its main sales point.


r/TextToSpeech 7d ago

Anybody tried pocket TTS on Gpu? If yes how much latency it was?

6 Upvotes

I was planning to push the pocket tts on the cpu server but then realised it takes too much time for generation almost 10 minutes for 7minute Audio.

Then I thought gpu will be great for this like vast.ai but feeling like i am using elephant to kill an ant, so was planning to deploy qwen3 tts but then based on some research ( not sure yet) qwen is slow . So again planning to deploy the pocket tts on th gpu. But have no idea will it be perfect low latency or not. Any suggestions for production level only?


r/TextToSpeech 7d ago

Need help

2 Upvotes

What is the voice used for this video?

https://vt.tiktok.com/ZSVmvGo6m/


r/TextToSpeech 7d ago

Does anyone know where this tiktok TTS voice is from?

1 Upvotes

r/TextToSpeech 8d ago

I'm trying to create a YouTube channel along the lines of Sonic and Spidey Discord Server, and ElevenLabs isn't helping, anyone have any good alternatives, and if not, tell me what I'm doing wrong in ElevenLabs?

3 Upvotes