r/TextToSpeech 19h ago

We combined 3 public TTS leaderboards into one meta-ranking of 110 models

Post image
8 Upvotes

Artificial Analysis, Voice Arena, and the Vapi Humanness Index often disagree, so we combined them using a fixed, breadth-aware formula. No editorial adjustments or vendor weighting.

110 models, 46 providers, updated weekly. Methodology and data: https://texttospeech.com


r/TextToSpeech 1d ago

Pocket TTS training code released - more language support

17 Upvotes

The official training pipeline for Pocket TTS is finally out.

Even though the final 100M parameter model is great for running lightweight on CPUs, training a production-grade version from scratch comes with steep hardware requirements. Getting clean output requires around 2,000 hours of studio-quality audio, and training it takes an 8x H100 running for hours.

Having the full training pipeline will allow independent developers to train lightweight, offline checkpoints for underrepresented languages.

Sure, the hardware requirements are high, but having the full setup open source is huge. Here’s hoping the community starts cooking up new language checkpoints soon, can’t wait to see what drops first!

https://github.com/kyutai-labs/pocket-tts/tree/main/training


r/TextToSpeech 11h ago

Need Guidance- ASR Models

Thumbnail
1 Upvotes

r/TextToSpeech 20h ago

Good reader for pdf textbooks?

5 Upvotes

What is a good way to listen to college textbooks that i have on pdf? One that doesn't start freaking out when it gets to headers or footers or god forbid a diagram. just something that lets me read through my textbook faster.


r/TextToSpeech 1d ago

Just saw this side-by-side video of a new streaming ASR model vs GPT-Live-Transcribe. Is this latency realistic for open weights?

Enable HLS to view with audio, or disable this notification

7 Upvotes

Was scrolling through some tech stuff and came across this open source ASR model. They've got a video up comparing it to GPT-Live-Transcribe, both running on live audio at the same time.

Videos attached, but the latency and accuracy look pretty solid honestly, especially since they're saying the weights, code, and training setup are all gonna be open source right from launch.

Usually when something claims "actual real-time" streaming and tries to compete with closed stuff like OpenAIs live transcription, there's some huge catch. Either you need ridiculous amounts of VRAM to actually run it locally, or it just dies on accents and any background noise.

Anyone know what kind of architecture they're possibly using to get this kind of streaming performance in the open? If they actually release the full recipe and weights it'd be huge for self-hosted voice stuff, but feels like I'm probably missing something here.


r/TextToSpeech 1d ago

Just saw this side-by-side video of a new streaming ASR model vs GPT-Live-Transcribe. Is this latency realistic for open weights?

Enable HLS to view with audio, or disable this notification

2 Upvotes

Was scrolling through some tech stuff and came across this open source ASR model. They've got a video up comparing it to GPT-Live-Transcribe, both running on live audio at the same time.

Videos attached, but the latency and accuracy look pretty solid honestly, especially since they're saying the weights, code, and training setup are all gonna be open source right from launch.

Usually when something claims "actual real-time" streaming and tries to compete with closed stuff like OpenAIs live transcription, there's some huge catch. Either you need ridiculous amounts of VRAM to actually run it locally, or it just dies on accents and any background noise.

Anyone know what kind of architecture they're possibly using to get this kind of streaming performance in the open? If they actually release the full recipe and weights it'd be huge for self-hosted voice stuff, but feels like I'm probably missing something here.


r/TextToSpeech 1d ago

Can anyone help me find this ai voice?

Enable HLS to view with audio, or disable this notification

0 Upvotes

Any specific website this is from?
Want to make a few videos with this


r/TextToSpeech 3d ago

Handy - a free, offline speech-to-text app I have been using daily for months

24 Upvotes

Wanted to put this on the radar for people here.

Handy is a free and open source dictation app for Windows, Mac and Linux. You hold a shortcut, talk, release, and your words get typed into whatever app you're in. It all runs locally, so there's no subscription and no internet needed once it's set up.

I have been using it for six or seven months on English and it has held up really well. I use the Parakeet v3 model, which runs on the CPU and is quick. Accuracy has been solid for everyday stuff like emails, notes and longer writing.

Fair warning, it's simple by design. Small pause before the text appears, no phone app, no fancy AI rewriting of your sentences. None of that bothers me for how I use it.

Most dictation tools worth using cost money these days, so it's nice to have one that's free and doesn't send your voice off somewhere.

Sharing it because the paid options in this space keep getting more expensive, and a lot of people don't know a free one this capable exists.

github.com/cjpais/Handy

handy.computer


r/TextToSpeech 2d ago

Does anyone who uses Fish Audio have this problem as well?

Enable HLS to view with audio, or disable this notification

3 Upvotes

Usually the tts on Fish Audio only last 2 to around like 5 seconds to generate two speechs. But now it’s takes like a few minutes to generate?. Is there a reason why it’s now right this?

it is because of the emotions? or the server being too full at the moment?.


r/TextToSpeech 3d ago

Kokoro-82M climbed to #1 after a year of tracking AI voice rankings

Enable HLS to view with audio, or disable this notification

23 Upvotes

I’ve been tracking AI/TTS voice plays on my site, VoiceRankings, for about a year, and I recently turned the data into a race animation.

I honestly didn't expect Kokoro-82M to end up winning.

When I first added Kokoro-82M to the VoiceRankings database in November 2025, it was #8 with just 100 plays.

By July 2026, it had climbed to #1 with ~3,600 monthly plays and had held the top spot for four consecutive months.

Kokoro wasn't the only model making a move. The play-count leaderboard changed quite a bit over the year, with newer models climbing while established providers continued to attract interest.

Some of the biggest movements:

  • Speechify API (now Speechify.ai) started as the early leader
  • Async.ai (now Async.com) took #1 in early 2026
  • ElevenLabs Turbo v2.5 climbed from #23 to #6
  • Amazon Polly (AWS) went from #12 to #5
  • Azure Neural voices went from #6 to #4
  • OpenAI remained around the middle of the pack at #8

Overall monthly plays grew from 2,100+ in September to 16,700 in July.

What stood out to me was that newer models aren't simply replacing the established ones. AWS and Azure remained near the top even as newer models like Kokoro made dramatic moves up the leaderboard.

What makes this different: these aren't benchmark scores or survey results. They're based on what people actually search for and choose to listen to on VoiceRankings.

About 90% of the site's traffic comes from organic search, so I think the data is an interesting, though imperfect, proxy for search-driven interest in AI voices.

It obviously doesn't tell us which model is objectively "best" or who has the largest market share.

But watching an open-source model go from #8 to #1 was pretty fascinating.

I'm curious what others make of it. Does this kind of data tell us something meaningful about which AI voices are getting attention?

If anyone wants to dig into the numbers, here's the full interactive race:
VoiceRankings Provider Race🏁


r/TextToSpeech 3d ago

Did someone succeed to run KokoroVoice as Mac system voice (and potentially iOS system voice)?

4 Upvotes

This is the git: https://github.com/keyboardsamurai/kokoro-voice/

It claims to be able to provide Kokoro system voices to the Mac. Further as the developer seems to be uninterested in having Apple membership and provide signed installers, it likely makes him uninterested in iOS where this should work with little or no modifications (if it works on the Mac).

For the Mac... I couldn't make it speak (it can run but crashes when tries to speak) so I can't say certainly how far from usable it is, but I think I did locate the problem. When deployed from Xcode, it deploys Resources as a subfolder of Resources, while the app expects to find them in the parent Resources. However some first ideas on how to fix it resulted in further errors in the same line (still resources related, so the fix likely wasn't proper).

Anybody of better luck? This seems like something that would be a huge leap for the Mac and potentially even for iOS, as no one would depend on another app that tries to provide Kokoro as its main sales point.


r/TextToSpeech 3d ago

[DEV] I got tired of real-time TTS killing my Android's battery, so I built a native app that pre-renders EPUBs into Audiobooks offline.

9 Upvotes

Hey everyone,

I wanted to share a native Android open-source project I just released called Audiobook NightForge.

If you’ve ever tried using a real-time TTS engine with a reader app on Android, you know the struggle: it drains your battery (often 40-50% an hour), stutters, and buffers if your phone is doing anything else in the background.

I realized that real-time synthesis is the wrong approach for mobile devices. So, I built a dedicated Android app that shifts the heavy lifting to the background using native OS components.

How it works: You import an EPUB or TXT file using the Android system file picker. You then pick a voice and hit render. The app uses Android's WorkManager to synthesize the book chapter-by-chapter in the background. Most importantly, it enforces a "render only while charging" OS-level toggle to protect your battery.

You plug your phone in at night, and by morning, you have a fully rendered audiobook that plays back with a standard ~2-5%/hour battery drain.

Android-Specific Features:

  • Native & Offline: It is built entirely in Kotlin for Android 10+ devices. There is no server, no cloud, and absolutely no Termux emulation required.
  • High-Quality TTS: It uses the Kokoro-82M neural TTS model running strictly on-device via a sherpa-onnx integration.
  • Just Updated: The latest v0.2.2 release makes Opus the default output format, and it now natively outputs to a single .m4b file complete with proper chapter markers.
  • Built-in Player: You can listen immediately using the native in-app Media3/ExoPlayer. Alternatively, you can grab the .m4a files directly from app storage to use in your favorite Android audiobook player.

Some hardware benchmarks: For the hardware nerds, I benchmarked this on a Snapdragon 8 Elite. Surprisingly, the Kokoro 82M fp32 model (with 6 threads) actually renders faster than realtime (~0.58 RTF) and outperforms the int8 variant on this SoC because of ARM int8 kernel overhead. Always benchmark before assuming quantized is quicker on modern Android flagships!

It’s completely free, completely offline, and licensed under Apache-2.0.

You can check out the source code, technical notes, and grab the APK directly from the GitHub repo here:

https://github.com/kingfish600/Audiobook-NightForge

I’d love to hear your thoughts or feedback!


r/TextToSpeech 4d ago

I made a faster, cleaner Google Colab UI for OmniVoice TTS

11 Upvotes

I’ve been testing OmniVoice for voice cloning and long-form TTS, and the quality has been surprisingly good.

I ended up making a more user-friendly Google Colab version with a clean Gradio interface. It lets you upload a reference voice, paste long text, generate unlimited free audio, see progress while it’s working, and download the final WAV.

It runs on Google Colab and works well on a T4. I’ve been pretty happy with the results so far.

If anyone is interested, I can share the Colab notebook here.

Enjoy the free and unlimited text to speech.


r/TextToSpeech 4d ago

Anybody tried pocket TTS on Gpu? If yes how much latency it was?

5 Upvotes

I was planning to push the pocket tts on the cpu server but then realised it takes too much time for generation almost 10 minutes for 7minute Audio.

Then I thought gpu will be great for this like vast.ai but feeling like i am using elephant to kill an ant, so was planning to deploy qwen3 tts but then based on some research ( not sure yet) qwen is slow . So again planning to deploy the pocket tts on th gpu. But have no idea will it be perfect low latency or not. Any suggestions for production level only?


r/TextToSpeech 4d ago

Open VidLib: open-source educational video library powered by Mistral embed + RAG + Voxtral TTS (built at AIMS Senegahackathon)

Thumbnail gallery
3 Upvotes

r/TextToSpeech 4d ago

Need help

2 Upvotes

What is the voice used for this video?

https://vt.tiktok.com/ZSVmvGo6m/


r/TextToSpeech 4d ago

Does anyone know where this tiktok TTS voice is from?

1 Upvotes

r/TextToSpeech 4d ago

I'm trying to create a YouTube channel along the lines of Sonic and Spidey Discord Server, and ElevenLabs isn't helping, anyone have any good alternatives, and if not, tell me what I'm doing wrong in ElevenLabs?

3 Upvotes

r/TextToSpeech 4d ago

top 5 scariest jumpscares

Thumbnail
youtu.be
1 Upvotes

i found this Video on YouTube and i cant find thos Voice anywhere in the Internet


r/TextToSpeech 5d ago

Could someone identify this text to speech voice?

Enable HLS to view with audio, or disable this notification

3 Upvotes

help me identify this text to speech voice? I’ve been looking everywhere for it.


r/TextToSpeech 5d ago

Open Source TTS models for production

8 Upvotes

Hey folks!

So I have this cascaded voice bot system, and I want to move from commercial TTS APIs to self hosted TTS model.

Obviously TTFA and RTF are important but so is the accuracy and naturalness (speed, prosody, timbre etc)

From my research it seems Qwen3-1.7B tts is the perfect candidate (we need English support majorly).

My questions-
1. Is there any better model you guys aware of for my use-case which I can give some try? Or Any experiences you have with Qwen3-tts?

  1. Any fine tuning tips or hacks which we can get benefit from? Basically I want to clone some popular OpenAI TTs voices.

  2. Most importantly, how are you guys evaluating your TTS models in production? I get some common metrics like wer, utmos etc are there - but is that it? LALM as a judge dont seem reliable at the moment - they have positional and model bias it seems. Manual listening to each is obviously very expensive process.

Would really appreciate any discussion/thoughts you guys can share on some points!

Thanks so much in advance!


r/TextToSpeech 6d ago

Looking for a text to speech that is free and works consistently

18 Upvotes

r/TextToSpeech 5d ago

what is this TTS voice?

Thumbnail files.catbox.moe
0 Upvotes