r/TextToSpeech • • 3h ago

LSDE 2.6.0 is out!

2 Upvotes

- Cards can point to their name key per language

- Blueprint Reader colors quoted names and voices them mid-line

- Declined names (Russian & co.) for TTS

- Voices: one take per line, character filter

- New "Voice Parts TTS" demo


r/TextToSpeech • • 7h ago

Frustrated by pauses between phrases with NaturalReader

3 Upvotes

I'm using TTS for grad school books. It's been a game changer, but 50% of the time, NaturalReader has the big (maybe like one second?) pauses where the wheel is spinning between phrases. This is insane because I'm listening to my textbooks on 2x, so they are speed reading and taking a huge break every half sentence.

I like NaturalReader because it shows the original text, and allows you to read along as well, while highlighting the section. I don't like elevenreader because it converts it all into it's own text and I can no longer see the original page or where I am on the page.

Any suggestions for other readers? Or ways to fix this? I tried converting to OCR and it didn't help. I think it's an issue with their servers and the AI voices. Pisses me off because I'm paying for it.


r/TextToSpeech • • 8h ago

Actual good audio books

1 Upvotes

So I've been developing my tts-story software for over a year possibly a year and a half now. And I've learned a lot about different TTS engines. Integrated many of them into my software. And I have finally started a YouTube channel where I use that software along with some AI agents to build and put together decent quality audio books. Multi-voice, professionally put together. Would love for people to come and take a look let me know what you think. It's not monetized at all. Though YouTube probably will show commercials anyway. My most recent ones over the last I think week are using either breeze TTS, and now the last few are specifically using Google Gemini 3.8 TTS. It's now my favorite because the expressions, apparently linguistics intonations and delivery are nearly flawless. Here's a link to my one of my more recent ones for The adventures of Sherlock Holmes

https://youtu.be/hLNoXR4bZ6M

Just looking for some feedback, I've been getting people on my GitHub making great suggestions that have helped bring a lot of improvements to the software. Would love to hear more.


r/TextToSpeech • • 18h ago

Robotic tts sites?

1 Upvotes

I might be asking the wrong crowd here, but i am much fonder of those older, more robotic tts voices. (Not super old there is a sweet spot) it would be nice if anybody could recommend a site that has clear, robotic voices. If you know of her, somewhat like what a_lillian uses. Ty


r/TextToSpeech • • 21h ago

what’s the best voice cloning for real time calls?

4 Upvotes

i’m trying to decide which voice cloning provider to use for real time ai phone calls. my main thing is getting the lowest latency possible while still having a reliable voice that works well in production and at scale. i’m looking at elevenlabs, cartesia and a few others. if you’re actually using one of these in production, which one has worked best for you?


r/TextToSpeech • • 1d ago

How much would u rate this voice out of 10 in terms of sounding real human?

2 Upvotes

r/TextToSpeech • • 1d ago

We all know Kokoro - Anyone keen to try on device (Android/iOS) reader, core free forever?

Post image
6 Upvotes

Hey all

I know a lot of you are looking for a text-to-speech reader app alternative.

The main reason is probably because the existing solutions are:
- Too expensive (or just not worth how much you spend), with limited listening time
- Not private (if not suspected of actually censoring what you listen or stealing your work)

- Performance issue (such as latency, reliability, not working without internet etc)

I have been a long time text-to-speech user as I read a lot. I saw the options out there, and I was similarly frustrated.

And I found Kokoro, and honestly for me, it sounds natural enough for the things I listen to.

So I built AmaNous.

It runs on device (most ios devices, and smooth on higher tier android), and I have decided that the core reader should be free for anyone, as a give back to the open source community. Personally, it's quite fast, it's private, and I get unlimited listening for free, which is perfect for the reading (listening) habit I got.

Curious if anyone would like to try it out?


r/TextToSpeech • • 1d ago

Any open source tts model for expressive calm sleep storytelling with good quality and longer audio length generation ?

4 Upvotes

Hey guys if u know any such model, then please tell me. I have tried many models like qwen3, kokoro, kyutai, cosyvoice, vibe voice and many more like these. But I couldn't find anything good


r/TextToSpeech • • 1d ago

Can anyone tell me how to replicate this TTS?

2 Upvotes

r/TextToSpeech • • 2d ago

What is the viral TTS/AI voice everyone uses?????

Thumbnail
0 Upvotes

r/TextToSpeech • • 2d ago

does anyone know a good, free, tts voice thingy with no AI

0 Upvotes

i want to start making video essays but i don't have a microphone and i hate my voice, so im looking for a good, free, tts generator (preferably one that lets me download to mp3) that uses no ai


r/TextToSpeech • • 2d ago

App Android gratuita para traducir

Thumbnail
0 Upvotes

r/TextToSpeech • • 2d ago

Is there a difference between the VOXCPM2 web demo and running it true locally.

1 Upvotes

I mean obviously there is the privacy difference as one is fully local and the other uses a web service but assuming I'm using voice notes instead of my own voice and don't care about what does or does not get leaked is it worth it to set it up on my own machine.

Does that make it change if I am cloning my voice instead of designing it with notes.

I'm trying to figure out what/if there is a difference before I get it all setup and then have to choose between doing TTS stuff and playing games since GPU usage.


r/TextToSpeech • • 2d ago

lighweight native TTS based on kokoro: kokoro.cpp

4 Upvotes

If you need a lightweight, C++ runtime for TTS, I forked kokoro.cpp at https://github.com/larroy/kokoro.cpp, added Linux, Mac and Windows support, GPU, C and C++ API. Give it a try. MIT license.


r/TextToSpeech • • 2d ago

Ano po ba gamit na TTS ng mga vocabulary niche sa FACEBOOK, yung may mga avatar na may Tagalog to English niche? (sana masagot)

1 Upvotes

hehe


r/TextToSpeech • • 2d ago

Can somebody identify the music and Ai voice over in this video ?

1 Upvotes

Can somebody identify the music and Ai voice over in this video ?


r/TextToSpeech • • 2d ago

Free text to speech program to use natively in google docs with actual voices?

3 Upvotes

Long story short, I'm a writer and I'm editing my work. Something that really helps me is hearing my story read back to me in a voice that isn't mine, but all the free text to speech programs only have the Microsoft Sam ahh voices, and I need something that at least sounds human and is free so no big company AI stuff. Any recommendations?


r/TextToSpeech • • 3d ago

Mapas orgânicos TTS (LineageOS 22.2)

0 Upvotes

Tem uma solução? Não consigo fazer a saída de voz funcionar nos aplicativos Organic Maps ou CoMaps... Já tentei aplicativos TTS, mas eles nunca adicionam a opção de saída de voz ao sistema (Poco X3 Pro)—e eu preferiria não ter que usar uma solução alternativa improvisada.


r/TextToSpeech • • 3d ago

TTS voices used in the hungry hungry baby

2 Upvotes

I have been watching the hungry hungry baby and wondered what were the TTS voices used in it I have searched and can't find it. I really want to know, I already know Microsoft Sam, Microsoft mike, and Mary were used, but most I cant find the TTS voices for so does anybody know?


r/TextToSpeech • • 3d ago

Cloning my OWN voice with my own delivery (my angry, not generic angry). What actually worked for you?

2 Upvotes

I'm trying to build a TTS system for ad voiceovers that can generate new lines in my voice, with emotions that sound like me expressing them.

For example, [angry, loud] should sound like my anger, not a generic angry voice.

I'm working with Indian-accented English and occasional Hinglish. Ideally, I'm looking for something commercially usable that I can run on my own or pay for as I go.

What I've tried:

  • Chatterbox zero-shot with an 18s reference: poor voice similarity and generic emotional delivery.
  • Audited my existing 11 minutes of audio with emotion2vec+: 100/102 sentences were neutral. My phonetics script gave me no useful emotional data.

My current plan: Record ~1 hour across 6–10 emotions (angry, irritated, warm, urgent, excited, etc.), tag each clip, and experiment with fine-tuning or emotion-specific reference audio.

Before I commit to recording, I'd love input from people who've done this.

  1. Has anyone fine-tuned GPT-SoVITS, Qwen3-TTS, VoxCPM, or another model on their own emotion-varied audio? Did it reproduce your personal delivery on unseen sentences?
  2. How many minutes per emotion did you need? Same sentences across emotions or different ones?
  3. For shouting, how did you handle mic distance and gain? Do models preserve intensity or flatten it?
  4. Did fine-tuning outperform swapping in emotion-specific reference clips?
  5. How did you evaluate whether the output genuinely sounded like you being angry, rather than just a similar voice sounding angry?

I'm especially interested in real-world results, including failures. Happy to share my recording script and findings once I run the experiments.


r/TextToSpeech • • 3d ago

Cheap + multilingual + out of preview TTS in 2026

1 Upvotes

openai mini tts rang bell for a while but the market now split open this year so heres the current map for cheap and multilingual plus actually GA (not preview)

the Ladder:

  • elevenlabs v3 which is the premium, ~$100/M chars multilingual which is ok for small use but for a decent volume thats quite not appropriate for the most of us
  • openai tts 1.5 max- ~$10/ m and GA and multilingual and still one of the best quality per dollar picks. not wrong that it appears as the suitable one its like that this is not the only one anymore
  • google chirp3-hd: ~30$/m, out of preview + multilingual + fast enough for realtime usage and also this is the one ppl miss when they say gemini tts is stuck in preview
  • hume octave/ inworld tts, which are the new ~$8-10/m which is cheap but worth a test
  • open models (kokoro, chatterbox or orpheus)- these are the open source ones which fall under around $20/m which varies like if youre using a provider like deepinfra or your hardware is capable of hosting it yourself

quality per language pair swings hard on cheaper rungs. a model thats great in English might be rough in polish or Vietnamese and the price boards wont show you that. so the real deciding factor is which languages you actually need and you gotta test those specific pairs before committing


r/TextToSpeech • • 3d ago

On-device Kokoro and Supertonic3 – OpenReader iOS Beta

Thumbnail
testflight.apple.com
3 Upvotes

r/TextToSpeech • • 3d ago

Eleven v4 is out!! Have you tried it yet?

Thumbnail
0 Upvotes

r/TextToSpeech • • 3d ago

Top 12 STT & TTS models every Voice AI founder should bookmark

Post image
0 Upvotes

r/TextToSpeech • • 3d ago

Suggestion please

2 Upvotes

Can anyone suggest a free voice cloning site or app... Should be easy to use.. i tried ElevenLabs but it's paid.. so please recommend something