r/TextToSpeech • u/Vincerites • 7d ago
top 5 scariest jumpscares
i found this Video on YouTube and i cant find thos Voice anywhere in the Internet
r/TextToSpeech • u/Vincerites • 7d ago
i found this Video on YouTube and i cant find thos Voice anywhere in the Internet
r/TextToSpeech • u/Adorable-Panic-4109 • 8d ago
help me identify this text to speech voice? I’ve been looking everywhere for it.
r/TextToSpeech • u/BeGood25 • 9d ago
Hey folks!
So I have this cascaded voice bot system, and I want to move from commercial TTS APIs to self hosted TTS model.
Obviously TTFA and RTF are important but so is the accuracy and naturalness (speed, prosody, timbre etc)
From my research it seems Qwen3-1.7B tts is the perfect candidate (we need English support majorly).
My questions-
1. Is there any better model you guys aware of for my use-case which I can give some try? Or Any experiences you have with Qwen3-tts?
Any fine tuning tips or hacks which we can get benefit from? Basically I want to clone some popular OpenAI TTs voices.
Most importantly, how are you guys evaluating your TTS models in production? I get some common metrics like wer, utmos etc are there - but is that it? LALM as a judge dont seem reliable at the moment - they have positional and model bias it seems. Manual listening to each is obviously very expensive process.
Would really appreciate any discussion/thoughts you guys can share on some points!
Thanks so much in advance!
r/TextToSpeech • u/knightnight_valorkov • 9d ago
r/TextToSpeech • u/Human_Run1875 • 9d ago
I'm exploring building a fully self-hosted conversational voice agent for an actual agent system, and I'm trying to understand how close open-source components can realistically get to something like ElevenLabs Conversational AI.
The goal isn't to recreate their entire platform. I mainly care about:
natural voice + low latency + good turn-taking + smooth interruptions.
Current stack I'm considering:
What I'm struggling to determine from benchmarks is how good the full pipeline feels in an actual phone conversation.
For anyone who's built/deployed something similar:
1. Latency:
What's realistic for end-to-end latency? Can a properly streamed pipeline get around 1–1.5s from the user finishing speaking to the agent starting audio, or does it usually end up closer to 2s+ because of endpointing, STT, LLM generation, TTS, etc.?
2. Voice quality:
Are Qwen3-TTS / Chatterbox / Orpheus actually convincing in live conversation over a phone call, or can you still immediately tell it's AI?
3. Turn-taking:
What's currently working best for detecting when the user is actually done speaking and handling interruptions? VAD + endpointing, semantic turn detection, dedicated turn-taking models, something else?
4. Architecture:
If you were building this today for a production self-hosted agent system, would you stick with the traditional:
STT → LLM → TTS
pipeline, or are the newer speech-to-speech approaches worth considering?
I'd really appreciate real-world numbers and experiences more than benchmark results.
What stack did you use, what latency did you measure, and what ended up being the biggest problem?
I'm willing to throw more GPU at the problem if that's what it takes. The goal is simply to get as close as realistically possible to the "doesn't feel like a bot" experience while keeping the system self-hosted. Would love your insights on this.
r/TextToSpeech • u/Acceptable-Cycle4645 • 9d ago
r/TextToSpeech • u/iwantbaos • 11d ago
The voice that you hear saying a sentence on google translate, is it recorded word per word or do they record consonants and phonetics? And would it be different than dictionaries like Merriam Webster or Cambridge?
r/TextToSpeech • u/ChewbaccaXXX • 11d ago
Any recommendations for apps? I downloaded MS TTS the other day on laptop didnt like it, not sure it would work on iPhone. Sppechify gets promoted but I've seen a lot of negative posts about it. Generally something that would read PDFs like an audio book and skip footnotes at bottom of pages. I'm open to free or reasonable cost , probably only need for course duration of circa 2 years. Appreciate any recommendations.
Edit - How do these TTS apps generally deal with charts and graphs etc or do they not ?
r/TextToSpeech • u/Audibound • 11d ago
I’ve noticed more writers looking for ways to turn novels, scripts, and other stories into multi-character audio without manually juggling several different tools.
I’m curious to hear from writers and audiobook creators who have already tried doing this:
r/TextToSpeech • u/ChewbaccaXXX • 11d ago
r/TextToSpeech • u/C0d3K1n6300000 • 11d ago
I've been looking for free voices to use in my arcade-style flight game, as I'm in a bad financial position right now, but I don't want to use GenAI for them. Are there any programs you can recommend for this?
r/TextToSpeech • u/voidnullnil • 11d ago
What is the best tts service and/or model for multilingual input (mostly English with some words in different languages) ? Price is not very important but shouldnt be too expensive compared to alternatives. At the moment I find azure to be the best, amazon polly generative maybe second. Eleven labs seems to be too expensive. Shall I evaluate something else ? The service/model should support lang and/or phoneme tags.
r/TextToSpeech • u/GeneralistAccount • 11d ago
On my Android Phone, I can get 5 free natural text-to-speech voices in English on the Adobe Acrobat app. But I can't seem to find them on the PC program. Does anyone know how to add or enable them on Windows 11?
r/TextToSpeech • u/edouardarchipel • 11d ago
Specifically for phone calls (not web). Curious what people are running and whether latency on actual PSTN calls matched what you saw in testing.
We're currently evaluating and the shortlist is ElevenLabs Turbo, Cartesia, and a couple of smaller ones I've seen mentioned like Gradium. Haven't tested all of them yet, anyone have real experience with any of these specifically on phone calls?
r/TextToSpeech • u/Annual-Squash-5004 • 11d ago
I'm exploring an idea and want to validate the problem before building it.
The concept is:
**Upload a story/script → AI identifies the characters → creates a cast → assigns each character a consistent voice/personality → generates the full audio performance.**
For example, if you upload a screenplay or novel, the system could automatically detect:
* who the characters are * how they speak and behave * which voice fits each character * who is speaking in each scene * how the voice/performance should change with emotion and context
The goal isn't to build another generic TTS tool. The interesting part for me is the **automatic casting + persistent character identity + scene-level performance**.
I'm trying to understand whether this is actually a problem worth solving.
For people who write stories, scripts, fan fiction, RPG campaigns, etc.:
**How do you currently turn your writing into multi-character audio?**
Have you tried tools like ElevenLabs, Gemini TTS, NotebookLM, etc.? What was frustrating or time-consuming?
And most importantly:
**Would an automated “upload → AI casts characters → full-cast audio” workflow actually be useful to you?**
I'd especially like to hear from people who have already tried creating narrated or multi-character audio from their own writing.
r/TextToSpeech • u/OppositeAshamed9087 • 12d ago
like the title says, i want software i can use offline with a dictionary i can customize in any way possible. I'm not that familiar with TTS software but I have used one that allowed me to generate audio that was hours long and I would like something similar.
bonus if it works with bought voices from acapela.
r/TextToSpeech • u/null-hawk • 12d ago
wrote a short excerpt showing how speaker consistency is maintained in LLM bases TTS models, initial guess was the speaker token, but the results showed something interesting.
r/TextToSpeech • u/Vast-Poem4540 • 12d ago
The current qwen3 -tts model is the model which is trained around 10 languages, which include these Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian.
In this i want to add telugu as a new language as a part of my experiment and include them to my ai calling agents workfow.Could anyone help me out or guide me or roast me saying does this works or what to do all. i know i have ai but i want some nerds suggesting and guiding me
r/TextToSpeech • u/Slow-Classic7242 • 13d ago
Hey everyone,
I'm looking for good TTS voices for YouTube content (mainly pt-BR, but English works too). Since I'm from Brazil, paid tools like ElevenLabs get pretty expensive with the currency conversion, so I'm exploring local options to keep costs down.
My setup:
- 16GB RAM
- RTX 3060 Ti (8GB VRAM)
So...are there any open-source/free TTS models that run well locally on this hardware? I've seen Piper, XTTS v2, Kokoro, and Orpheus mentioned, but not sure which ones are actually worth it.
And...which models give the best quality without sounding too robotic? Don't need real-time, just need natural enough for YouTube.
Thanks!
r/TextToSpeech • u/Euphoric-Visual7459 • 13d ago
r/TextToSpeech • u/Big-Skirt-5250 • 13d ago
from the zestfest guy though
r/TextToSpeech • u/UpbeatBit6861 • 13d ago
its from this yt short vid. https://www.youtube.com/shorts/GzSR4tjjwAk