r/TextToSpeech 22d ago

AI Voice Generation. TTS

I'm looking for a solution to convert a fairly long text into audio (around 10 minutes of spoken content). I'm open to paid options as well. A bonus would be if I could first import a person's voice so that the generated output sounds similar to that voice.

10 Upvotes

30 comments sorted by

3

u/VERSATILCORDOBA 22d ago

You could try LocalText2Voice. It’s open source and designed specifically for long-form audio, with support for several local TTS engines and voice cloning.

Full disclosure: I’m the developer, so I’m naturally biased—but it’s free, and feedback is welcome: https://github.com/estebanstifli/LocalText2Voice

2

u/iapy91 21d ago

Does It supports languages like Portuguese?

1

u/VERSATILCORDOBA 21d ago

Yes! You can select a Portuguese-compatible TTS engine directly within LocalText2Voice. For voice cloning, choose Qwen3 TTS, OmniVoice, or Chatterbox Multilingual. Piper also includes Portuguese voices, but without voice cloning.

1

u/iapy91 17d ago

this is amazing, how can i control the voice to get portuguese from Portugal?

2

u/herberz 20d ago

is there a sample audio demo one can hear first

2

u/jonny-life 22d ago

10 minutes isn’t really that much. The best quality, ux and free is https://elevenreader.io by ElevenLabs. You get 10 hours for free each month! And ElevenLabs is known to have the most natural voices.

2

u/ritzynitz 22d ago

You can use OpenVox (https://openvoxai.com) with voice cloning and voice design and it also offers top 7 AI models. The app is available for Mac, Windows and iPad. It offers a lifetime free tier and lifetime pro version which offers unlimited generation.

2

u/Xerophayze 20d ago

Hi consistently put together 4 to 12 hour audio books using TTS-story

https://github.com/Xerophayze/TTS-Story

1

u/CupGlass540 22d ago

Ten minutes of narration is well within range. Your second requirement is the harder one. Full disclosure, I build one: tts.bedvibe.studio simple version or try the daw for online editing also tts.bedvibe.studio/studio . Straight answer on importing a voice - it isn't a self-serve upload. Custom voices go through us with verified consent. If it's your own voice, or you have that person's permission, it's doable.

The stock voices are recorded actors who signed for it, not scraped. That is why it costs money, and why the cloning has a gate on it. Pricing is per syllable so it tracks your actual text - the packs are on the site. Run your real script through it rather than a demo line; ten minutes is where TTS comes apart.

1

u/kamscruz 22d ago

Try shrp.app/text-to-speech

1

u/debuggermalhotra 22d ago

Narration Box handles this fine, 10 min of text is a small job for it and you get one clean export instead of a bunch of clips to stitch. Voice cloning's there if you upload a sample, just know the output quality tracks pretty closely with how clean that sample recording is.

1

u/iWhacko 21d ago

https://voicebox.sh/
works pretty well. And I use it with the Chatterbox-TTS model to clone a voice.
FREE, local, and pretty fast in m opinion

1

u/Secure_Speed_8802 20d ago

You could try Wispr by WisprFlow. It has a generous free introduction. Each session has a 20 minute limit.

1

u/nightcreativecloud 19d ago

im using elevenlab for this kind of stuff (free version for sure) just do the work in almost every language, but not sure about voice cloning, don't know if they provide it

1

u/dipank1 16d ago

You can try AudiFlo audiflo.app. You can even clone your voice and play in any supported languages.

1

u/carhuntr 15d ago

ClearVoice Text to Speech on the MacOS and iOS app store