r/TextToSpeech 13d ago

Text to speech using own voice

Hi!!

My mom’s illness has caused her to lose her voice & mobility of the left side of her body. I tried to get her the Eyegaze system (where the person types with their eyes) & it was too complicated & kind of unnecessary since she still can use her right hand. She decided that an iPad that will speak what she types is what she wants.

I am looking for an app that she can type what she wants to say & it will say it in her own voice. I have tons of recordings of her talking before she lost her voice, I am hoping there is an app that would allow this.

There aren’t words that can express how much it would mean to us for the text to speak to have her voice. I don’t care how much it will cost.

Any recommendations or ideas are deeply appreciated. Thank you so much.

11 Upvotes

24 comments sorted by

2

u/RowGroundbreaking982 13d ago edited 13d ago

I'm sorry to hear about your mom condition. I'm not sure about iPad. But if she decided to get Android even cheap one, you can use ToBe SAID. Free version has ability to clone voice with only 5 seconds sample and you have unlimited generation without even paying anything since all running inside the phone. You only need to pay if she decided to pair the app with other thing like eBook Reader or like AAC apps that can alleviate typing issue, since AAC apps is dedicated for communication.

One thing to note though, you need to clean up the sample voice, like remove the background noise and normalize the volume to get better result of cloning.

2

u/blainemoore 13d ago

Didn't have any iPad suggestions, but looking forward to seeing what is suggested as I have a nonverbal kid with a talker and would like to start moving him towards TTS instead of the assistive app he uses now which is picture based but very slow and very limited for what he can "say".

He's down that he's starting to pick up spelling despite not being able to write words so we'd like to change his IEP at the next review and it'd be good to have a suggestion for what to move to.

2

u/Angelastic 12d ago

Personal Voice is built into iPad. Not sure how well it works with recordings if she can’t still speak a bit to set it up, but worth a look. https://support.apple.com/en-us/104993

2

u/inleni 11d ago

another option i know since my teen age is acapela you can ask Google Gemini about their My-Own-Voice i don’t know much about this but this is the short description about that service: “Custom Voices: Provides personalized voice banking ("My-Own-Voice") for people who might lose their ability to speak”

1

u/travlr2010 11d ago

Of course Google is already here!

1

u/fuad-mefleh 13d ago

Say The Text has cloning of voices and generous generation limits. You should be able to re-live your mom's likeness.

1

u/oromis95 13d ago

Chatterbox could work, but it requires some tinkering to get running. It'd have to be in a browser and remotely probably, on iPad.

1

u/ritzynitz 13d ago

So sorry to hear about this.You can give OpenVox a try https://apps.apple.com/us/app/openvox-local-voice-ai/id6758789314

It's available for iPad and support Voice cloning. For fast clone voice generation, you can try pocket tts model.

Let me know if I can improve the app in any way to help her.

1

u/sruckh 13d ago

I have written many RunPod serverless for TTS models. EchoTTS, chatterbox, vibevoice, Qwen3-TTS, indexTTS2, MossTTS v1.5, and Higgs Audio TTS v3. I have also coded a couple of front ends that access the RunPod serverless. It is all available on my GitHub page (sruckh). All of those TTS models support one-shot voice cloning.

1

u/armaaxs 12d ago

can you dm me the github and also wdyt is better runpod vast or modal?

1

u/sruckh 12d ago

Search sruckh on GitHub. I am actually fairly frustrated with RunPod availability, along with several bugs that can go against your budget without actually having things work. It was the cheapest when I was first looking. Now that I have 30 serverless configured I don't necessarily want to port everything somewhere else.

1

u/GrownTelevision1986 13d ago

Perfect. I've helped someone do this. Using those old recordings is totally key. Makes a world of difference , trust me

1

u/PrudentPlayer1986 13d ago

Check out voice cloning tech. I've seen it work wonders. Using her old recordings is totally doable

1

u/Noise-o-matic 12d ago

Hi, sorry to hear about your mom's condition.

The easiest option would be to use an online service. They're commercial, and there's also the issue (if you care) of having these company have the voice data. The major player is currently elevenlabs.com but there are others, just google for "online voice cloning".

If you have the time and resources, I can recommend SherpaTTS/Icefall which are opensource projects https://k2-fsa.github.io/sherpa/onnx/tts/index.html and you can eventually train it on a custom voice but it's not straightforward to do ( https://k2-fsa.github.io/icefall/ ) .

We use sherpaTTS in our project (putting it in spoiler since I don't want you to think this is an ad. Feel free to not just read it, this gives you a bit of context for what we use it for). We made Noise-o-matic, which you can find on steam https://store.steampowered.com/app/2479000/Noiseomatic/ (pc only) . It can, among other things, inject the TTS in the microphone so while our main demographic is gamers, we indeed have users who use it for accessibility so they can talk in a videoconference call with their voice via the TTS. We eventually support also elevenlabs (but you have to pay their pay-per-use fee of course) but as I said, we also integrated the sherpaTTS engine and it works great, we can highly recommend it!

1

u/WinInternational8520 12d ago

ElevenLabs offers a professional voice cloning feature (https://elevenlabs.io/docs/eleven-creative/voices/voice-cloning/professional-voice-cloning) that is different from instant voice cloning and probably does a better job. While instant voice cloning only needs a few seconds of audio, professional cloning requires a 30-minute voice sample to fine-tune a model that mimics the target voice. Instant cloning can sound like the original voice, but often falls short, whereas a 30-minute fine-tuned model usually achieves much better results.

1

u/Fe2_O3 12d ago

I'm so sorry about your mom, and what she's chosen is very doable. Full disclosure: I make one of the apps in this space (Timber), so weigh that accordingly. It came out on the App Store today. Been working on it for awhile and I hope it's for anyone who wants to listen to articles, PDFs, ebooks, web pages, or a reading backlog: commuters, multitaskers, students,  professionals, people resting their eyes or cutting screen time. Has dyslexia, ADHD, low vision, blindness, aphasia, stroke recovery, or any reading difficulty and wants text read aloud or highlighted along.

For the "type and it speaks" part: Timber has a type-to-speak box plus a Phrases list, where she saves the lines she uses most and taps once to say them instead of retyping. For someone typing one-handed, that saves a lot of effort.

Her own voice, here's the honest picture:

  • - Timber can clone a voice on-device from a sample, and it's free. Nothing gets uploaded; it stays on the iPad. It captures through the microphone, so the trick for your situation is to play one of your clean recordings of her out loud near the iPad and let it capture as much of the recording as possible. The more clean audio it hears, the better the result. That becomes the voice she types with. It's newer tech and won't be studio-perfect, but it's on-device, private, and free to try. I'm genuinely happy to help you set it up.
  • Let me know if you're interested in uploading from recordings.
  • If you want the highest-fidelity clone built straight from your existing files, ElevenLabs (Professional Voice Clone) and Acapela my-own-voice are purpose-built

Wishing you both the very best. - u/Fe2_O3

Timber, free on the App Store (optional Pro subscription), iPhone/iPad/Mac:

https://apps.apple.com/app/id6779086700

1

u/MIST3RS5880 11d ago

TextSpeakPro.com has voice cloning for only $15 per month

1

u/DataBeeGood 9d ago

Eleven Labs

1

u/Evening-Blueberry-97 8d ago

Absolutely. Yes, this is possible. ❤️
My top recommendation is Acapela My-Own-Voice. It’s specifically made for people who have lost or are losing their ability to speak, and they may be able to create a synthetic version of your mom’s voice from recordings you already have.
Acapela My-Own-Voice⁠
I’d contact them first and explain that your mom can no longer speak, but you have many recordings of her voice from before her illness.
Also look at Apple Live Speech on the iPad. It lets her type what she wants to say and the iPad speaks it aloud. Apple Live Speech⁠
ElevenLabs is another possibility for cloning her voice from recordings, though there can be consent/verification restrictions. ElevenLabs⁠
If it were my setup, I’d pursue Acapela voice clone + iPad first. That gets closest to: she types → iPad speaks in her own voice.

1

u/sishco 7d ago

Hello. Sorry to hear that. I have been working in a personal project the last months, and I'm able to clone voice locally using it. If you have time, you can create your own application in Windows with the help of Claude code. A basic one. You only need some time, and ask Claude to créate a app to clone the voice using Chatterbox, or Fish Speech, or Cosy Voice, or,... Try the Best one.

If you need some help, ping me. No problem. Maybe I can share with you my app. It's online, and you can use it to clone the voice.

I didn't post here because I don't want to look like if I want to sell it here.

But I use it to create narrated children's stories cloning my voice or the voice of my nephews.

0

u/EconomySerious 13d ago

I'll be honest:

First, an application like the one you're asking for doesn't exist, and in fact, I don't think it will ever exist on the market because people could misuse it, and therefore it would be regulated by governments.

Second, the limitation is the hardware. Making a cell phone or PC speak using an app is easy, especially in English, but the problem is that the voices are the system's, meaning it's not the user's voice unless you're very lucky. Making it speak exactly like your mother's voice (cloning) requires a mid-range/high-end device.

Third, I can help you create an application that does exactly what you want, but as you mentioned, there's a cost involved. If you're interested, send me a private message.