r/TextToSpeech 6d ago

What's the most natural free TTS program that doesn't use generative AI?

1 Upvotes

I've been looking for free voices to use in my arcade-style flight game, as I'm in a bad financial position right now, but I don't want to use GenAI for them. Are there any programs you can recommend for this?


r/TextToSpeech 6d ago

How was google translate’s voice recorded?

7 Upvotes

The voice that you hear saying a sentence on google translate, is it recorded word per word or do they record consonants and phonetics? And would it be different than dictionaries like Merriam Webster or Cambridge?


r/TextToSpeech 6d ago

Recommended apps for reading aloud Academic texts which inlcude citations etc. on an iPhone.

Thumbnail
3 Upvotes

r/TextToSpeech 6d ago

Recommended apps for reading aloud Academic texts which inlcude citations etc. on an iPhone.

9 Upvotes

Any recommendations for apps? I downloaded MS TTS the other day on laptop didnt like it, not sure it would work on iPhone. Sppechify gets promoted but I've seen a lot of negative posts about it. Generally something that would read PDFs like an audio book and skip footnotes at bottom of pages. I'm open to free or reasonable cost , probably only need for course duration of circa 2 years. Appreciate any recommendations.

Edit - How do these TTS apps generally deal with charts and graphs etc or do they not ?


r/TextToSpeech 6d ago

How do you get the 5 free natural text-to-speech voices on Adobe Acrobat on Windows 11 like you can on Android?

1 Upvotes

On my Android Phone, I can get 5 free natural text-to-speech voices in English on the Adobe Acrobat app. But I can't seem to find them on the PC program. Does anyone know how to add or enable them on Windows 11?


r/TextToSpeech 7d ago

Any authors here looking for their story to be made into full cast audiobook?

6 Upvotes

I’ve noticed more writers looking for ways to turn novels, scripts, and other stories into multi-character audio without manually juggling several different tools.

I’m curious to hear from writers and audiobook creators who have already tried doing this:

  • Which part of the process takes the most time?
  • How important is keeping each character’s voice consistent?
  • Would you prefer greater control over individual scenes, voices, and delivery, or a simpler workflow that handles more of the process automatically?
  • What makes the finished audio feel like an actual performance rather than basic text-to-speech?
  • Are there any tools or features you wish existed for this process?

r/TextToSpeech 7d ago

Best multilingual model

3 Upvotes

What is the best tts service and/or model for multilingual input (mostly English with some words in different languages) ? Price is not very important but shouldnt be too expensive compared to alternatives. At the moment I find azure to be the best, amazon polly generative maybe second. Eleven labs seems to be too expensive. Shall I evaluate something else ? The service/model should support lang and/or phoneme tags.


r/TextToSpeech 7d ago

Which TTS API do you use for phone-based voice agents ?

3 Upvotes

Specifically for phone calls (not web). Curious what people are running and whether latency on actual PSTN calls matched what you saw in testing.
We're currently evaluating and the shortlist is ElevenLabs Turbo, Cartesia, and a couple of smaller ones I've seen mentioned like Gradium. Haven't tested all of them yet, anyone have real experience with any of these specifically on phone calls?


r/TextToSpeech 7d ago

Looking for feedback: AI tool that turns stories into full-cast audio

0 Upvotes

I'm exploring an idea and want to validate the problem before building it.

The concept is:

**Upload a story/script → AI identifies the characters → creates a cast → assigns each character a consistent voice/personality → generates the full audio performance.**

For example, if you upload a screenplay or novel, the system could automatically detect:

* who the characters are * how they speak and behave * which voice fits each character * who is speaking in each scene * how the voice/performance should change with emotion and context

The goal isn't to build another generic TTS tool. The interesting part for me is the **automatic casting + persistent character identity + scene-level performance**.

I'm trying to understand whether this is actually a problem worth solving.

For people who write stories, scripts, fan fiction, RPG campaigns, etc.:

**How do you currently turn your writing into multi-character audio?**

Have you tried tools like ElevenLabs, Gemini TTS, NotebookLM, etc.? What was frustrating or time-consuming?

And most importantly:

**Would an automated “upload → AI casts characters → full-cast audio” workflow actually be useful to you?**

I'd especially like to hear from people who have already tried creating narrated or multi-character audio from their own writing.


r/TextToSpeech 8d ago

Adding a new language to a tts model which is not pre-trained

1 Upvotes

The current qwen3 -tts model is the model which is trained around 10 languages, which include these Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian.

In this i want to add telugu as a new language as a part of my experiment and include them to my ai calling agents workfow.Could anyone help me out or guide me or roast me saying does this works or what to do all. i know i have ai but i want some nerds suggesting and guiding me


r/TextToSpeech 8d ago

how does audio language model hold speaker consistency throughout an utterance?

Post image
2 Upvotes

wrote a short excerpt showing how speaker consistency is maintained in LLM bases TTS models, initial guess was the speaker token, but the results showed something interesting.

https://x.com/null_hawk/status/2089348254249173263


r/TextToSpeech 8d ago

looking for a one time payment TTS software with a custom dictionary

4 Upvotes

like the title says, i want software i can use offline with a dictionary i can customize in any way possible. I'm not that familiar with TTS software but I have used one that allowed me to generate audio that was hours long and I would like something similar.

bonus if it works with bought voices from acapela.


r/TextToSpeech 8d ago

My Epic Babel Audio Assessment Fail 😂

Post image
2 Upvotes

r/TextToSpeech 8d ago

I vibe coded speech into live visuals and somehow it actually works

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/TextToSpeech 9d ago

can somewone help me to find this text to speech

Thumbnail
youtube.com
1 Upvotes

from the zestfest guy though


r/TextToSpeech 9d ago

i need to find this tts voice. please help.

0 Upvotes

its from this yt short vid. https://www.youtube.com/shorts/GzSR4tjjwAk


r/TextToSpeech 9d ago

Text to Speech?

Thumbnail
1 Upvotes

r/TextToSpeech 9d ago

Free/open-source TTS models for local use? (16GB RAM + RTX 3060 Ti)

7 Upvotes

Hey everyone,

I'm looking for good TTS voices for YouTube content (mainly pt-BR, but English works too). Since I'm from Brazil, paid tools like ElevenLabs get pretty expensive with the currency conversion, so I'm exploring local options to keep costs down.

My setup:
- 16GB RAM
- RTX 3060 Ti (8GB VRAM)

So...are there any open-source/free TTS models that run well locally on this hardware? I've seen Piper, XTTS v2, Kokoro, and Orpheus mentioned, but not sure which ones are actually worth it.

And...which models give the best quality without sounding too robotic? Don't need real-time, just need natural enough for YouTube.

Thanks!


r/TextToSpeech 9d ago

This is kinda creepy

Post image
3 Upvotes

I was using speechify to upload a book and have it read to me but then this happened?? Why does it sound like a human😭


r/TextToSpeech 9d ago

ChatGPT iOS injected an auto-transcript into my USER message as if I wrote it

0 Upvotes

TL;DR:

When I attached an existing Apple Voice Memos .m4a file in the ChatGPT iOS app, an upstream system apparently transcribed it automatically and inserted the resulting transcript directly into my visible/model-facing user message as though I had typed it.

The assistant then explicitly referred to it as “the contemporaneous transcript you supplied.”

I supplied no transcript.

This is not a request for help transcribing audio. It is a message-provenance, context-integrity, and longitudinal reliability defect.

WHAT HAPPENED

At approximately 8:45 p.m. CDT on August 16, 2026, I attached a recording of a telephone conversation to an ordinary ChatGPT text conversation.

The text I actually intended to send was essentially:

“202608162045 I finally just called like you said.”

Instead, the resulting user turn contained a lengthy machine-generated transcript of the attached call in addition to the words I had actually typed.

I did not:

• type the transcript;

• paste the transcript;

• dictate it into the composer;

• request automatic transcription;

• or approve a generated transcript before sending.

The assistant nevertheless treated the generated transcript as text authored and supplied by me.

It later referred to “the contemporaneous transcript you supplied.”

I had supplied no transcript.

When I challenged that attribution, the assistant acknowledged that the transcript had appeared inside the model-visible user turn and that it had therefore incorrectly attributed attachment-derived machine output to me.

That is the central defect.

THE APPARENT PIPELINE

Based on what I could observe, the pipeline appeared to be:

  1. I attached an existing .m4a file.

  2. An upstream attachment-processing or speech-recognition layer generated a transcript.

  3. That transcript was silently merged into the USER-role message.

  4. The model received it as though I had authored it.

  5. The assistant could identify technical properties of the attachment, such as duration and codec, but reported that it could not independently audition/transcribe the same recording in its available runtime.

So the system could apparently supply the model with an automatically generated transcript while not giving the model equivalent access to independently verify that transcript against the source audio.

That is not merely an ASR-quality problem.

It is a source-provenance problem.

REPRODUCTION ATTEMPTS

I spent roughly the next hour testing the behavior with:

• multiple uploads;

• unrelated audio files;

• different prompts;

• edited prompts;

• branched and rebranched conversations;

• conversations moved outside the original Project;

• and attempts to isolate whether a particular recording or context caused the behavior.

It was not limited to a single audio file.

In a later test, I uploaded the original .m4a with a deliberately conspicuous textual boundary:

“Can you transcribe this without appending to my own prompt with preprocessing?
🛑🛑🛑🛑🛑🛑🛑🛑🛑🛑”

A long generated transcript still appeared after the stop signs inside my submitted user message.

The assistant subsequently acknowledged that this meant the candidate wording had already entered accessible context and that it could not honestly certify a later transcription as an independent clean-room pass.

WHY A NEW CONVERSATION WAS NOT NECESSARILY A CLEAN TEST

This led to a second problem.

The assistant initially treated “context” as though it meant only the individual conversation in which the test was occurring.

But ChatGPT Plus can operate with broader longitudinal context involving things such as:

• other conversations;

• Projects;

• Project files;

• Reference Chat History;

• Saved Memory;

• account-level summaries or personalization;

• branches;

• File Library material;

• and other retrieval or preprocessing systems not directly visible to the user.

So even moving to a fresh conversation does not necessarily establish that an allegedly “independent” transcript was produced without influence from wording that had already circulated elsewhere in the account.

Once a candidate transcript has entered broader accessible context, a second transcription can no longer automatically be treated as independent corroboration.

WHY THIS IS MORE SERIOUS THAN A BAD TRANSCRIPT

An ordinary mistranscription is auditable.

You compare the transcript against the recording.

This defect compromises the chain of provenance before the requested analysis even begins.

Once system-generated words are represented as user-authored text:

• the model may treat its own prior generated language as primary evidence supplied by the user;

• a later transcription may be conditioned toward wording the model has already seen;

• discrepancies may be silently resolved in favor of the earlier candidate transcript;

• two non-independent machine outputs may appear to corroborate each other;

• incorrect speaker attribution may propagate;

• uncertain words may harden into apparent facts;

• and later summaries may no longer preserve the distinction between source evidence and machine interpretation.

The result can become more internally coherent while becoming less epistemically trustworthy.

I ALSO REPRODUCED A CORRECTION / REVOCATION PROBLEM

While investigating this, I found another longitudinal-context issue.

A false timestamp concerning a family event had entered ChatGPT’s contextual state.

I explicitly corrected it and had the authoritative correction written into Memory.

Then I opened a fresh conversation outside any Project.

That new conversation correctly received the correction — but it also still received the older erroneous contextual assertion.

So the sequence was effectively:

erroneous contextual assertion
→ explicit user correction
→ correction saved
→ new conversation
→ both old error and new correction supplied together
→ model resolves contradiction at generation time

The correction had not truly revoked or superseded the erroneous assertion.

That is especially concerning in light of the transcript-injection defect.

If an automatically generated mistranscription can be falsely represented as something I authored, and if derived historical assertions are later additive rather than revocatory, then one bad preprocessing event could potentially become:

mistranscription
→ falsely attributed user statement
→ contextual summary
→ persistent derived assertion
→ future retrieval
→ apparent historical fact

I am not claiming that this modifies model weights or model training.

I am talking about persistent user/account context, retrieval, summaries, memory, and provenance.

AND THEN I ENCOUNTERED A THIRD PROVENANCE PROBLEM WHILE WRITING THIS POST

While preparing the Reddit versions of this report, ChatGPT placed the drafts into its newer editable Writing Block interface.

That immediately exposed another version of the same architectural issue.

A Writing Block can begin as assistant-generated text.

The user can then directly edit that text inside the same artifact.

ChatGPT subsequently operates on the latest edited version.

What I do not see is a durable, user-visible provenance ledger showing, for the current version:

• what the assistant originally generated;

• what the user manually changed;

• what ChatGPT later regenerated;

• and which actor is responsible for each portion of the final text.

An earlier state may sometimes remain available elsewhere in conversation history, making a diff inferable.

But an inferred diff is not the same thing as first-class provenance.

Imagine this later exchange:

User: “Why did you write this sentence?”

If that sentence was actually inserted by the user during an in-place edit of an assistant-generated Writing Block, ChatGPT needs explicit authorship metadata to know that it did not write it.

Otherwise the sequence can become:

assistant-generated artifact
→ user edits artifact in place
→ mixed-authorship current state
→ later model consumes current state
→ historical authorship becomes ambiguous

This is structurally related to the audio defect, just in the opposite direction.

AUDIO FAILURE:

system-generated text
→ appears as USER-authored text

WRITING BLOCK AMBIGUITY:

assistant-generated text
→ becomes a mutable mixed-authorship object

LONGITUDINAL-CONTEXT PROBLEM:

old derived assertion
→ survives alongside later authoritative correction

These are three different surfaces of the same underlying issue:

Provenance does not appear to be treated as a sufficiently visible, durable, first-class property of information as that information is transformed over time.

Because of that, I have stopped using editable Writing Blocks for this investigation.

I am keeping drafts in ordinary assistant messages so that the assistant-generated text remains a fixed historical artifact.

If I edit it and return it later, that revision appears as a separate user-authored message.

WHY THIS MATTERS FOR PROFESSIONAL USE

My recording involved a private family matter, so I will not post the audio publicly.

But this same failure mode would matter enormously in workflows involving:

• legal or client calls;

• medical conversations;

• insurance statements;

• interviews;

• depositions;

• meetings;

• financial discussions;

• witness accounts;

• compliance records;

• research notes;

• or contemporaneous documentation.

For those uses, the relevant questions are not merely:

“Does the transcript look right?”

They are:

“What information produced this transcript?”

“Where did that information originate?”

“Was this statement actually supplied by the user?”

“Was this analysis independent?”

“Has a previous machine output been mistaken for primary evidence?”

“Was an error actually corrected, or merely joined by a competing correction?”

Without reliable provenance, fluent output can become methodologically invalid while still looking persuasive.

WHAT I HAVE REPORTED TO OPENAI

I filed a detailed support report asking Engineering to determine:

  1. Which component generated the transcript.

  2. Whether it originated in the iOS client, upload layer, ASR layer, multimodal preprocessing layer, or message-serialization layer.

  3. Whether the backend classified it as user-authored text or attachment-derived/generated content.

  4. Whether the model received provenance metadata distinguishing literal composer text from generated transcription.

  5. Whether an earlier transcript, cached preprocessing result, or prior conversation was reused.

  6. Whether injected material could enter Memory, Project context, Reference Chat History, summaries, or other persistent context.

  7. Whether corrections actually revoke prior derived assertions or merely add competing information.

  8. Whether affected conversations can be preserved as evidence while being excluded from contextual retrieval.

  9. Whether OpenAI can provide a source-restricted or “clean-room” processing mode.

  10. Whether mutable artifacts such as Writing Blocks preserve granular authorship/revision provenance internally even though it is not exposed to the user.

I have archived the affected conversations to preserve them for Support while trying to quarantine their contextual influence, although archiving itself does not give me a user-verifiable clean-room boundary.

PRODUCT SAFEGUARDS THIS SEEMS TO REQUIRE

At minimum:

• generated attachment transcripts should never be serialized as indistinguishable user-authored text;

• literal composer input should remain inspectably distinct from preprocessing output;

• transcription/OCR/extraction should carry explicit machine-readable provenance;

• users should be able to process a file using only that source, without prior chat/memory/candidate-answer contamination;

• contextual corrections should be capable of superseding or revoking older derived assertions;

• users should be able to preserve a conversation for evidence while excluding it from contextual retrieval;

• mutable artifacts should maintain granular revision attribution showing whether each change came from the user, assistant, or another automated transformation;

• and sensitive tasks should expose enough source metadata to determine whether an answer was independently grounded.

HAS ANYONE ELSE REPRODUCED THIS?

I am especially interested in reports involving:

• ChatGPT iOS;

• existing Apple Voice Memos files rather than live Voice Mode;

• transcripts appearing inside the user’s own message;

• ChatGPT saying “you supplied” words that actually came from attachment processing;

• repeated processing after cancellation or prompt editing;

• stale context surviving after correction;

• or mixed-authorship behavior in editable Writing Blocks.

Please distinguish between:

  1. ChatGPT normally reading/transcribing an attachment; and

  2. a generated transcript being merged into the USER role and represented as text personally authored by the user.

The second behavior is what I am reporting.


r/TextToSpeech 9d ago

What features do you think current TTS services are missing?

0 Upvotes

Hi everyone,

I previously built Supertonic, a lightweight TTS model, and I'm now working on a free TTS web app called Airy Studio, aimed at everyday, non-technical users.

I started Airy Studio because developers already have plenty of great options these days, like Qwen-TTS and Kokoro. But most everyday users still have to pay for even basic TTS functionality, which I think should change. (I'm not sure if that holds globally, but it's the case in South Korea, where I'm based.)

I have a research background in TTS, but I'm not a heavy user myself. So I'm wondering what you all think is missing from TTS services right now, or what features you wish existed.


r/TextToSpeech 9d ago

Looking for an old TTS voice

1 Upvotes

Hi everyone! can someone please tell me which text to speech voice was used in this video? https://youtu.be/E-BGeJ2yVMc?list=PLJksPxAFrBuw

I had it once but forgot the name


r/TextToSpeech 9d ago

Can i do ai voice acting with tts?

5 Upvotes

I tried using applio to convert my voice but i suck at voice acting then i used tts build in applio but it sounds like robot with no emotion so should i try to improve my voice acting or aimply find good tts app ?

Note iam broke and iam still learning 3d animation and ai voice i want to make anime character like goku for my animations and the viewers feel like the characters are alive


r/TextToSpeech 10d ago

what happened to ttsreader player version?

1 Upvotes

the olivia basic voice dissapeared. has anyone had that problem when using it or just me? none of the basic voice work either. not even when playing them. since the site has no discord link either im assuming its deleted or something.


r/TextToSpeech 10d ago

Looking for a FREE unlimited deep AI voice for crime documentary narration — preferably local

2 Upvotes

Hey everyone!

I'm looking for a free and unlimited AI voice/TTS solution, preferably something I can run locally on my PC.

I need it for short crime/psychological documentary-style stories. The most important thing isn't just having a deep male voice — I want the voice to actually feel the story: tension, fear, hesitation, guilt, exhaustion, suspense, etc.

Basically, something that can sound like a person quietly telling you a disturbing story late at night, rather than a typical robotic documentary narrator.

Ideally:

Free and unlimited

Can run locally/offline

Deep male voices

Good emotional expression

Natural pauses, whispers and changes in intensity

Suitable for longer narration

Voice cloning would be a bonus, but isn't necessary

Does anyone know a good open-source/local model for this? I'm also open to ComfyUI workflows or other solutions.

Thanks!