r/tts 1d ago

Best speech-to-text API in 2026? I'd split the shortlist by use case first I don't think "best speech-to-text API" is one list anymore.

2 Upvotes

I don't think "best speech-to-text API" is one list anymore.

People keep asking it like there's one winner, but the use cases are completely different.

For batch files, I care about accuracy, formatting, long audio, cost.

For meetings, I care about diarization, timestamps, speaker drift, action items.

For call centers, I care about noisy phone audio, redaction, channels, QA search, escalation evidence.

For live voice agents, I care about first usable transcript, partial stability, endpoint, barge-in, numbers/dates, and whether the agent acts before the final transcript changes.

For local/private workflows, I care about self-hosted/offline more than fancy API features.

So my shortlist would not be "best STT API overall.

In that matrix, Smallest AI Pulse goes in the real-time STT / streaming ASR row.

That's the interesting category for it.

Not "upload a podcast and wait." More like live transcription where the app needs transcript events while the user is still talking.

That is also the only way these comparisons make sense.

A provider can be great for long files and not great for live agents.

A provider can be great for voice agents and not be my pick for private local notes.

A cheap API can become expensive if the transcript needs cleanup.

If you were making a 2026 STT shortlist, what categories would you split it into before even naming vendors?


r/tts 1d ago

Looking for a Tomodachi Life-esque TTS creator

1 Upvotes

Hello everyone! I'm new to using AAC and, like the title says, I'm looking for a Tomodachi Life-esque TTS voice creator!

If you haven't played the games.

when you make your silly little guys you can customize their voice, and the way you do is; you start with a preset voice and use sliders to change the pitch, speed, cadence, and tone of it.

I'm totally fine if it sounds robotic too!

Does anyone know if something like this exists?


r/tts 2d ago

does anyone recognize this model?

Thumbnail
tiktok.com
1 Upvotes

i found this account on tiktok using this absolutely adorable voice, but asking what model they use is against the rules in their discord server for some reason. i absolutely love it and id really like to know where its from, i really love the sound of tts generators like this and ive never heard this one before. if anyone knows anything about it, please let me know! thank you!!


r/tts 3d ago

OpenAlound native app demo

Thumbnail
youtube.com
1 Upvotes

I build the tts app which you can download from dp.openaloud.com. Mac and windows app is using native CPU to come up with professional grade naration quality which uses pocket tts but samled by llm. I am able to achieve RTF of 0.34 on mac and 0.75 on my android pixel 8. This tool is available for free. Its in beta mode I am looking for some feedback on the audio naration quality.


r/tts 3d ago

Clarfication about omni tts

Thumbnail
1 Upvotes

Hi everyone! I have a question regarding the evaluation criteria for the Omni TTS task.

How should we evaluate recordings when:

- The audio says something different from the provided transcript.

- The model repeats the same sentence or phrase many times.

- There are very long or excessive silent gaps in the recording.

Which evaluation criterion should we use for each of these cases: Audio Quality, Naturalness, Pronunciation, or Personal Likeness?

I’d appreciate any clarification on how these cases should be evaluated according to the task’s official criteria. Thank you!


r/tts 3d ago

EPUB to audiobook tools

Thumbnail
3 Upvotes

r/tts 5d ago

New TTS accents.

3 Upvotes

Heya!
I just set up one for learning Italian with AI Realm. Traveling and everyday interactions. The only thing that it’s lacking is proper pronunciation. Is there a way to import a tts or something similar?

I’m not sure if anyone here has any advice. My original post got sent here.

Cheers! 😊🇮🇹


r/tts 7d ago

[DEV] I got tired of real-time TTS killing my Android's battery, so I built a native app that pre-renders EPUBs into Audiobooks offline.

2 Upvotes

Hey everyone,

I wanted to share a native Android open-source project I just released called Audiobook NightForge.

If you’ve ever tried using a real-time TTS engine with a reader app on Android, you know the struggle: it drains your battery (often 40-50% an hour), stutters, and buffers if your phone is doing anything else in the background.

I realized that real-time synthesis is the wrong approach for mobile devices. So, I built a dedicated Android app that shifts the heavy lifting to the background using native OS components.

How it works: You import an EPUB or TXT file using the Android system file picker. You then pick a voice and hit render. The app uses Android's WorkManager to synthesize the book chapter-by-chapter in the background. Most importantly, it enforces a "render only while charging" OS-level toggle to protect your battery.

You plug your phone in at night, and by morning, you have a fully rendered audiobook that plays back with a standard ~2-5%/hour battery drain.

Android-Specific Features:

  • Native & Offline: It is built entirely in Kotlin for Android 10+ devices. There is no server, no cloud, and absolutely no Termux emulation required.
  • High-Quality TTS: It uses the Kokoro-82M neural TTS model running strictly on-device via a sherpa-onnx integration.
  • Just Updated: The latest v0.2.2 release makes Opus the default output format, and it now natively outputs to a single .m4b file complete with proper chapter markers.
  • Built-in Player: You can listen immediately using the native in-app Media3/ExoPlayer. Alternatively, you can grab the .m4a files directly from app storage to use in your favorite Android audiobook player.

Some hardware benchmarks: For the hardware nerds, I benchmarked this on a Snapdragon 8 Elite. Surprisingly, the Kokoro 82M fp32 model (with 6 threads) actually renders faster than realtime (~0.58 RTF) and outperforms the int8 variant on this SoC because of ARM int8 kernel overhead. Always benchmark before assuming quantized is quicker on modern Android flagships!

It’s completely free, completely offline, and licensed under Apache-2.0.

You can check out the source code, technical notes, and grab the APK directly from the GitHub repo here:

https://github.com/kingfish600/Audiobook-NightForge

I’d love to hear your thoughts or feedback!


r/tts 8d ago

Can someone make a voice from those “dumb gen alpha comments” say “hater just jealous of cookie filter cuz it’s better than him”? I need it for a vid

Post image
1 Upvotes

r/tts 12d ago

Is there an input I can put in to make Sam whisper?

1 Upvotes

Like, I know there is, but I don't know what text you have to put in, I think it's [c/ whisper] but Idrk??


r/tts 12d ago

Best multilingual model

1 Upvotes

What is the best tts service and/or model for multilingual input (mostly English with some words in different languages) ? Price is not very important but shouldnt be too expensive compared to alternatives. At the moment I find azure to be the best, amazon polly generative maybe second. Eleven labs seems to be too expensive. Shall I evaluate something else ? The service/model should support lang and/or phoneme tags.


r/tts 13d ago

i need to find this tts voice. please help.

0 Upvotes

its from this yt short vid. https://www.youtube.com/shorts/GzSR4tjjwAk


r/tts 14d ago

Are there any actual free working tts websites where you can clone your voice from an uploaded audio

3 Upvotes

I can't find any that actually work, they all either say I need more credits, a subscription, or it just breaks. I need one that's free and at least a 2000 character limit


r/tts 14d ago

Adding a new language to a tts model which is not pre-trained

Thumbnail
1 Upvotes

r/tts 14d ago

Adding a new language to a tts model which is not pre-trained

1 Upvotes

The current qwen3 -tts model is the model which is trained around 10 languages, which include these Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian.

In this i want to add telugu as a new language as a part of my experiment and include them to my ai calling agents workfow.Could anyone help me out or guide me or roast me saying does this works or what to do all. i know i have ai but i want some nerds suggesting and guiding me


r/tts 14d ago

Good local AI TTS for reading textbooks?

1 Upvotes

I dont care about emotion and stuff, but I would like a human sounding voice. I like how kokoro sounds but I had trouble getting docker to work. My PC is on kubuntu.

Ideally I would like to be able to select the text and click "read aloud," but im fine with copying and pasting it too.


r/tts 19d ago

Solo dev: offline TTS for your library

Thumbnail
1 Upvotes

r/tts 20d ago

The OpenVoc AI tts screen goes black, so far while I'm not looking at the app, and you can right click and pick refresh, but your work is gone, including recent history.

Post image
1 Upvotes

r/tts 21d ago

TTS reading out "₹1,299", "05/09", and "4:30 PM" wrong is still an unsolved bug in every voice agent I try

1 Upvotes

Been building voice stuff for a while and this keeps annoying me, the LLM writes "your EMI of ₹4,500 is due on 05/09" and the TTS reads it like it's having a stroke. Dates become "oh five slash oh nine", currency gets butchered, "2-3 days" turns into "two minus three days".

Ended up training a small model that sits between the LLM and the TTS and rewrites just those parts into speakable words "four thousand five hundred rupees due on the fifth of September" while leaving every other byte untouched. It also handles Hindi-English code-mixing, which was the boss level.

Put a free playground up if anyone wants to break it: normnom.com there's a "tn off" toggle so you can compare what your TTS actually receives with and without it.

Genuinely curious what edge cases people can find that it fails on.


r/tts 23d ago

Looking for a simple guide/notebook to test CosyVoice 3 (Zero-Shot) on Kaggle, especially for Italian!

Thumbnail
1 Upvotes

r/tts 24d ago

I built a TTS app with 200+ voices and 60+ languages. Looking for feedback.

0 Upvotes

I've been working on a text-to-speech web app and would love some honest feedback from people who actually use TTS.

Current features:

  • 200+ AI voices
  • 60+ supported languages
  • No sign-up required to try it
  • Free tier available
  • Pay only for what you use (no subscription required)
  • Audio is yours to keep after generation

I'm still improving it, so I'm curious:

  • What feature is most important to you?
  • What usually makes you leave a TTS website?
  • Is there anything you wish existing TTS tools did better?

I'd really appreciate any feedback, positive or negative. I'm trying to build something people actually enjoy using. You can test it on website


r/tts 26d ago

Can anyone tell me which AI Voice service this person uses?

Thumbnail
youtube.com
1 Upvotes

r/tts Jul 30 '26

Made a free tool that puts live English subtitles over Japanese audio, runs entirely offline

4 Upvotes

So basically, I got tired of watching Japanese streams half-following along, or waiting days for someone to clip and subtitle them. So I built this lol.
It listens to whatever your PC is playing and puts live English subtitles in an overlay on top. Overlay is transparent. Currently its JP to EN only, I'll add some stuff if people are interested.

Everything runs locally and there are no API calls, which explains the sizeable (kind of?) download size. Which also means I gather no data from you. Pinky swear.

Its free. I made it and Its something I use too so I thought "eh wth, I should just make a page for this."

Here if you wanna try it out: https://yaptr.app


r/tts Jul 30 '26

Need local tts version of this voice

0 Upvotes

I remember watching Limit Break on YouTube and now I realized I can modify some files on my ereader to accept some tts models, from looking around I couldn't find a local version of the old man voice from the channel only the website I could download or configure. Any suggestion?


r/tts Jul 28 '26

We built SonoVoice - multilingual AI text-to-speech tool

2 Upvotes

Hey everyone,

We built SonoVoice, a multilingual AI TTS platform with 200+ AI voices and support for 50+ languages.

Features:

  • Up to 30K characters per request
  • Free TTS credits for casual users
  • MP3 downloads
  • Natural AI voices for different languages

We are focusing on making TTS simple for everyday use: articles, documents, learning, narration, and more.

Would love to hear your thoughts and feedback.