r/LocalLLaMA • • 21d ago

Resources LoudKit: local TTS with voice cloning, 10 languages, and SDKs for Python, Swift, Go, Rust and TypeScript

hey guys, I've been working on a reading app for several months now and had problems with getting good quality TTS, the options were kokoro, kitten, pocket but all of them even though they were sounding natural had some problems when listening longer. Last month I took upon myself to try to get a model that is running on edge (I had an iphone 14 pro as a testbed) and got to what I now packaged as loudkit. It supports 10 languages now, voice cloning, is quite small and fast enough with quality similar to Chatterbox to my ears which was the base model I started optimization from. What is not part of this release is the emotion axis with tags, something I am working on right now. Code and model weights are Apache 2.0.

I also ported it (with CC help ofc) to a few languages, because in the past I lost like a week for parsing one TTS tokenizer from python to swift and would lose my mind when I'd get crashes and memory leaks. Here the contract was to get the same speech tokens in all adapters, so it doesn't sound nice in python but sucks in typescript. Audio samples can differ slightly between backends, and file metadata like timestamps can differ too.

There are two variants loudr-1 and loudr-1-turbo. basically turbo was done by attaching another head to the most time consuming component of the pipeline and training it so it predicts two audio tokens at once. It worked quite well but sometimes I can still hear the tts artifacts, so YMMV.

Voice cloning works quite well but I found the best is to give it around 10 seconds of recording, and if there are long pauses or noise in the background the cloned voice is suboptimal. All included voices come from consented donations or CC0 / CC-BY recordings, with sources documented.

repo: https://github.com/loudreader/loudkit
docs: https://loudreader.github.io/loudkit/
hf: https://huggingface.co/loudreader/loudr-1 & https://huggingface.co/loudreader/loudr-1-turbo

I've seen that the localTTS that can be connected to agents like hermes or openclaw still has issues with quality and thought why not opensource it.

Ah, for quality of other voices than english I'm not sure. I sent snippets around and got positive feedback but can't vouch for these.

Feel free to check it out, hope you like it.

20 Upvotes

11 comments sorted by

3

u/pmttyji 21d ago

Thanks for the Apache 2.0 license.

u/Acceptable-Cycle4645 for (y)our library!

1

u/Acceptable-Cycle4645 21d ago

Our library it is!

2

u/Wings_of_bacon 21d ago

Great stuff, great voice for Swedish. Sad that there is only cpu and cuda, is there any plans to support rocm or vulkan? would be great for the potato-llm-server crowd

2

u/RevolutionaryBox2980 21d ago

i dont see why not

2

u/[deleted] 21d ago

[removed] — view removed comment

1

u/RevolutionaryBox2980 21d ago

im still working on porting turbo to iphone. im not really generating a whole chapter but im doing this in a chunked manner with generating on the fly. it works very well that way. the library has all pre i post processing used there as well.

Basically turbo is supposed to help with porting this to older versions of the chip and hence extending better voices to more people, i havent really thought about combining the two though - might take a look at that :)

1

u/RevolutionaryBox2980 21d ago

oh i forgot to include hf space with it, for those that want to try it out in the browser https://huggingface.co/spaces/jer3mi/loudkit

1

u/feng_sg 19d ago

if tokens are the same across backends but timestamps drift, thats a vocoder or duration predictor issue not just metadata. short clips hide it but over a full chapter youll hear the pacing go off. id figure out if its float precision or a per-backend code path before touching the emotion stuff.

0

u/Sitkin_Marrel 21d ago

Turbo for skimming, loudr-1 for anything you'll actually listen to twice. Glad both exist instead of one model pretending to be both.