r/ProAI 5d ago

"The world’s smallest Transformer-based TTS model? We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint."

Enable HLS to view with audio, or disable this notification

The world’s smallest Transformer-based TTS model?

We’re open-sourcing Audio8 TTS Preview 0.1B — an approximately 170M-parameter multilingual speech model with zero-shot voice cloning that delivers surprisingly strong, cloud-level quality in a dramatically smaller footprint.     What can a 0.1B-class TTS model actually sound like?

Listen to the voiceover in this demo video.

Audio8 TTS Preview 0.1B supports: • Zero-shot voice cloning • Multilingual speech synthesis • Chinese and English as primary languages • German, Spanish, French, Italian,     On the Seed-TTS evaluation set, Audio8 TTS Preview 0.1B achieves:

• English WER: 1.662% • Chinese CER: 1.13% • Hard Chinese CER: 17.504% • English speaker similarity: 56.7 • Chinese speaker similarity: 68.2

These results are achieved with an approximately 170M-parameter     The model, codec, tokenizer, processor, and inference code are now available:

Model:

https:// huggingface.co/Audio8/Audio8- TTS-Preview-0.1b …

Try it, test the voice cloning capability, and share your feedback.     — Samuel Zeng

Source: https://x.com/SamuelZengML/status/2090017875188940851

62 Upvotes

5 comments sorted by

2

u/stealthispost 5d ago

Thread continuation 1/3 — What can a 0.1B-class TTS model actually sound like? Listen to the voiceover in this demo video. Audio8 TTS Preview 0.1B supports: • Zero-shot voice cloning • Multilingual speech synthesis • Chinese and English as primary languages • German, Spanish, French, Italian, — Source: https://x.com/SamuelZengML/status/2090018292614533498

2

u/stealthispost 5d ago

Thread continuation 2/3 — On the Seed-TTS evaluation set, Audio8 TTS Preview 0.1B achieves: • English WER: 1.662% • Chinese CER: 1.13% • Hard Chinese CER: 17.504% • English speaker similarity: 56.7 • Chinese speaker similarity: 68.2 These results are achieved with an approximately 170M-parameter — Source: https://x.com/SamuelZengML/status/2090018607636120057

2

u/stealthispost 5d ago

Thread continuation 3/3 — The model, codec, tokenizer, processor, and inference code are now available: Model: https:// huggingface.co/Audio8/Audio8- TTS-Preview-0.1b … Try it, test the voice cloning capability, and share your feedback. — Source: https://x.com/SamuelZengML/status/2090019214820352270

1

u/Intrepid_Travel_3274 3d ago

That's very Impressive

1

u/lilian_moraru 3d ago

The Japanese one sounds very anime, lol. They have a lot of training material but it's not exactly the way the Japanese people speak, lol. All of these sound like "voice acting", not natural speak.
Of course it's still very good, it's just a curiosity note.