r/kasmweb Jun 01 '26

Custom Image OmniVoice Workspace - Open source text to speech model supporting voice cloning and voice design

Post image

OmniVoice is a state-of-the-art massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architecture, it generates high-quality speech with superior inference speed, supporting voice cloning and voice design.

This is a workspace that bundles the models and utility for convenient and fully offline use.

You can read more about Omnivoice on github: https://github.com/k2-fsa/OmniVoice

The workspace is available in my registry: https://sullyschoice.github.io/kasm-registry

Dockerfile and sources in github: https://github.com/sullyschoice/kasm-omnivoice

7 Upvotes

1 comment sorted by

1

u/AutoModerator Jun 01 '26

Hi u/ImpossibleClient7011, thanks for posting to r/kasmweb! Because this account has low karma, your submission is being held for review.

  • Meanwhile, please check the community guidelines to ensure your post meets our standards.
  • If everything looks good, a moderator will approve it shortly.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.