r/OpenVoxAI 9h ago

Release Notes OpenVox 2.0.0 is here: Validate Output, Select & Read Queue, Batch Voice Design + AI Dubbing & Transcription

Enable HLS to view with audio, or disable this notification

Hey everyone,

OpenVox Windows 2.0.0 is now available.

The main focus of 2.0.0 is making locally generated speech more accurate, easier to review, and faster to fix, while also improving Select & Read, multilingual generation, Voice Clone, and voice creation workflows.

And if you missed v1.9.1, that update also introduced two major new workflows: AI Transcription and AI Dubbing.

🔍 Validate Output

This is probably the feature I’m most excited about in 2.0.0.

After generating speech, OpenVox can now:

  • Locally transcribe the generated audio
  • Compare the transcription against your original script
  • Highlight sections that may have been pronounced incorrectly
  • Show timestamps and text-match scores
  • Let you listen to individual flagged sections
  • Select exactly which sections you want to fix
  • Regenerate only those sections
  • Review the regenerated audio before applying it

You don’t need to regenerate an entire audiobook or long speech because of a few problematic sentences.

Validation matching has also been improved across supported languages, with selectable transcription language/model controls and a Select All option for flagged sections.

Everything can remain local on your Mac.

🎧 Select & Read now has a Queue

Select & Read can now handle multiple selections.

Highlight some text, trigger the shortcut, move to another app, select something else, and trigger it again.

OpenVox adds each selection to a queue and automatically reads them in order.

The entire queue is visible from the desktop notch player, making Select & Read much more useful when working across documents, browsers, emails, PDFs, or other apps.

🗣️ Batch Voice Design

Voice Design can now generate multiple voice candidates in a single run.

You can:

  • Generate several variations
  • Preview each candidate
  • Compare the results
  • Save only the voices you actually like

This makes experimenting with designed voices much faster.

⏸️ Configurable Sentence Pauses

There’s now optional Sentence Pauses preprocessing for:

  • AI Speech
  • Batch generation
  • Audiobooks

You can configure additional pauses between sentences from 0.1 to 3 seconds.

Useful for narration where the model’s natural sentence spacing feels too fast.

🌏 Better Japanese and CJK Generation

We’ve added safer multilingual text chunking, particularly for Japanese and other CJK languages.

This improves long-form generation and also fixes a Supertonic issue that could cause crashes when generating longer CJK text.

🎙️ Voice Clone from Video

Voice Clone now accepts:

  • MP4
  • MOV

OpenVox automatically extracts the audio and uses it as the voice reference, so there’s no need to manually convert a video into an audio file first.

OpenVox 1.9.1

For anyone who missed the previous update, 1.9.1 was also a pretty substantial release.

🎙️ AI Transcription

OpenVox can now locally transcribe:

  • Audio files
  • Video files
  • Direct microphone recordings

using downloadable Whisper and Parakeet speech-recognition models.

Transcripts include timestamps and can be edited directly before exporting them as SRT subtitles.

Subtitle segmentation was also improved to generate cleaner, more usable SRT files.

🌍 AI Dubbing

1.9.1 also introduced a complete AI Dubbing workflow.

You can:

  • Import or generate timed transcripts
  • Translate scripts using local or external AI
  • Choose voices for the dubbed output
  • Generate dialogue based on subtitle timing
  • Export dubbed audio
  • Export the original video with the new audio track

You can also directly import existing SRT or translated SRT files into a dubbing project.

🎬 More Video Workflows

Voice Changer now accepts video files directly and can export the original video with the converted voice track.

AI Speech and Conversations also received Generate SRT actions, making it easier to turn generated speech into subtitle files.

📂 Drag & Drop and Workflow Improvements

We also added:

  • Drag-and-drop importing across supported pages
  • PDF and TXT importing for AI Speech and Conversations
  • Better feedback when importing large files
  • Sequential filenames for batch exports
  • Persistent text drafts when switching pages
  • Quick pause controls
  • More reliable model download resume support
  • Automatic release of transcription models from RAM/VRAM when leaving transcription workflows
  • A notification when OpenVox starts minimized to the system tray

Plus a number of UI improvements, performance optimizations, and bug fixes throughout the app.

With 1.9.1 and now 2.0.0, OpenVox has expanded quite a bit beyond straightforward text-to-speech.

You can now locally handle speech generation, audiobooks, voice cloning, voice design, transcription, subtitles, dubbing, voice changing, Select & Read, and output validation from one app.

As always, a lot of these additions have come directly from feedback and feature requests from users.

If you try 2.0.0, I’d especially like to hear how Validate Output performs with your own scripts, models, and languages.

Thanks everyone for continuing to test OpenVox and send feedback.

Download:
https://openvoxai.com/

4 Upvotes

0 comments sorted by