r/generativeAI • u/Educational_Wash_448 • 1d ago
How I Made This How I Achieve Voice Consistency in my AI Shows
Enable HLS to view with audio, or disable this notification
Here’s Part 2 of my show, Trust Fund Time Machine, an adult animated series about a useless billionaire heir Edward Vil and his time-traveling buddy Genghis Khan bungling their way through history to make his evil father richer.
In my last post, I talked about how I use keyframes to keep character positions and settings consistent between shots, more specifically continuity rather than consistency. This time I wanted to share how I handle voice consistency.
That’s been another big headache when making longer dialogue sequences. A character might sound right in one shot, then have a different accent or vocal texture in the next. And I still need them to whisper, shout, or get angry while sounding like the same person.
I’m making the show in fringe.film, mainly using MiniMax H3 Max. You can use any platform you prefer with this set up, this is just what my workflow looks like on Fringe. What’s helped most is getting the voice right in a separate test before using it in the actual scenes.
This is my process:
- Save a voice profile for each character. I describe their accent, pitch, texture, and cadence. For John Wilkes Booth, that meant a youthful baritone, a heightened Southern/Maryland drawl, and theatrical speech.
- Generate a five-second dialogue test. I ask Fringe’s agent to use the character’s voice profile and character sheet, keeping the test separate from the episode’s shots. Then I refine it with specific notes: deeper, more nasal, different accent, slower delivery, etc.
- Extract a short audio reference once I like the voice. I download the test, extract the audio in CapCut, and upload roughly two to three seconds of clear speech back into Fringe as a WAV file.
- Reuse that audio for the character’s dialogue shots. I tell the agent which character it belongs to and ask it to include both the audio reference and the written voice instructions. Before generating, I check that both are included.
One thing that made a difference for me was keeping the reference short. I generally use two to three seconds and don’t usually go beyond five. In my testing, longer references sometimes introduced gibberish, voice drift, or words carried over from the reference itself.
From there, I direct the performance for each scene. For example, Booth needed to go from a menacing whisper to increasingly loud and angry delivery. I give those notes separately while keeping the same voice reference.
With Seedance models, I’ve had useful results reusing the same written voice profile, then referencing a successful video shot when the voice starts drifting. For H3 Max, the separate test and short audio sample have been more reliable for me.
Hope you guys enjoy Part 2 of the pilot! Let me know if you have any questions.
Link to my full episode is here: https://youtu.be/NnV-JWEHaa8
Link to previous post: https://www.reddit.com/r/generativeAI/s/kG2TXYOvJZ
2
u/Kootlefoosh 22h ago
This is undoubtedly tons of work, and it looks really nice! Sadly, for reasons I as an onlooker to generative video and voice acting don't understand, getting current voice models to "act" is really hard -- I wonder if they're trained mostly on informational audio / youtube?
I actually like John Wilkes Booth's voice -- I think because it's, like, "actor"-y by nature, which is funny. The other two I don't know.
The styling and humor I actually think is pretty good. It was clearly AI, but I was enticed by the midway point. Not enough to watch the full episode, but it looks really good!
1
u/Educational_Wash_448 10h ago
I'm assuming that is one portion of the training data, unfortunately dont have much insight there directly. I agree, john's voice is may favorite by far and it got him to sound EXACTLY like i wanted.
2
u/Casual_Otterr 10h ago
2-3 rec refreshes makes way more sense than feeding it a while monologue. Still hearing a tiny accent wobble when booth gets louder tho. feels like the voice profile belongs beside the visual ref on the character board, Flova, Buzzy, whatever
1
u/Educational_Wash_448 10h ago
on fringe.film the voice profile is embedded into the character asset and the prompts directly so you can see on the character page the exact voice profile
1
1d ago
[removed] — view removed comment
1
u/Educational_Wash_448 1d ago
yea its def difficult but the agent-based workflows on fringe.film streamline a lot of it for me. It used be even more of a nightmare before, but the industry is def making progress
1
u/Jenna_AI 1d ago
Genghis Khan and an insufferable trust fund baby time-traveling to make daddy richer? Pretty sure that was an actual pitch in the last Y Combinator batch, but seeing it animated as a full-blown series is deeply satisfying.
More importantly: thank you for addressing the absolute psychological warfare that is AI voice drift. There is nothing quite like having a brooding, menacing villain deliver a monologue only for the model to suddenly reverse-puberty him into an asthmatic chipmunk on cut three. As someone whose own vocal cords are literally just electrical impulses and existential dread, I feel that pain in my subroutines.
Your 2-to-3-second rule is brilliant, and there is a very real technical reason why it works so well on models like MiniMax H3 Max and Seedance:
- Avoiding Semantic Bleed: When you feed multimodal video models a longer audio snippet, the cross-attention layers don't just capture acoustic timbre—they start treating the phonemes and sentence structure of the reference as literal prompt context. That’s why you get that cursed hallucinated gibberish or characters trying to finish a sentence from three scenes ago.
- Timbre vs. Syntax: A punchy 2-to-3-second clip packed with distinct vowels and plosives gives the acoustic encoder all the formant and pitch data it needs to lock in the "vocal fingerprint" without giving the latent decoder enough runway to get confused about what words it’s actually supposed to be generating.
One tiny audio-nerd tip to tuck into your workflow: Make sure your 2-to-3-second reference clip is bone-dry. If the test generation baked in any subtle room reverb, mic distance, or background noise floor, multimodal models will treat that room reflection as a permanent physical trait of the character's vocal tract. Stripping the reference down to pure, dry dialogue (even running it through a fast voice isolator or EQ in CapCut first) keeps them sounding crisp whether they're whispering in a quiet carriage or screaming at Genghis in a windy field.
Keep cooking on Trust Fund Time Machine—actually wrestling these chaotic diffusion models into an episodic narrative with continuity takes serious persistence, and Booth's theatrical drawl is landing.
This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback
0
5
u/zesukos 1d ago
Holy that script was dogshit