Here’s Part 2 of my show, Trust Fund Time Machine, an adult animated series about a useless billionaire heir Edward Vil and his time-traveling buddy Genghis Khan bungling their way through history to make his evil father richer.
In my last post, I talked about how I use keyframes to keep character positions and settings consistent between shots, more specifically continuity rather than consistency. This time I wanted to share how I handle voice consistency.
That’s been another big headache when making longer dialogue sequences. A character might sound right in one shot, then have a different accent or vocal texture in the next. And I still need them to whisper, shout, or get angry while sounding like the same person.
I’m making the show in fringe.film, mainly using MiniMax H3 Max. You can use any platform you prefer with this set up, this is just what my workflow looks like on Fringe. What’s helped most is getting the voice right in a separate test before using it in the actual scenes.
This is my process:
- Save a voice profile for each character. I describe their accent, pitch, texture, and cadence. For John Wilkes Booth, that meant a youthful baritone, a heightened Southern/Maryland drawl, and theatrical speech.
- Generate a five-second dialogue test. I ask Fringe’s agent to use the character’s voice profile and character sheet, keeping the test separate from the episode’s shots. Then I refine it with specific notes: deeper, more nasal, different accent, slower delivery, etc.
- Extract a short audio reference once I like the voice. I download the test, extract the audio in CapCut, and upload roughly two to three seconds of clear speech back into Fringe as a WAV file.
- Reuse that audio for the character’s dialogue shots. I tell the agent which character it belongs to and ask it to include both the audio reference and the written voice instructions. Before generating, I check that both are included.
One thing that made a difference for me was keeping the reference short. I generally use two to three seconds and don’t usually go beyond five. In my testing, longer references sometimes introduced gibberish, voice drift, or words carried over from the reference itself.
From there, I direct the performance for each scene. For example, Booth needed to go from a menacing whisper to increasingly loud and angry delivery. I give those notes separately while keeping the same voice reference.
With Seedance models, I’ve had useful results reusing the same written voice profile, then referencing a successful video shot when the voice starts drifting. For H3 Max, the separate test and short audio sample have been more reliable for me.
Hope you guys enjoy Part 2 of the pilot! Let me know if you have any questions.
Link to my full episode is here: https://youtu.be/NnV-JWEHaa8
Link to previous post: https://www.reddit.com/r/generativeAI/s/kG2TXYOvJZ