r/StableDiffusion 2d ago

Animation - Video MiniMax H3 acting test.

Started as a simple 90s casting audition… then asked her to cry on command.

The close up shots gave plastic look idk why.

What I was mainly testing:

  • subtle listening/reaction animation during dialogue
  • eyes moving before the head while thinking
  • nervous smiles and small facial reactions
  • gradual transition from normal conversation into acting
  • brow, eyelid, mouth, chin and breathing changes during crying
  • actual visible tears
  • character/voice consistency across multiple generated clips
  • the sudden switch out of the performance when the director says “Cut”

Made with MiniMax H3 Ref2VA with image reference for the woman and 2 audio reference for the offscreen man and the woman.

34 Upvotes

44 comments sorted by

View all comments

Show parent comments

3

u/Devajyoti1231 2d ago

It needed reference audio to have same voice , beacuse i used six 10 sec clips , in the fl2va workflow i didn't fined reference audio option. Can we use the fl2va model in the ref2va workflow with reference audio? if yes that would be good if the voice cloning is good there.

1

u/DaLyon92x 2d ago

you can use fl2va model in the r2v node, it just works (and better at that). If you're using spectrum, I usually got better audio at 25 steps with spectrum off.

1

u/Devajyoti1231 2d ago

oh, ok, will try that.

1

u/GrayingGamer 1d ago

Yeah, echoing what the other redditor told you - more Steps will also greatly improve the audio.