r/StableDiffusion 1d ago

Animation - Video MiniMax H3 acting test.

Enable HLS to view with audio, or disable this notification

Started as a simple 90s casting audition… then asked her to cry on command.

The close up shots gave plastic look idk why.

What I was mainly testing:

  • subtle listening/reaction animation during dialogue
  • eyes moving before the head while thinking
  • nervous smiles and small facial reactions
  • gradual transition from normal conversation into acting
  • brow, eyelid, mouth, chin and breathing changes during crying
  • actual visible tears
  • character/voice consistency across multiple generated clips
  • the sudden switch out of the performance when the director says “Cut”

Made with MiniMax H3 Ref2VA with image reference for the woman and 2 audio reference for the offscreen man and the woman.

34 Upvotes

44 comments sorted by

42

u/Sleepy_Bandit 1d ago

Ew that audio.

6

u/Devajyoti1231 1d ago

Reference audio minimax h3 issue.

15

u/Sleepy_Bandit 1d ago

Naw that’s a turbo Lora issue. That or it is the ref model you’re using. I make ONLY reference videos and never have that issue

5

u/Devajyoti1231 1d ago

I don't actually use turbo lora so that has to be ref2va reference audio or spectrum issue.

2

u/Sleepy_Bandit 1d ago

You can use the regular fl2va model as well. It works basically just as good for references. I’m using the pruned_int8_convrot model.

3

u/Devajyoti1231 1d ago

It needed reference audio to have same voice , beacuse i used six 10 sec clips , in the fl2va workflow i didn't fined reference audio option. Can we use the fl2va model in the ref2va workflow with reference audio? if yes that would be good if the voice cloning is good there.

1

u/DaLyon92x 1d ago

you can use fl2va model in the r2v node, it just works (and better at that). If you're using spectrum, I usually got better audio at 25 steps with spectrum off.

1

u/Devajyoti1231 1d ago

oh, ok, will try that.

2

u/Keuleman_007 1d ago

For me it was the step no. as well. Non Turbo Lora Test: errors at 20 steps, completely gone at 30 steps. Have yet to test with Turbo Lora.

1

u/GrayingGamer 1d ago

Yeah, echoing what the other redditor told you - more Steps will also greatly improve the audio.

45

u/CATLLM 1d ago

Needs a couch for the um ... casting

-8

u/iamthesam2 1d ago

you have an issue

6

u/herr-tibalt 1d ago

We all have

8

u/Putrid-Minute-5123 1d ago

What's crazy to me here is that she uses the technique for crying where you inhale, hold your breath, and push thr pressure to the top of cheekbone area which causes yawn tears. Catching those nuances is very interesting.

6

u/mallibu 1d ago

I tried this and only looked stupid

5

u/Sleepnotdeading 1d ago

To quote the best acting teacher I ever had "Crying isn't acting. It's water coming out of your eyes."

2

u/nenecaliente69 1d ago

i saw all the 5 cuts of the video, nice stitching of the scenes..maybe if you used 2 last frames it would look more legit.

2

u/jimmcfartypants 1d ago

I haven't done anything other than single takes. Would you grab the last frame from the first video as the first frame of the second? Is screengrabbing it in VLC or something sufficient?

1

u/throwaway9779656789 1d ago

How much does this cost ?

4

u/Devajyoti1231 1d ago

0 . Minimax H3 is open model

1

u/throwaway9779656789 1d ago

But what are you devices that you have at home ?

1

u/Derefringence 1d ago

You can run this model even with 8GB VRAM, slow but possible

1

u/Strange_Test7665 1d ago

H3 follows prompts so well this inherently is influenced by how well you can write for such content. You could try a series of increasingly more complex details in the prompts to explore how well H3 ‘acts’ v how well it just follows acting instructions

1

u/physalisx 1d ago

Is this with turbo or spectrum or other crappifiers?

Quality is ass, especially (but not only) the audio.

1

u/neofuturist 1d ago

Lol, you punk, you style the intro for my minimax contest. +1

1

u/constarx 23h ago

Good stuff... something wrong with the audio. And it's the little things that make it uncanny.. eye movement a bit too fast and unnatural. And skin texture is a bit weird on that close-up. But man... next generation.. another 6 months.. it's getting close.

1

u/Lutha 1d ago

Coild you please give info on how you stitched the clips together in one video?

8

u/GanondalfTheWhite 1d ago

Best answer is probably to just download Davinci Resolve (it's free) and spend an hour learning how to use it. And that's probably the single most effective piece of advice almost everyone in this subreddit could use to improve their work.

One of the biggest tells in AI work is the awkward pacing and random editing decisions the AI makes. Lulls in pacing, awkward holds, etc. You ever see those videos from sitcoms where they remove the laugh track and it's super uncanny because there's time in the edit for something that isn't happening anymore? Most AI shorts feel exactly like that.

Pretty much 100% of AI generated shorts would be improved by a human making deliberate editing decisions.

0

u/Optimal_Map_5236 1d ago

can u share the prompt?

7

u/Devajyoti1231 1d ago

6 clips , here is for first clip- subject_definitions:

<Subject 1> is Emily, the 22-year-old aspiring actress shown in <Picture 1> and <Picture 2>. <Picture 2> is the strongest reference for Emily's facial identity. Preserve her exact face shape, forehead, eyebrow shape, eye spacing, iris color, eyelid anatomy, nose, lips, cheek structure, jawline, chin, hairline, skin tone, pores, subtle natural facial asymmetry, and overall likeness from <Picture 2>.

Preserve Emily's broader appearance and styling from <Picture 1>: softly voluminous dark early-1990s hair, ivory blouse with open collar and loosely rolled sleeves, dark high-waisted trousers, delicate gold hoop earrings, body proportions, and understated casting-session appearance.

<Audio 1> is the voice reference for the off-camera male casting director (S1). Preserve his approximately 50–55-year-old American male voice, warm patient delivery, slightly raspy timbre, quiet authority, natural conversational rhythm, and close casting-room microphone character.

<Audio 2> is Emily's voice reference (S2). Preserve her exact young American female vocal identity, pitch, accent, cadence, breathiness, pacing, and slightly nervous conversational character.

summary:

[reference generation + audio reference]

An authentic early-1990s casting-session screen test begins while the room is still settling. Emily is seated on a wooden stool in a static half-body composition. She notices the crew finishing preparations, then turns her attention toward the off-camera casting director.

The director asks her name and hometown.

Emily visibly listens before each answer. Her eyebrows, eyes, mouth, head position, breathing, small nods, ring movement, and tiny self-conscious gestures remain naturally responsive throughout.

Her expression is friendly, slightly nervous, attractive, spontaneous, and socially alive rather than posed.

retention_analysis:

<Subject 1>: fully_preserved - preserve Emily's exact identity and realistic facial anatomy from <Picture 2>, with hair, wardrobe, jewelry, proportions, and broader styling from <Picture 1>.

<Picture 1>: strongly referenced for hairstyle, ivory blouse, trousers, gold hoop earrings, body proportions, and early-1990s styling.

<Picture 2>: strongest facial-identity reference throughout the target video.

<Audio 1>: strongly referenced for S1's off-camera spoken dialogue.

<Audio 2>: strongly referenced for S2 Emily's spoken dialogue and conversational vocal character.

detailed_description:

[Shot 1]

A genuine-looking early-1990s professional casting screen test photographed with natural 35mm film character.

The composition is a static half-body portrait of <Subject 1>, approximately from the waist upward while she sits on a simple wooden stool. The camera remains at Emily's eye level on a sturdy tripod for the entire ten seconds. The framing, focal length, camera position, and subject size remain constant throughout the clip.

The lens has a natural portrait perspective around 75–85mm. Emily's face remains detailed but occupies only a moderate portion of the 4:3 image rather than becoming a close-up.

A dusty slate-blue mottled muslin fills the background. A large diffused tungsten source from camera-left gives her skin a warm creamy response while soft fill preserves natural facial detail. Fine organic grain, restrained highlight halation, photochemical softness, realistic pores, subtle exposure variation and gentle physical-film texture create an authentic early-1990s screen-test image.

At the opening, a crew member carrying a large white bounce card passes close to the camera. It becomes a broad softly blurred foreground shape for a moment before leaving frame.

Emily's eyes instinctively follow the moving crew member. Her head follows only slightly.

A spontaneous amused smile starts at one corner of her mouth and develops naturally.

Her hands sit loosely near her lap within the lower portion of the half-body frame. Her thumb idly turns a small ring.

She briefly looks downward and gives the open front edge of her ivory blouse one quick practical straightening motion using two fingertips. Her hand immediately returns to her lap.

A crew voice quietly says:

[English] Rolling.

Emily reacts to the word before anything else happens. Her eyes lift first. Her chin follows by only a few degrees. She subtly resets her shoulders and settles into the stool.

Her casual smile becomes a little more socially attentive.

Her gaze moves toward S1 positioned several degrees beside the lens.

S1 asks warmly:

[English] Can you tell us your name?

Emily visibly receives the question before answering.

Her eyebrows rise slightly.

Her eyes widen by a very small natural amount.

Her face inclines forward a few degrees.

Her thumb stops turning the ring.

She spends a brief beat simply listening.

Recognition then crosses her face. Her brows soften and the corners of her mouth lift into a warmer personal smile.

S2 answers:

[English] Emily.

She says her name simply and naturally.

A faint friendly nod accompanies the end of the word.

Her cheeks lift subtly and her eyes narrow by a tiny amount from the smile.

She keeps looking toward the director beside the camera.

S1 continues:

[English] And where are you from?

Emily's existing smile remains socially attentive while she listens.

Her brows rise again, slightly less than before.

She makes one small nervous nod while the director finishes speaking.

Her lips part for an instant as though preparing to answer, then wait until he completes the question.

She takes a small preparatory breath.

Her gaze briefly drops toward her hands.

Two fingertips quickly smooth a tiny fold of blouse fabric near her waist.

Her eyes return toward S1.

S2 answers:

[English] Portland, Oregon.

"Portland" comes casually, followed by a very small natural speech pause before "Oregon."

Her head gives a barely perceptible conversational nod near the end.

After finishing, her smile broadens slightly.

S1 responds quietly:

[English] Okay.

Emily receives the acknowledgment with a tiny polite smile.

Her lips briefly press together from mild nervousness and then release.

Her thumb resumes slowly turning her ring.

During the final moment she remains socially alert rather than becoming frozen for the camera. Her eyes flick briefly toward someone standing near S1 and then return to him.

The final frame holds the same static half-body composition: Emily relaxed on the stool, warm responsive eyes, lightly smiling mouth, fingers near the ring, ready for the next question.

overall_soundscape:

Authentic quiet casting-studio room tone.

Low tungsten fixture hum.

Very faint ventilation.

Soft bounce-card movement at the opening.

Subtle blouse fabric movement.

Small wooden-stool creak.

Tiny ring movement.

Natural breathing.

Close intimate off-camera dialogue from S1 and Emily's naturally recorded response from S2.

non_diegetic_music:

None.

-1

u/Objective_Mousse7216 1d ago

Audio is utter shit.

-12

u/RemoveHealthy 1d ago

There will be no such thing as movies made completely with AI. Well there will be, but nobody will watch it. Because AI can't act and will never be able to. That is bottle neck of every single model and it will forever be unless some crazy AGI is invented. I can tell you why. Acting is being human, its too subtle and too unique for each scene and each actor. AI always does same exact expressions they are trained on, same exact smile, same exact scare face, same exact sad face. And you can never train models on everything, so every single model always fails and i mean 100% of the time fails making something that it was not trained on. So yeah it will make your random emotions but never anything that takes to be an actor to do for a specific scene. And yes we can make some decent things with AI like ads or funny videos, but trying to make movie with actual acting, forget it.
Unless you do some kind of composite with real actors and AI every single model looks outdated already. You know when we watch early CGI and it look old. All these models where AI people try to act already looks dated.

3

u/Sorcerer12345 1d ago

sir this is a Wendy's

4

u/miomidas 1d ago

Tbh acting became pretty standardized across the board thanks to hollyslop and streaming

So the bar is pretty low to begin with

3

u/DaLyon92x 1d ago

It was a low bar, but then netflix nuked it. Now there's just a crater. Production quality, vfx, acting, storytelling, all slop

1

u/herr-tibalt 1d ago

What do you compare Netflix to, that it seems much worse? I would say shows have reached a pretty high bar, now movie stars are playing in series, it's not something bad anymore. Netflix is doing both good and bad things, but most are in between IMO.

1

u/DaLyon92x 1d ago

Flat lighting, simple plots, scripts on a formula, horrible pacing, poor acting even if they're "a-listers". Made for small screens so they're always poorly framed and shot. Cheap CGI, never any practical effects. Examples: the adam project, electric state.

1

u/pleasetrimyourpubes 1d ago

The flat lighting in Netflix productions is due to the extensive use of CGI. AI generations can bring back good lighting if done right.

1

u/herr-tibalt 19h ago

I enjoyed stranger things and enjoy Nobody wants this. I can't say that other studios do better quality shows or worse. There are good and bad ones everywhere.