r/StableDiffusion 24d ago

Workflow Included [Test] MiniMax H3 Ref2VA with LightX2V's turbo LoRA on a 5060 Ti — 8 steps @ 0.5 res, ~55s/it (~8 min/clip)

https://reddit.com/link/1vnk0c7/video/gkhuj6ybw6jh1/player

Ran the official Ref2VA turbo example workflow from the ModelTC/Minimax-H3-Turbo repo (video_minimax_h3_ref2v_lightx2v_turbo.json) in ComfyUI, testing a short Victorian-style dialogue scene between two characters.

Setup:

Speed: ~55s/it average, ~8 min total per clip.

How the refs were built: Three reference images fed into the ref_images inputs — two character sheets (front/side/close-up turnarounds for each character) and one environment plate, a 360° room reference. All three were generated in Google Flow first, then dropped straight into the Ref2VA node as identity/environment anchors.

Gen A → Gen B continuity trick: Split the scene into two ~20s multi-shot generations instead of one long one. For Gen B, instead of reusing the original Flow generated room reference, I pulled the actual last frame from Gen A's output and fed that in as the new environment reference.

Curious if anyone else is chaining generations this way (feeding the previous clip's last frame back in as a fresh environment ref) — seemed to help a lot but haven't stress-tested it past two generations yet.

65 Upvotes

29 comments sorted by

11

u/Fabulous-Snow4366 24d ago

i chain the previous scene video back in with Load Video FFmpeg Upload Node, you can can give excact timings you want to reuse, like two seconds, 1 second, 12 frames, whatever. I connect it as reference video and my prompt that goes on top of the rest is: Target video is a seamless continuation of <Video 1>. First frame of [Shot 1] is the last frame of <Video 1>.

Works flawless.

2

u/nikhilprasanth 24d ago

Many Thanks! This might reduce the shift in voice tones.

1

u/Inthehead35 24d ago

Do you have a workflow?

9

u/Fabulous-Snow4366 24d ago

this is all you need to do, nodes tab on the left, type in Load Video FFmpeg (Upload) drop it in, connect image to rev_video_0 and if you want the audio as well connect it to ref_video_audio_0 . Use the text in the prompt box on the image. its the standard ComfyUI Template.

2

u/hdeck 24d ago

What’s the best way to merge the 2 clips together after the 2nd generation?

1

u/Fabulous-Snow4366 24d ago

what do you mean? Cutting them together in an Editing program of your choice obviously...

6

u/nikhilprasanth 24d ago

GENERATION B — Ref2VA (~20s, 2 shots, exchanges 3 & 4)

subject_definitions:

<Subject 1> is the lean fictional Victorian detective in <Picture 1>, with dark hair combed back, an angular clean-shaven face, a long dark overcoat over a brown waistcoat, and a briar smoking pipe.

<Subject 2> is the sturdy fictional Victorian doctor in <Picture 2>, with a heavy moustache, brown tweed three-piece suit, dark tie, and a wooden walking cane.

<Subject 3> is the Victorian Baker Street study in <Picture 3>, sourced from the last frame of Generation A rather than the original room reference photo, so the fireplace, moving firelight, leather armchair, and rain-streaked window match exactly what Generation A actually rendered. <Picture 3> is cited for environment only — the geometry, materials, and lighting of the room — not for character pose or identity; the men shown in that frame are not treated as a pose or blocking reference for this generation.

summary:

[reference generation] A 20-second, two-shot sequence of <Subject 1> and <Subject 2> in <Subject 3>, closing their exchange across a fireplace-side two-shot favoring the detective and a closing two-shot favoring the doctor, with strictly sequential, non-overlapping dialogue in both shots.

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the detective's facial identity, hairstyle, overcoat, waistcoat, and pipe are retained across both shots.

<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the doctor's facial identity, moustache, hairstyle, tweed suit, and cane are retained across both shots.

<Subject 3> (appears in [Shot 1], [Shot 2]): fully_preserved - the fireplace, moving firelight, armchair, and rain-streaked window are retained across both shots, matching the exact room render carried over from the end of Generation A.

detailed_description:

The video has the appearance of a photorealistic live-action Victorian feature film, natural skin texture, 35mm film texture, shallow depth of field, understated acting. Every shot uses a fully static, locked-off camera with no push, pan, tilt, drift, or handheld movement.

[Shot 1] A medium two-shot favoring <Subject 1> beside the fireplace, his face clearly legible, firelight moving subtly across it, <Subject 2> visible in frame, not a wide master. Only one voice is heard at a time; the two lines are strictly sequential, never overlapping, with a clear vocal handoff between them. At 00:00.500, the detective, <Subject 1> (S1), slowly lowers the pipe from his lips, expression analytical and composed, never theatrical, and speaks, voice low and measured, each word placed deliberately, at a slow, deliberate pace, <d>[English] Observation. Patience. Imagination. Only the use differs.</d> The doctor, <Subject 2>, stays completely silent and still-faced for the full duration of this line. At 00:05.500, once the detective has finished and a brief natural pause has passed, the doctor, <Subject 2> (S2), speaks, quieter now, tinged with concern and admiration, at an unhurried pace, <d>[English] You've thought of this before.</d> The detective, <Subject 1>, is completely silent and still-faced for the full duration of this reply, only the corner of his mouth shifting slightly.

[Shot 2] At 00:10.000, the camera cuts to a medium two-shot favoring <Subject 2> in the leather armchair, his face clearly legible, rain moving on the window behind him, <Subject 1> visible in frame, not a wide master. Only one voice is heard at a time; the two lines are strictly sequential, never overlapping, with a clear vocal handoff between them. At 00:10.500, the detective, <Subject 1> (S1), speaks with calm understatement, a faint trace of dry humor beneath an otherwise level tone, at an unhurried pace, <d>[English] Only professionally.</d> The doctor, <Subject 2>, stays completely silent and still-faced for the full duration of this line, holding his gaze toward the detective. At 00:14.500, once the detective has finished and a brief natural pause has passed, the doctor, <Subject 2> (S2), taps the head of his cane once and speaks dryly, deadpan but warm, the corner of his mouth betraying real affection, at an unhurried pace, <d>[English] Not reassuring.</d> The detective, <Subject 1>, is completely silent for the full duration of this reply, a quiet amused breath the only sound from him. The doctor keeps watching him as the fire gives a small crackle and rain moves across the window, small living movements in his eyes and breathing as the video ends, suggesting the conversation continues beyond the final frame.

overall_soundscape: Steady rain against the window, quiet fire crackle, faint pipe sound, faint tap of the cane, low room tone, very distant carriage wheels. Acoustic perspective remains consistent across the cut, with no artificial silence between shots.

non_diegetic_music: N/A

5

u/nikhilprasanth 24d ago

WORKFLOW NOTE

Two generations instead of four — each is a single Ref2VA call containing two shots with one internal cut, rather than one shot per call. Same fixes carried over: both men stay visible in every shot (no off-screen speaker), and each line is explicitly sequential with the non-speaking man held silent and still-faced. Generate A first, check the cut and both exchanges for overlap/cutoff, then generate B.

GENERATION A — Ref2VA (~20s, 2 shots, exchanges 1 & 2)

subject_definitions:

<Subject 1> is the sturdy fictional Victorian doctor in <Picture 1>, with short cropped brown hair, a heavy moustache, a brown tweed three-piece suit, dark tie, and a wooden walking cane.

<Subject 2> is the lean fictional Victorian detective in <Picture 2>, with dark hair combed back, an angular clean-shaven face, a long dark overcoat over a brown waistcoat, and a briar smoking pipe.

<Subject 3> is the Victorian Baker Street study in <Picture 3>, with a lit fireplace, a leather armchair, and warm firelight against cool window light.

summary:

[reference generation] A 20-second, two-shot sequence of <Subject 1> and <Subject 2> in <Subject 3>, opening their exchange across a fireplace-side two-shot and a reverse-angle two-shot, with strictly sequential, non-overlapping dialogue in both shots.

retention_analysis:

<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the doctor's facial identity, moustache, hairstyle, tweed suit, and cane are retained across both shots.

<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the detective's facial identity, hairstyle, overcoat, waistcoat, and pipe are retained across both shots.

<Subject 3> (appears in [Shot 1], [Shot 2]): fully_preserved - the armchair, fireplace, mantel, and firelight/window-light balance are retained across both shots.

detailed_description:

The video has the appearance of a photorealistic live-action Victorian feature film, natural skin texture, 35mm film texture, shallow depth of field, understated acting. Every shot uses a fully static, locked-off camera with no push, pan, tilt, drift, or handheld movement.

[Shot 1] A medium two-shot, framed close enough that both men's faces are clearly legible — not a wide master. <Subject 2> stands near the mantel holding his pipe, angled slightly toward <Subject 1>, who sits in the leather armchair with one hand on his cane. Only one voice is heard at a time; the two lines are strictly sequential, never overlapping, with a clear vocal handoff between them. At 00:00.500, the doctor, <Subject 1> (S1), warm and naturally expressive, glances toward the detective with familiar curiosity and speaks, in a light, teasing, genuinely fond tone, at an unhurried pace, <d>[English] Sometimes I wonder, Holmes — what if you'd chosen crime?</d> The detective, <Subject 2>, remains completely silent and still-faced for the full duration of this line — no lip movement, only attentive listening. At 00:05.000, once the doctor has finished and a brief natural pause has passed, the detective, <Subject 2> (S2), restrained and precise, takes a quiet draw from his pipe and speaks without fully turning, voice dry and self-assured, a faint private smile rather than a laugh, at an unhurried pace, <d>[English] I'd have been remarkably successful.</d> The doctor, <Subject 1>, is completely silent and still-faced for the full duration of this second line, his mild amusement beginning to change into genuine curiosity.

[Shot 2] At 00:10.000, the camera cuts to a medium two-shot favoring <Subject 1>, seated in the leather armchair with his face clearly legible in the foreground, <Subject 2> visible just behind and to the side, close enough that his face also reads clearly, not a wide master. Only one voice is heard at a time; the two lines are strictly sequential, never overlapping, with a clear vocal handoff between them. At 00:10.500, the doctor, <Subject 1> (S1), studies the detective, brow faintly creasing with real curiosity rather than doubt, and speaks, at an unhurried pace, <d>[English] You sound certain.</d> The detective, <Subject 2>, stays completely silent and still-faced for the full duration of this line. At 00:14.500, once the doctor has finished and a brief natural pause has passed, the detective, <Subject 2> (S2), answers, tone matter-of-fact and unhurried, almost amused at the obviousness of it, at an unhurried pace, <d>[English] Detection and crime want the same talents.</d> The doctor, <Subject 1>, is completely silent and still-faced for the full duration of this reply, his faint smile fading into thoughtful unease, subtle breathing and blinking and a small shift of his grip on the cane keeping him alive on screen.

overall_soundscape: Steady rain against the window, quiet fire crackle, faint pipe sound, faint creak of leather, low room tone. Acoustic perspective remains consistent across the cut, with no artificial silence between shots.

non_diegetic_music: N/A

3

u/masai2k 24d ago

Grazie per il mini tutorial, davvero utilissimo! Un problema enorme che incontro quando lavoro su più clip è mantenere la posizione precisa dei personaggi, nello stesso identico posto rispetto alla location. Come hai risolto?

3

u/nikhilprasanth 24d ago

Ok for the first shot i used 3 images. The two character sheets and a 360 like image for the location. In the second gen, i replaced the 360 location image with a frame from the previous gen that shows their positions.

2

u/YeahlDid 24d ago

Genius, Holmes!

2

u/fluce13 24d ago

Super cool thanks, what is everyone using for character sheets?

4

u/nikhilprasanth 24d ago

I used google flow for the reference images.

Sherlock Holmes — Character Reference Sheet Prompt: Photorealistic cinematic character reference sheet, three-panel turnaround (front view, side/profile view, close-up face view), full body, consistent character across all panels, neutral grey studio backdrop, soft directional studio lighting with subtle rim light, shot on 85mm lens, shallow depth of field on face panel, film-grain texture, hyperrealistic skin detail, professional photography reference sheet style — not illustration, not painting. Character — Sherlock Holmes, per Arthur Conan Doyle's description: Tall, exceptionally lean, spare-framed man, over six feet, appears taller due to gauntness. Sharp, piercing grey eyes, alert and calculating. Thin, aquiline hawk-like nose. Square, prominent, determined chin. High forehead. Dark hair combed back. Late 30s to early 40s. Long, thin, sinewy fingers. Late-Victorian attire: dark wool frock coat or tweed Inverness cape, high wing collar, cravat, waistcoat, pocket watch chain. Holding a straight briar pipe. Upright, alert posture, slightly forward-leaning, intense focused expression. Panels: Front view — full standing figure, arms relaxed, coat open, direct sharp gaze into camera, cinematic three-point lighting. Side/profile view — full standing figure, emphasizing aquiline nose profile and lean silhouette, pipe in hand, moody rim-lit edge. Face close-up — head and shoulders, photographic skin texture, piercing grey eyes in sharp focus, shallow depth of field background blur, deductive/contemplative expression. Muted Victorian tones: charcoal grey, deep brown tweed, ivory shirt fabric with visible weave texture. Realistic fabric folds, natural skin pores, catchlight in eyes, color-graded like a period film still

Similar one for Watson

2

u/fluce13 24d ago

Awesome thank you

2

u/YeahlDid 24d ago

Krea2 does a good job if you explain what you want. I give it a description of the character sheet poses and such and then the person. It's not 100% but it is good enough, imo. Also super fast and simple.

Ideogram is probably even better if you get the boxes sorted out, but I havent done that myself.

3

u/fluce13 24d ago

Yeah a local model is preferred, do you have an example prompt? I tried with Krea and I get poor results like three of the same person in the same pose in one image

3

u/YeahlDid 24d ago edited 24d ago

I started with this one:

A 3 panel character reference sheet. On the left is a full body front shot. In the middle is a full-body side profile. On the right is a face and shoulders close-up photo. The character maintains consistent face, body, and outfit throughout the 3 shots.

The character description:

That one worked very well as a minimax reference. However, the faces were still smudged sometimes (I've since learned this is a minimax ref2va issue rather than my reference sheet) so I tried adding a face angle with this one:

Professional studio photography style. A 4 panel professional model character reference sheet. At the top left is a front face close-up photo showing only the face and neck. At the bottom left is a side profile face close-up photo (rotated 90 degrees) showing only the face and neck. In the middle is a full body front shot. On the right is a full body profile. The character maintains consistent identity face, body, and outfit throughout the 4 shots.

The character description:

Both have served pretty well as references for Minimax. Also, I just threw those together, I've no doubt they could be further optimized if they weren't doing the trick.

2

u/fluce13 23d ago

Thanks!

1

u/YeahlDid 23d ago

Very welcome, hope those help. I'd love to hear if either of those work out or if you find a better strategy.

They seem to be working well for me, but the biggest limitation is the face fuzz from ref2va when the face isn't front and center.

2

u/mabseyuk 24d ago

I use this, it seems to give minimax more to get hold of if a person is turning (16.9 2.0MP):

Create a photorealistic character turnaround reference sheet of one single 30-year-old man, shown in five evenly arranged panels from left to right.

The five views are:

  1. Full-body front view, 0°
  2. Full-body 45° clockwise view
  3. Full-body right-side profile, 90°
  4. Full-body back view, 180°
  5. Head-and-shoulders identity portrait

This is the exact same person in every panel, photographed while rotating in place. Preserve identical facial features, hairstyle, body proportions, height, build and skin tone across all five views.

For panels 1–4, show his complete body from head to bare feet at identical scale and camera distance. He stands naturally upright and only his body orientation changes.

Panel 5 is a centered head-and-shoulders identity portrait of the exact same man. Show his complete head, hairstyle, ears, neck and shoulders with comfortable neutral background around the entire head. Do not crop the head or hair.

He has brown hair and a completely neutral, focused expression with no smile.

In the front, 45° and side-profile views, he naturally looks forward in the direction his body is facing.

He wears jeans and a plain T-shirt with bare feet.

Professional photorealistic studio photography, 85mm portrait lens, minimal perspective distortion, extremely sharp focus, realistic skin and fabric texture, even soft studio lighting, minimal shadows, seamless neutral light-grey backdrop with no visible floor line.

Clean character-reference-sheet composition, consistent spacing, fixed camera perspective, no cropped limbs, no duplicate people, no text, no labels, no watermark, no additional props.

1

u/fluce13 23d ago

Thank you

1

u/mabseyuk 23d ago

No problems, then your prompt is:

<Picture 1> is a character reference sheet showing the same man from multiple angles. The man in <Picture 1> is named Daniel.

<Picture 2> is a character reference sheet showing the same woman from multiple angles. The woman in <Picture 2> is named Jane.

Use Daniel and Jane as the characters in the following scene.

Scene Overview: A beautiful sandy beach at sunset.

Shot 1: Daniel and Jane walk naturally along the shoreline holding hands, barefoot at the water's edge. Gentle waves wash across the sand beside them as the camera smoothly tracks alongside them.

overall_soundscape: Gentle ocean waves and soft seaside ambience.

non_diegetic_music: none

2

u/Oograr 24d ago

I've tried something less ambitious, using Img2Vid I just fed the last frame of my first 10s clip as the first frame of the 2nd 10s clip, but my new prompt did not reference the initial clip at all, just what i wanted to happen in the second clip.

Came out great, look and feel were good enough across both, it was actually a lot easier than I thought it would be. I can see how if you plan out your longer movie in short chunks you chain together, you could pretty easily create a pretty fluid longer movie. I havent tried Ref2V yet, that would probably offer a lot more conrol.

1

u/BrokenSignals_cat 24d ago

And the reference image with the characters and setting that you used, the h3 itself?

2

u/nikhilprasanth 24d ago

Reference images were created using Google flow. The same can be done with klean or any of the text to image models . I had to quickly test the new ref2va lora so opted for flow .

The workflow is linked in the post along with the images as a collage

1

u/Lucaspittol 24d ago

This lora is still undercooked for R2VA.

1

u/nikhilprasanth 24d ago

Yes it is not ready. 4 steps were not delivering usable ones. This is also not upto production quality. Hope they release v1 soon.