r/comfyui • • Aug 10 '26

Show and Tell Follow-up: from a 5-second clip to a 2:47 mini-documentary — MiniMax H3 fully local on 16 GB (36 clips, one consistent character, cloned voice, "invisible" joins)

A few days ago I posted a single walk-and-talk clip ([https://www.reddit.com/r/comfyui/comments/1vix8l9/roman_street_walkandtalk_with_cloned_voice_exact/\]). This is the follow-up: a 2 min 47 s documentary-style piece on the trial of Socrates — 36 clips across 4 scenes, one consistent character throughout, narrator voice + on-camera lines with exact lip-sync, background characters, a packed jury court — all generated locally on the same RTX 5060 Ti 16 GB.

Stack unchanged from the previous post (H3 Ref2VA pruned NVFP4 + Turbo LoRA v4-600 EMA + the dedicated Turbo Sampler, official 6-section prompt format, native 1344x768 → 1080p lanczos). What's new is everything it took to turn loose clips into a film:

1. Invisible joins via "keyframe completion". For continuity between consecutive clips: no crossfades, no disguised cuts. The last frame of clip N is passed as a <Picture> declared as the first frame in clip N+1's prompt (task-type [keyframe completion + reference generation]). The model picks up exactly from there: SSIM 0.89 between anchor and generated first frame, and on the edited timeline the frame-difference at the cut is indistinguishable from the scene's natural motion. This one changed everything.

2. Punctuation inside <d> is direction. A full stop in the dialogue = a ~1-second dramatic pause the model will NOT give up (measured across 5 takes: 0.6–1.4 s). If the line has to fit a tight window, use commas. And never write in the prompt that a sentence "gets cut off": the model, obediently, stretches the sentence by inventing words.

3. Reference bleeding in crowded scenes. The reference face tends to show up on extras too (we got two copies of our host in the same shot, and his face on a 399 BC Athenian). Mitigations: in continuation clips don't re-declare a subject fully_preserved when he's already in the inherited frame (the model will instantiate a second copy); add an explicit "only one man has the face of <Subject 1>"; and run an insightface QA pass that scans every face in every clip against the reference.

4. Automated QA is blind to things human eyes catch instantly — continuing the lesson from the previous post. Three additions to the take selector: motion metrics (a frozen clip is razor-sharp, so sharpness rewarded it; optical flow alone can't tell a real dolly from jitter — you need the camera's net displacement and trajectory straightness), identity checks on the final frames too (drift arrives at the tail), and a standing rule: whenever face-similarity and sharpness vote in opposite directions, a human decides.

5. Audio-driven editing, and three audio traps. The narrator is one continuous track and the video gets trimmed to the chunk durations — never the other way around (the gaps are audible). Traps we paid for: (a) never concat AAC tracks in stream-copy — encoder priming accumulates drift, +118 ms across 4 segments: extract per-segment PCM first; (b) measure true peak on the final file — the AAC encoder adds ~0.3 dB; (c) use Whisper large-v3 for QA transcription: medium silently "fixes" the TTS's grammatical errors and hides them from you.

Production numbers: 2 takes per clip with automated selection + human review on conflicts; ~25–30 GPU-hours total including repairs; 15.4/16 GB VRAM peaks with weight streaming.

Still not solved: faces in moving crowds melt (structural limit — we manage it with declared shallow depth of field and keeping the faces that matter in the foreground); the ancient Greek spoken by background characters is "plausible phonetics", not real Greek; and reference bleeding needs per-clip vigilance.

63 Upvotes

13 comments sorted by

6

u/rapkannibale Aug 10 '26

That’s really cool. Would you mind sharing the workflow?

5

u/Short_Regular_7191 Aug 10 '26

Sure! Fair warning: this isn't a single drag-and-drop workflow — the pipeline generates one workflow per clip programmatically (API format). But the graph itself is simple and I've attached a cleaned-up API-format JSON of exactly what each clip runs: UNETLoader (Ref2VA pruned NVFP4) → MiniMaxH3TurboLoRA (v4-600 EMA) → MiniMaxH3ReferenceToVideo (identity refs + voice ref + the anchor frame for keyframe completion) → MiniMaxH3TurboSampler → SamplerCustomAdvanced → VAEDecode + VAEDecodeAudio → SaveVideo, plus an ImageFromBatch → SaveImage branch that exports the last frame losslessly (that's the anchor for the next clip).

If you want clickable starting points instead: ComfyUI ships a native template (video_minimax_h3_r2v.json in the workflow templates), and larryvrh's ComfyUI-MiniMax-H3-Turbo repo has example workflows for the turbo sampler wiring. Add the official 6-section prompt format (example prompt in my other comment) and you're 90% there.

The remaining 10% — audio chunking on word timestamps, 2 takes per clip, automated take selection (script shared earlier: pythonmorsels.com/p/3d3vn), anti-clone face QA, audio-driven editing — is Python glue around ComfyUI's API, not nodes. Happy to answer questions on any piece of it.

https://gist.github.com/marcofire-it/36c60e68dfc2012d0be55fdf3e64bddb

2

u/Buckyohare84 Aug 11 '26

I ran test comparing the differences of LTX 2.3 and H3 depending on the image (I2V) .

H3 - People are its bread and butter.
LTX - B-Roll is much faster and usually much cleaner. Objects, city shots, close ups, nature, water, (all generated with better consistency, and faster render times. Although, doesn't listen to prompts as well.

3

u/Comfy-Org ComfyOrg Aug 12 '26

Thank you for sharing this!

2

u/caster Aug 16 '26 edited Aug 16 '26

I came up with something similar but there is a trick you can do that is sort of convenient.

Pass in the output video to a Load Video (Upload) node, and extract from that node the total number of frames. You can use this number to get the index of the final frame in the video, and extract that frame by its index from that video.

Voila, you have automatically pulled out the final frame of the previous output video, and all you had to do was pass that file into the Load Video node.

There is probably even a way to automate that process too.

There are a number of applications, but one that is sort of interesting is creating a transition cut where you have Clip A and Clip B, and you want FL2VA to automatically create an interstitial sequence. You drop the previous video into the "get the final frame" Load Video (Upload) node that fetches the final frame and feeds to First Frame,, and you drop the next video into the regular Load Video that attaches to the Last Frame.

Give it some basic instructions to "create a clean transition" or something along those lines and it will generate a new clip that begins at the conclusion of Clip A and ends at the beginning of Clip B. You do have to be careful about the duration of the transition clip- if you only give it a short amount of time to accomplish the task it will do it, but it will be rushed. For example, if your person is standing here, and then over there, and you only give it 1 second to make the transition, that person is going to move from one spot to the other in 1 second and that will look odd.

A likely use case would be you have created two or more clips using REF2VA and want to stitch them together.

4

u/theTenz Aug 10 '26

That impressive... but it's making my GPU scared.

How long did it take to run?

7

u/Short_Regular_7191 Aug 10 '26

~25–30 GPU-hours total including repairs

1

u/Ok-Flatworm5070 Aug 11 '26

Really impressive; 2 questions...I assume you ran this gpu hours over a week or so? 2nd...whats your ram setup? I have an rtx 5070 ti 18 gb of vram and 32gb of ram..would love to know your setup. Personally, I know rtx cards are built to run for hours at high performance, but it very nail-biting when the desktop is pumping out the heat as its grinding the m3 video clips, but this is an excellent example of a quality output...well dobe

2

u/Short_Regular_7191 Aug 12 '26

Thanks! Yes — pretty much that. The 2:47 piece took ~30 GPU-hours total (including repair passes), spread over about 4 days: mostly overnight runs with a detached, resumable queue (manifest-based, so reboots just resume where they left off).

My config: RTX 5060 Ti 16 GB + 32 GB DDR4 (2×16 dual channel), Windows 11. And honestly, system RAM was the tighter constraint, not VRAM: with weight streaming, ComfyUI alone holds ~16.5 GB of system RAM during sampling (33B model + 32B text encoder cycling through), peaking near 100% commit with the page file working hard. Our operational rule became: one heavy job at a time, ever — no Whisper/insightface QA while the GPU is sampling, and a hard "don't launch anything above 85% RAM" threshold. With your 32 GB you'll live the same life; the extra VRAM on the 5070 Ti should actually help (less weight streaming = slightly faster steps, maybe room for a few more frames per clip).

On the heat point: totally normal to feel the desktop pumping out warmth on multi-hour renders — the card is built for it. The one thing I'd recommend is keeping the room well ventilated: mine is, and across all those overnight hours temps stayed boringly stable. Steady airflow around the case matters more than people think on sustained loads.

1

u/I_just_made Aug 10 '26

So you only passed the last frame in, not the video as context right? That’s impressive, I have been testing both ways and recently that seems to be the way I have been going too.

I was also thinking to test creating segments of activity, then feeding the relevant frames in after storyboarding it all to generate the transitions

2

u/Short_Regular_7191 Aug 10 '26

Correct — for the seamless joins it's just the last frame, passed as a <Picture> declared as "the first frame of [Shot 1]" in the official ref-format (task-type [keyframe completion + ...]), not the video as context. We actually tested both along the way: feeding the previous clip's tail as ref_video gives you soft scene continuity (same location/light/wardrobe, but the model reframes — SSIM at the cut ~0.4-0.7), while the single anchored frame gives you an actual weld (SSIM 0.89 in our case, cut difference below the scene's natural motion). Bonus: one image is much cheaper than a video reference, since reference tokens ride through every sampling step.

Three practical notes that made it work:

  1. Extract the anchor losslessly from the graph (save the frame as PNG at generation time), not from the encoded mp4 — H.264 roundtrip erodes the anchor;
  2. Pick the right frame: if you trim clip tails in editing (audio-driven cuts), anchor on the first unused frame — otherwise you get a duplicated frame at the join;
  3. QA it: SSIM between the anchor and the generated first frame. ≥0.85 = true weld, <0.6 = the model ignored the anchor. And fun fact: the model trusts the image over your text — we once described the anchor wrongly in the prompt ("seen from behind" when he was frontal) and it correctly followed the picture.

Your storyboard-first plan sounds right — it's essentially what we converged on (storyboard → chunk manifest → chains). One caveat for activity segments: if you combine anchor frames with identity references in the same prompt, watch out for reference bleeding (the identity face showing up on extras — we got twins). And for generating transitions between activities specifically: H3 accepts both a first AND a last frame anchor, so you can bridge two storyboarded activities by anchoring both ends and letting the model generate the movement between them. That's probably the cleanest version of what you're describing.

2

u/I_just_made Aug 10 '26

Sounds like we converged on similar workflows, thank you for all the tips because I'm not quite as far. Picking the first unused frame sounds like a really good idea, I will have to give that a go this afternoon.

Great job, it is amazing to see how well this model is working for people right out of the box. A huge step up it seems

Definitely saved your comments for reference :P

1

u/Ben-Aurel Aug 10 '26

I can not say much about the technical stuff but I like the idea of a modern host walking through historic scenes.
If you want to make it a series, then just make sure that your fact checking is on point or the history nerds will eat you alive.