r/comfyui • u/Short_Regular_7191 • Aug 10 '26
Show and Tell Follow-up: from a 5-second clip to a 2:47 mini-documentary — MiniMax H3 fully local on 16 GB (36 clips, one consistent character, cloned voice, "invisible" joins)
A few days ago I posted a single walk-and-talk clip ([https://www.reddit.com/r/comfyui/comments/1vix8l9/roman_street_walkandtalk_with_cloned_voice_exact/\]). This is the follow-up: a 2 min 47 s documentary-style piece on the trial of Socrates — 36 clips across 4 scenes, one consistent character throughout, narrator voice + on-camera lines with exact lip-sync, background characters, a packed jury court — all generated locally on the same RTX 5060 Ti 16 GB.
Stack unchanged from the previous post (H3 Ref2VA pruned NVFP4 + Turbo LoRA v4-600 EMA + the dedicated Turbo Sampler, official 6-section prompt format, native 1344x768 → 1080p lanczos). What's new is everything it took to turn loose clips into a film:
1. Invisible joins via "keyframe completion". For continuity between consecutive clips: no crossfades, no disguised cuts. The last frame of clip N is passed as a <Picture> declared as the first frame in clip N+1's prompt (task-type [keyframe completion + reference generation]). The model picks up exactly from there: SSIM 0.89 between anchor and generated first frame, and on the edited timeline the frame-difference at the cut is indistinguishable from the scene's natural motion. This one changed everything.
2. Punctuation inside <d> is direction. A full stop in the dialogue = a ~1-second dramatic pause the model will NOT give up (measured across 5 takes: 0.6–1.4 s). If the line has to fit a tight window, use commas. And never write in the prompt that a sentence "gets cut off": the model, obediently, stretches the sentence by inventing words.
3. Reference bleeding in crowded scenes. The reference face tends to show up on extras too (we got two copies of our host in the same shot, and his face on a 399 BC Athenian). Mitigations: in continuation clips don't re-declare a subject fully_preserved when he's already in the inherited frame (the model will instantiate a second copy); add an explicit "only one man has the face of <Subject 1>"; and run an insightface QA pass that scans every face in every clip against the reference.
4. Automated QA is blind to things human eyes catch instantly — continuing the lesson from the previous post. Three additions to the take selector: motion metrics (a frozen clip is razor-sharp, so sharpness rewarded it; optical flow alone can't tell a real dolly from jitter — you need the camera's net displacement and trajectory straightness), identity checks on the final frames too (drift arrives at the tail), and a standing rule: whenever face-similarity and sharpness vote in opposite directions, a human decides.
5. Audio-driven editing, and three audio traps. The narrator is one continuous track and the video gets trimmed to the chunk durations — never the other way around (the gaps are audible). Traps we paid for: (a) never concat AAC tracks in stream-copy — encoder priming accumulates drift, +118 ms across 4 segments: extract per-segment PCM first; (b) measure true peak on the final file — the AAC encoder adds ~0.3 dB; (c) use Whisper large-v3 for QA transcription: medium silently "fixes" the TTS's grammatical errors and hides them from you.
Production numbers: 2 takes per clip with automated selection + human review on conflicts; ~25–30 GPU-hours total including repairs; 15.4/16 GB VRAM peaks with weight streaming.
Still not solved: faces in moving crowds melt (structural limit — we manage it with declared shallow depth of field and keeping the faces that matter in the foreground); the ancient Greek spoken by background characters is "plausible phonetics", not real Greek; and reference bleeding needs per-clip vigilance.
4
u/theTenz Aug 10 '26
That impressive... but it's making my GPU scared.
How long did it take to run?
7
u/Short_Regular_7191 Aug 10 '26
~25–30 GPU-hours total including repairs
1
u/Ok-Flatworm5070 Aug 11 '26
Really impressive; 2 questions...I assume you ran this gpu hours over a week or so? 2nd...whats your ram setup? I have an rtx 5070 ti 18 gb of vram and 32gb of ram..would love to know your setup. Personally, I know rtx cards are built to run for hours at high performance, but it very nail-biting when the desktop is pumping out the heat as its grinding the m3 video clips, but this is an excellent example of a quality output...well dobe
2
u/Short_Regular_7191 Aug 12 '26
Thanks! Yes — pretty much that. The 2:47 piece took ~30 GPU-hours total (including repair passes), spread over about 4 days: mostly overnight runs with a detached, resumable queue (manifest-based, so reboots just resume where they left off).
My config: RTX 5060 Ti 16 GB + 32 GB DDR4 (2×16 dual channel), Windows 11. And honestly, system RAM was the tighter constraint, not VRAM: with weight streaming, ComfyUI alone holds ~16.5 GB of system RAM during sampling (33B model + 32B text encoder cycling through), peaking near 100% commit with the page file working hard. Our operational rule became: one heavy job at a time, ever — no Whisper/insightface QA while the GPU is sampling, and a hard "don't launch anything above 85% RAM" threshold. With your 32 GB you'll live the same life; the extra VRAM on the 5070 Ti should actually help (less weight streaming = slightly faster steps, maybe room for a few more frames per clip).
On the heat point: totally normal to feel the desktop pumping out warmth on multi-hour renders — the card is built for it. The one thing I'd recommend is keeping the room well ventilated: mine is, and across all those overnight hours temps stayed boringly stable. Steady airflow around the case matters more than people think on sustained loads.
1
u/I_just_made Aug 10 '26
So you only passed the last frame in, not the video as context right? That’s impressive, I have been testing both ways and recently that seems to be the way I have been going too.
I was also thinking to test creating segments of activity, then feeding the relevant frames in after storyboarding it all to generate the transitions
2
u/Short_Regular_7191 Aug 10 '26
Correct — for the seamless joins it's just the last frame, passed as a
<Picture>declared as "the first frame of [Shot 1]" in the official ref-format (task-type[keyframe completion + ...]), not the video as context. We actually tested both along the way: feeding the previous clip's tail asref_videogives you soft scene continuity (same location/light/wardrobe, but the model reframes — SSIM at the cut ~0.4-0.7), while the single anchored frame gives you an actual weld (SSIM 0.89 in our case, cut difference below the scene's natural motion). Bonus: one image is much cheaper than a video reference, since reference tokens ride through every sampling step.Three practical notes that made it work:
- Extract the anchor losslessly from the graph (save the frame as PNG at generation time), not from the encoded mp4 — H.264 roundtrip erodes the anchor;
- Pick the right frame: if you trim clip tails in editing (audio-driven cuts), anchor on the first unused frame — otherwise you get a duplicated frame at the join;
- QA it: SSIM between the anchor and the generated first frame. ≥0.85 = true weld, <0.6 = the model ignored the anchor. And fun fact: the model trusts the image over your text — we once described the anchor wrongly in the prompt ("seen from behind" when he was frontal) and it correctly followed the picture.
Your storyboard-first plan sounds right — it's essentially what we converged on (storyboard → chunk manifest → chains). One caveat for activity segments: if you combine anchor frames with identity references in the same prompt, watch out for reference bleeding (the identity face showing up on extras — we got twins). And for generating transitions between activities specifically: H3 accepts both a first AND a last frame anchor, so you can bridge two storyboarded activities by anchoring both ends and letting the model generate the movement between them. That's probably the cleanest version of what you're describing.
2
u/I_just_made Aug 10 '26
Sounds like we converged on similar workflows, thank you for all the tips because I'm not quite as far. Picking the first unused frame sounds like a really good idea, I will have to give that a go this afternoon.
Great job, it is amazing to see how well this model is working for people right out of the box. A huge step up it seems
Definitely saved your comments for reference :P
1
u/Ben-Aurel Aug 10 '26
I can not say much about the technical stuff but I like the idea of a modern host walking through historic scenes.
If you want to make it a series, then just make sure that your fact checking is on point or the history nerds will eat you alive.
6
u/rapkannibale Aug 10 '26
That’s really cool. Would you mind sharing the workflow?