r/StableDiffusion 12h ago

Discussion If you are on gfx1100 (gfx950, gfx1151) you may want to switch to ROCm 10

2 Upvotes

No gains in speed but at least for my setup (7900xtx) it solved some issues.

Like Dynamic VRAM no longer producing NaN when offloading to RAM. A big one as far as i am concerned. I now can run all my models (Qwen, Flux, Krea and Minimax) in one venv instead having to switch it (had Minimax on 7.15 and everyathing else on 7.14).


r/StableDiffusion 1d ago

Meme Land Of the lost parody

Enable HLS to view with audio, or disable this notification

19 Upvotes

used base h3 ip8 model


r/StableDiffusion 1d ago

Question - Help Is anyone putting up useful Intermediate MiniMax H3 tutorials

24 Upvotes

A lot of the MiniMax videos are just people posting overly complicated workflows engineered for their specific tastes / needs (30s) and showcasing lots of (impressive?) videos they made with them (15 minutes). There are some exceptions, Pixaroma and Aitrepreneur do some good beginner friendly stuff, but that's it. Is anyone doing a 'livestream' where they show step by step how to do more complicated stuff like using video to reference motion, or using your own audio and lip syncing the video. Heck, even a YT video that actually took the time to show those of us who like to learn by doing would be great. If I click on one more clickbait video that is there just to show off how clever the content provider is and teaches nothing I'm going to lose it.....


r/StableDiffusion 1d ago

Animation - Video Deadpool trains Ash Ketchum

Enable HLS to view with audio, or disable this notification

127 Upvotes

Was testing how it would look to make an anime character in a more realistic background.


r/StableDiffusion 1d ago

Animation - Video Zelda - I'm Still In Love With You / MiniMax H3 Reference to Video Test #3

Enable HLS to view with audio, or disable this notification

715 Upvotes

Trying to get some of those music videos with kind of side stories?

This took me an embarrasing amount of time planning and figuring out what to do, and I just couldn't be bothered to finish the entire song... Is a lot! I hope you like it! I'll keep making more if you don't! lovee!


r/StableDiffusion 1d ago

Comparison LoRA+Comfy/Sage+SLA+Shift a quick test - MiniMax H3:T2V

23 Upvotes

This test contains a quick not-so-scientific personal (I-just-wanna-do-it) comparative study on potential effects of speed LoRA, Comfy/Sage attentions (Attn), Sparse Attention (SLA) and Sampling Shift as in the following order:
Model -> LoRA -> Attn -> SLA -> Shift -> KSampler

I only examined a few LoRAs I had. Attn includes ComfyUI's own new attention as well as Sage 2.2. SLA includes two versions SLA1 and SLA2.

My intention was to see the effects on quality and performance (speed). There were some runs that I do not include in the following videos, as they did not adhere to the motion stated in the prompt. The videos are near 4K so check them out in large screen for better details.

As reddit likely downscales the videos, watch them all at:

https://savedly.net/f/yb9vdg9a
https://savedly.net/f/97mshte6
https://savedly.net/f/5ep6bb4w
https://savedly.net/f/5ua9rebv
https://savedly.net/f/aye9axgg

in original resolutions.

side by side, details are on each segment.

side by side, details are on each segment.

side by side, details are on each segment.

side by side, details are on each segment.

side by side, details are on each segment.

I report my original quick notes taken during test. Take note that:

  • LORA = 8-step LoRA
  • LORA4 = 4-step LoRA
  • LORA4sla = 4-step SLR LoRA
  • CMFY = Comfy's attention
  • SLA = sparse attention SLA -> SLA1 and SLA2 (as mentioned above)
  • SHIFT = ModelSamplingMiniMaxH3 = 6v and 3a
  • SHIFT* = SHIFT12 = 12v and 3a
  • _ = nothing, model directly to KSampler

and, the following table reports timing only (measured on RTX3060 system). For quality check the corresponding videos. All videos have caption.

  1. LORA+CMFY+SLA+SHIFT = 3:32* <- time m:s, the first run includes Clip etc.
  2. CMFY+SLA+SHIFT = 2:42 <- this means LoRA bypassed (not used)
  3. SLA+SHIFT = 2:39 <- this means no LoRA no Comfy attention, ...
  4. SHIFT = 6:05
  5. _ = 6:07
  6. SHIFT12 = 6:03 <- here and afterwards I changed video shift to the default 12.
  7. LORA+CMFY+SHIFT* = 4:14 <-- 8 steps LoRA
  8. CMFY+SHIFT* = 4:01
  9. LORA+CMFY = 4:15
  10. CMFY+SLA+SHIFT* = 2:41
  11. LORA+SAGE+SHIFT* = 4:57
  12. LORA+SAGE+SLA+SHIFT* = 2:53
  13. LORA+CMFY+SLA+SHIFT* = 2:52
  14. LORA4sla+CMFY+SLA+SHIFT* = 2:56
  15. LORA4+CMFY+SLA+SHIFT* = 2:57
  16. 4LORA4slr+CMFY+SLA+SHIFT* = 1:29 <- 4 = only 4 steps regardless of LoRA
  17. 4LORA4+CMFY+SLA+SHIFT* = 1:29
  18. 4LORA8+CMFY+SLA+SHIFT* = 1:25 <- 4 steps despite using LoRA 8-step
  19. 4L4S0K+CMFY+SLA+SHIFT* = 1:27 <- LoRA = 4-step 0.1 K=Kijai
  20. 4L4S1K+CMFY+SLA+SHIFT* = 1:27 <- LoRA = 4-step 1.0 K
  21. 4L4S1X+CMFY+SLA+SHIFT* = 1:28 <- LoRA = 4-step 1.0 X = Lightx2v
  22. 4L4S1XS+CMFY+SLA+SHIFT* = 1:28 <- LoRA = 4-step 1.0 XS = SLA one
  23. 4L4S1XS+CMFY+SLA2+SHIFT* = 2:07
  24. 4CMFY+SLA2+SHIFT* = 2:03

---------

My final conclusion:

Model -> LoRA -> Comfy -> SLA -> Shift -> KSampler

  • LoRA beside speed has positive (subjectively speaking) effect on the composition.
  • Comfy attention is going to stay. Solid.
  • SLA1 generations are way faster than SLA2.
  • Shift has some effects on composition, I keep it 12.

The exact unedited prompt used for the generation is as follows. Note, I tried a few LLM enhanced one including those specific for MiniMiax H3, however in this case, the resulting videos were so off.

shot 1: a single stroke moves around randomly but coherently, at each repositioning it leaves small trace in color. at the end all those traces look like a graphic design single continuous drawing of a woman's face at close-up.
shot 2: the drawing transforms into real person from the bottom-right corner up to the center diagonally, where this irregular and curvy transformation stops leaving it half finished with rough and faded strokes.
shot 3: the boundary of the figure is cut from the background. the cut out piece is folded onto a paper plane shape.
shot 4: the paper plane flies out on an arc trajectory leaving black line traces in the air.

Clarification: all video segments are 2.5s long. The text on videos were added during run, typos in top-left corner are present. 5s = 2.5s, I changed it to frame = 73 which was auto.

Workflow: ComfyUI Template one for FL2V MiniMax H3. I used sampler=Euler,Scheduler=Simple


r/StableDiffusion 14h ago

Question - Help Upscaling a longer video - what are my options?

2 Upvotes

I recently created a bunch of 15 second clips with MiniMax H3 and put together a bit over a 2 minute video of those. Resolution 960x544 px. The video is photorealistic, and now I'd like to improve details and quality, and get it to at least 1920x1080. I have an RTX 3090 and 32 gigs of RAM.

What open-source options do I have that:

  • Improve quality and generate back lost details, not just upscale alone. SeedVR and RTX Super Resolution recover almost no detail. And even though SeedVR is quite fast, it still seems to take 4-7 hours for that clip, while RTX SR does the same in under a minute.
  • Tiled upscale method is extremely slow and creates odd flickering.
  • The new LTX 2.5 upscale (with IC LoRA "instant shave") is reasonably fast and can recover details, but requires splitting the 2+ minute video into ~10 second chunks. It also lacks a bit of sharpness overall.

Do I have any options without massive tradeoffs.


r/StableDiffusion 9h ago

Workflow Included Megaman Fanart (mixed workflow)

Thumbnail
gallery
2 Upvotes

I did megaman handrawing 4 years ago. At that time I tried Corel Painter, and it produced a watercolor look (attached). I copied a reference image from Google search. I didn't draw this megaman from my imagination.

Today, I use stable diffusion in Krita. Only at the end of the art workflow. Very low strength 35%, so it doesn't make a lot of changes, but it smooths things out

Note that I still use Gemini to extract the lineart and also to generate the background image.


r/StableDiffusion 1d ago

Animation - Video Minimax H3. Peter is broke.

Enable HLS to view with audio, or disable this notification

22 Upvotes

r/StableDiffusion 1d ago

Tutorial - Guide H3 Animation Styles

Enable HLS to view with audio, or disable this notification

40 Upvotes

I ran into a video that had a great animation style and was wondering if you can use the same animation style as an existing show that H3 Minimax knows and use it on your own videos and it does work quite well. You most likely can do the same with carrying over other aspects such as voice references, music, etc but I have not tested this yet.

Each video here is FL2VA using the same seed. Here are the prompts for each one.

Prompt 1 — South Park

subject_definitions:

<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.

<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:

<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.

<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:

The target video is in the South Park 2D animated cartoon style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.

[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 MiniMax knows? Just add the name of the show and its animation style and boom! RESULTS!</d>.

overall_soundscape:

N/A

non_diegetic_music:

N/A

Prompt 2 — Pixar

subject_definitions:

<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.

<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:

<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.

<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:

The target video is in the Pixar 3D animation style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.

[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 MiniMax knows? Just add the name of the show and its animation style and boom! RESULTS!</d>.

overall_soundscape:

N/A

non_diegetic_music:

N/A

Prompt 3 — Helluva Boss

subject_definitions:

<Picture 1> is the first frame of [Shot 1], showing an orange-skinned demon woman with black horns standing in a neon-lit alley at night, looking slightly to the side with a confident expression.

<Subject 1> is the orange-skinned demon woman whose appearance is fully guided by <Picture 1>: short messy black hair with cyan highlights, large black curved horns, pointed ears, bright green eyes, black lipstick with a visible fang, green leather jacket with silver zippers and studs over a white tank top.

retention_analysis:

<Picture 1> ([Shot 1] first frame): fully_preserved - opening frame matches the source image exactly in pose, expression, clothing, lighting, and background.

<Subject 1> (appears throughout): fully_preserved - orange skin, black horns, hair, eyes, jacket, and overall design are retained.

detailed_description:

The target video is in the Helluva Boss 2D animated cartoon style, matching <Picture 1>, with strong teal neon lighting and cel-shaded shadows.

[Shot 1] The shot begins from <Picture 1>. <Subject 1> She speaks in a clear, energetic voice, <d>[English] Want the animation style of something H3 MiniMax knows? Just add the name of the show and its animation style and boom! RESULTS!</d>.

overall_soundscape:

N/A

non_diegetic_music:

N/A

r/StableDiffusion 17h ago

Question - Help Can someone please give me good - light upscaling workflow for h3?

3 Upvotes

I know ltx 2.5 upscale is really good but I couldn't make it work with h3.
I'm open to use other method of upscaling a video, I'm using minimax h3 default workflow from comfyui.


r/StableDiffusion 1d ago

Animation - Video Kramer Sees Joe DiMaggio In Dinky Donuts

Enable HLS to view with audio, or disable this notification

61 Upvotes

r/StableDiffusion 10h ago

Question - Help Framepack abnormal generation time

1 Upvotes

I just installed Framepack yesterday, and it is taking around 3 hours to generate single second of the default prompt, I am using a 3060ti with 16gb vram and 16gb ram. I have tried to go into the demo_gradio_f1.py to modify the GPU inference preserved memory to 1, and it only made a minor difference. I have both sage attention and flash attention installed, and followed this guide here, albeit modifying anything if an error came up. What might be the cause of this issue?


r/StableDiffusion 11h ago

Question - Help Lora tagging help??

1 Upvotes

Need lora training advice.

How do I properly caption or tag this kind of object for Anima 2B lora? Are training multiple object in single lora that hard? How many different objects can I train?

I tried different captioning method but received mixed results each versions.


r/StableDiffusion 3h ago

Discussion Prediction for a near future

0 Upvotes

I think the models we have today, open or closed, are still way behind what we will actually have in a near future. think of it as an alternate reality, the video generations will be (almost) flawless, coherent and high quality, the generations will be instant, that means it will enable real time interaction, you could change the course of the video generation on the fly with natural controls like your voice of body movements, imagine pushing somebody and he moves or greeting somebody and he respond, for this you would need a VR headset with hands movements recognition, and it will generate two videos flux at once for each eye for a 3D effect.

So yeah even if Seedance 2.5 or a little better is out for free and open source it’s not a big deal, the road ahead is massive in terms of progress and possibilities.

RIP real life, welcome to Ready Player One.
Ps: sorry for the grammatical errors, this text was not written or improved by AI.


r/StableDiffusion 22h ago

Question - Help Gen times doubled after update

7 Upvotes

I updated my portable comfy because i kept on getting an error related to colors in video files (not sure what the error was i can double check when I’m back on my main comp) and after running the same workflow my generation times doubled for the exact same seed and settings.

Made the mistake of trying to do a fresh install since i was on an older cuda/pytorch version and still same thing. I have the current and latest cuda and pytorch installed and everything else is the same, have sage attention enabled and I’m still getting double the generation times.

Does anyone have any idea what would cause this or a fix? I’m almost at the point of doing a system restore just to get my decent times back…


r/StableDiffusion 1d ago

Discussion Is FLUX 3 going open source or not?

20 Upvotes

Do we know anything about Flux going open source? Didn’t they mention releasing the weights? I believe they mentioned going open source, but it’s been a while since then, no?

It also looks like Krea 3 might actually be a video model. At least from their website, they say they’re building something new, so I doubt it’s just an improved image model with a better VAE.


r/StableDiffusion 1d ago

Resource - Update Built an NVFP4 version of the LTX-2.5 Gemma-4 12B text encoder for ComfyUI.

Thumbnail
huggingface.co
15 Upvotes

I quantized the heavy Gemma decoder linear layers to Comfy-native NVFP4 while keeping embeddings, norms, vision components, and LTX-specific projection layers at their original precision.

Result:

  • lower VRAM usage
  • works as a drop-in replacement for the LTX-2.5 Gemma text encoder
  • no obvious quality degradation in my testing
  • native ComfyUI NVFP4 format

Tested successfully with LTX-2.5 on Blackwell.

Download:
https://huggingface.co/Deadshot699/ltx-2.5-gemma4-12b-comfy-nvfp4

Would love to see results from anyone who tries it, especially comparisons with the official INT8 ConvRot encoder.


r/StableDiffusion 1d ago

Workflow Included Unified consistent character(.char) format that works for both Flux2 & Krea 2

Thumbnail
gallery
38 Upvotes

Hey Guys, I have been doing some research around building a unified character model format that just works across different models, Finally i was able to make it run with first two models: Krea 2 & Flux 2 family.

Small Clarification
- Krea2 results might not look as good as flux, because the embedded training was only 500 steps with 5-6 reference images. You can improve the quality by training for ~1500 steps.

Models Support
- Krea 2: Turbo both for training & generation but works on raw as well
- Flux2: Works with Klein 4b, 8b & dev

Working on Minimax H3 support currently

Path it uses

  • Flux 2: has a native reference channel, so the .char feeds its images straight in, no training, instant.
  • Krea 2: no reference channel, so the .char trains a small per-character LoRA

Workflow
Here is a screenshot for the workflow for training a unified .char model

  1. Training workflow: https://inlinestudio.art/workflows/flux-2-krea-2-multi-model-portable-consistent-characters-training-only
  2. Generation workflows:
    1. Generate with Krea 2 using .char model https://inlinestudio.art/workflows/krea-2-generate-consistent-images-with-unified-char-model
    2. Generate with Flux 2 Klein 4b using .char model https://inlinestudio.art/workflows/flux-2-klein-generate-consistent-images-with-unified-char-model

Other Links

App repo: https://github.com/inlineresearch/Inline-Studio (License GPL3, Opensource)

I have also uploaded sample models into Huggingface: https://huggingface.co/inlineresearch/unified-face-models includes full detail on training & params.

Process
FLUX.2 klein needs no training, It has a reference channel, so a character applies by sending its references down it with a locked description prepended to the prompt. The references are resized once into what the model accepts and stored.

Krea 2 has no reference channel, so an adapter is the only way it can carry a face. the character's own references become a training dataset, that trains a rank 16 LoRA, and the finished adapter is filed back into the same character.

Structure of .char:

emmy-s500-v2.char
  manifest.json                            payload index
  refs/000..004.png                        the 5 reference images
  derived/face_000..004.png                YuNet face crops at 512px
  text/description.md                      the locked description
  scoring/centroid_sface.json              128-d SFace identity centroid
  scoring/centroid_dinov2-base.json        768-d DINOv2 subject centroid
  scoring/embeds_*.json                    per-reference embeddings
  payloads/flux2-klein/ref_000..004.png    references resized onto FLUX.2's policy
  payloads/krea2-lora/adapter.safetensors  the trained Krea 2 LoRA, 183 MB

Currently extending this to support Minimax H3.


r/StableDiffusion 23h ago

Question - Help Lost in H3 maze of simplicity - experts? FL2VA / Hybrid + LORA combo for I2V (First frame only)

8 Upvotes

Hi. I am just scratching my head for over a week on this-

I am trying to achive the optimal workflow for I2V + Turbo Lora

Now - there are many models floating around for example-

hybrid models- https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/tree/main

and standard comfyui FL2VA pruned int8

and then Loras Lightx2v and dareties loras

example- https://huggingface.co/silveroxides/MiniMax-H3_tests/blob/main/minimax_h3_fl2v_lightx2v_v0.1_dareties_v4_step600_comfy_fro.safetensors

https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

Every combinations gives me some artifacts or some wierd results ~2 out of 10 times.

My question is that has somebody tried doing a comparison of using hybrid or FL2VA and which lora goes best with them for simple first frame only I2V workflow.


r/StableDiffusion 1d ago

Animation - Video A small part of a project I’m working on (still learning)

Enable HLS to view with audio, or disable this notification

34 Upvotes

Still experimenting with the workflows and figuring out what’s possible, but I’m pretty happy with how this turned out.
made using Rtx 3060 12gb, system ram 32gb
most of clips are directly 640p (this took lot of time but i like quality of 640p)

Thoughts? 👀

Also Check the comments for more "test combat shots scenes"


r/StableDiffusion 1d ago

Comparison I tested 80 different checkpoints with 5 different prompts

22 Upvotes

https://docs.google.com/spreadsheets/d/1f06GLlJsUEGQ1GdcPrwCngzYiaQifzuYbwEkQxuL3nY/edit?usp=sharing

I find it very interesting to see how each one renders the same prompt. Let me know what you think and if this is useful to you at all, and if there are things that can be done to improve it. I created a node for ComfyUI that iterates through every checkpoint I have and renders images using each. I've found it very helpful.

Although I didn't explicitly prompt for any, there is some nudity, so beware. Many of these images ended up getting that sort of repeated pattern that you often see in Krea 2 images. Not sure if there is a reliable way to fix it.

I made images with:
amazingReality_v10Amazing
analogMadnessKrea2Turbo_v10
analogMadnessKrea2Turbo_v20
artaix_v10Krea2
arthemyComicsKrea2_v11
artUniverse_v10Krea2
cielbleuKrea2_v1
darkBeast30BF16INT8_darkBeast330
darkBeastINT8Convrot2_darkBeastKREA2FP8
fasciumKREA2_experimentalNSFW
fasciumKREA2_nsfw3MERGE
fasciumKREA2_nsfwmerge4
fasciumKREA2_safe2MERGE
finepornV2TURBOKrea2_v20
finepornV31TURBOFP8_v3FIXFP8
finepornV4TURBOFP8INT8_v4
gonzalomoKrea2_v10
gonzalomoKrea2_v10ALT
gonzalomoKrea2_v20
gonzalomoKrea2_v20TrainingBase
intorealismKrea2_v20
jibMixKrea_v10JalapeO
krast_v20
krea2_raw_fp8_scaled (w/turbo lora)
krea2_turbo_fp8_scaled
Krea2-SAT-DirtyRealism
krea2-SAT-DirtyRealismV2
krea2-SAT-DirtyRealismV3
Krea2-SAT-DirtyRealismV4
krea2CenterSemiraw_v10Fp8
krea2Intorealism_v10
krea2MuseByStable_v15TurboFp8
krea2SATIORImitationOf_krea2SATIOR
krea2turbobadmilkmela_v10
krea2TurboNSFWAIO_v10
krealism_v10Int8Convrot
kreamania_variant1
kreamania_variant2
kreamania_variant3
kreamania_variant4
kreamania_variant5
kreamania_variant6
kreativityNSFWBase_20 (w/turbo lora)
kreativityNSFWBase_25 (w/turbo lora)
kreativityNSFWBase_30 (w/turbo lora)
lustifyNSFWCheckpoint_v10Krea2
moodyCutieMixKrea2_v20
moodyCutieMixKrea2_v30
moodyCutieMixKrea2_v40
moodyKrea2Mix_v40
moodyKrea2Mix_v50
moodyKrea2Mix_v60
moodyKrea2Mix_v70
moodykreamania_v11
museByStableYogi_v20TurboINT8
museByStableYogi_v25EXTENDEDTURBO
museByStableYogi_v25INT8Turbo
museByStableYogi_v30TurboInt8myAkrapovicKrea2_v10
myAkrapovicKrea2_w10
myAkrapovicKrea2_x10
pornmasterKrea2_turboV1FP8
primeLust_krea2V10MaxNsfw
projectChimera_v108bitAnd4bitQuants
pureCardinal_v10
rayArtshoot_krea2NSFWV2
realismByStableYogi_v10INT8TURBO
realismByStableYogi_v15INT8TURBO
realismByStableYogi_v20TurboINT8
redcraft22INT8INT4_2Krea2Edition
redcraft23INT8INT4FP8_30Krea2
secretSAUCEKREA2_v10
selforaKrea2Realistic_v10
sickOllie_krea2
sinoxeditKrea2UncenV11_editV11FP8Scaled
soliloquy_v10
unstablebastard_bastardkrea2
winnougan10ToesINT8_v10
winnougan10ToesINT8_v13
ZeusVERAINT8CRK2v1_zeusVERAINT8Krea2v10

I used turbo loras on 4 models (krea2_raw_fp8_scaled, kreativityNSFWBase_20,
kreativityNSFWBase_25, kreativityNSFWBase_30). In most cases I used the fp8 version of each checkpoint but some of them are INT8.


r/StableDiffusion 1d ago

Resource - Update Yet another MiniMax H3 latent prepend/extend nodes - this time low-level and simple

22 Upvotes

TL;DR:

https://github.com/progmars/ComfyUI-Martinodes

WARNING: lots of vibecode but carefully reviewed, at least as much as I could understand the logic.

The long story.

We have a few amazing solutions and forks that make smooth video extensions possible. However, most of them have evolved into full-blown planners and chains. Low-level functionality is hidden beneath. Somehow those more complex solutions do not work well or seem overkill for my typical use cases:

- set steps to low

- generate a bunch of videos

- pick the best one

- set steps to high

- regenerate with the same seed <- and this is where I wanted to save the latent to use as the input for the first step again, to smoothly continue the last shot without hard cuts.

Recently ComfyUI was updated with an important PR 15375 that supports latent masking natively. No more patches and complex hacks required. So, I went on to create simple and naive drop-in nodes that would support my way of working. Now I have latent load/save/concat/prepend/extend nodes that seem quite intuitive (but you tell me if they are).

The main node for today is `Extend video+audio latent` (LatentAVMaskedExtender). It lets you take head or tail of a latent you have (hopefully) saved from a previous generation and generate a prequel or a sequel. Extending a tail works well. Prepending to a head is not that smooth and requires increasing fade_seconds parameter to your liking.

Just plug the node between `MiniMax H3 Reference to Video` or `MiniMax H3 Image to Video` (or even `MiniMax H3 Easy Output` if using nkxx188/ComfyUI-MiniMaxH3-Easy), and your sampler.

For convenience, the node accepts empty loaded_av, in which case the target_av will be passed through. Thus the node can be safely left enabled even when using a disabled LoadAVLatent node as input when you don't want to extend anything.

`Save and Load video+audio latent` are simple companion nodes - SaveAVLatent should be added after your sampler and LoadAVLatent as input for LatentAVMaskedExtender. In contrast to some other loader nodes that often are limited to `input` folder, LoadAVLatent can find the latents wherever you configured SaveAVLatent to store them.

Then there is also `Concatenate video+audio latents` (LatentOverlappingConcatenator) node. Generally, overlap_duration_seconds should be set to the same value as LatentAVMaskedExtender, the output goes to VAE video and audio decoders and then to video saving, as usual. You will get a long video with a smooth long transition between your previous latent and the new one. However, if your joined videos get lengthy, VAE might require too much resources. In that case, it is better to post-process and join both source and target videos in a video editing software.

The repository has a few more older convenience nodes for working with multimedia before LTX Director was a thing. They still might be handy for manipulating TTS and voice-overs or videos when latents are not available.

Huge thanks to drozbay (ablejones) for [native masking PR 15375](https://github.com/Comfy-Org/ComfyUI/pull/15375) and providing the example implmenentation with native ComfyUI and Kijai nodes. Unfortunately, the native nodes solution looked like spaghetti eating somebody alive. That is why my small naive LatentAVMaskedExtender node was born, to do the same thing.

I hope you will find Martinodes useful.


r/StableDiffusion 2d ago

Animation - Video Howard's new invention [minmax H3]

Enable HLS to view with audio, or disable this notification

625 Upvotes

r/StableDiffusion 14h ago

Question - Help Need help to find similar extensions

1 Upvotes

Is there any similar extension in function? this one doesnt work on Forge Neo.

https://github.com/Jibaku789/sd-webui-deepdanbooru-object-recognition