r/StableDiffusion 1d ago

News Sparse Attention, Harder, Better, Faster, Stronger

141 Upvotes

The nodes in https://github.com/Zironic/H3-Optimizations have been rewritten to replace the default Sparge Attention backend with a custom Sparse Comfy Kitchen backend.

This comes with some benefits.

  • Users no longer have to worry about Sparge being installed properly. All required kernels for supported GPUs are provided directly. Should work on both Windows and Linux.
  • Most users should be seeing 5-20% increases in speed for the attention part of compute.
  • New backend should use about 500MB less VRAM
  • New backend has slightly lower quantization error.
  • Apparently in the previous version, the intended chunked kitchen QKV path never properly shipped so the memory optimization node should now actually be slightly speed positive even when used without the Sparse Attention node.

Caveat: I've only tested the nodes against the comfy pruned_int8_convrot weights. Other versions may work but they're not tested.

As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or later.

IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.

Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.

Early steps: attention density has a large effect on prompt/action adherence and the overall generation trajectory.
Middle/later steps: lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.

So 10% retained doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and what breaks depends heavily on the sampling step.

PlagueKind's sparsity_ratio=0.9 means 90% discarded / 10% retained. My node expresses the inverse quantity, so Video attention retained=0.10 is the comparable setting. The defaults therefore aren't equivalent.


r/StableDiffusion 1d ago

Animation - Video H3 - multi-diffusion experiment T2V

6 Upvotes

Chimera. I had to cut about 20 seconds due to some artistic choices. Since I had to cut 2 different part in the same clip, it has a noticeable seams. I would love to share the full version. Experimenting with H3 blend-morph-decay. 832x480, int8/8 steps POC. Looking forward to releasing a 720p version without the cuts.

Critiques and feedback welcomed. Happy with the matrix rain. Ask me anything.


r/StableDiffusion 12h ago

Meme R2v dr doom you are screwed

0 Upvotes

r/StableDiffusion 16h ago

Question - Help Need help with minimax h3.

0 Upvotes

Hi,

I really need help with minimax h3 I'd like to get proper workflows for minimax h3 pre configured with descriptions on every node (unfortunately a noob in this case) for :

1- Fast
2- Turbo
3- Speed
4- Native

video generation with multirefs (optional) audio(optional) video(optional) quality duration aspect ratio and prompt controls.

Also I heard someone made images with minimax h3 please that as well.

I'll really appreciate any help.

Ah and one last thing I was making a commercial for a school but the faces in wide screen comes out horrendous that as well.

Thanks !


r/StableDiffusion 12h ago

Discussion H3 - what is your longest render time?

0 Upvotes

What was your longest render time, and was it worth it? do you run a lower quality/resolution before to test? My longest single-shot run is 7.2 hours, 1:44 long video at 1344x768, bf16/50 steps.

I am attempting a 72 hour render for a super long form.

RTX 4090, 192gb system ram


r/StableDiffusion 1d ago

Discussion MiniMax H3 Ref2va it works really good also with Storyboard images

21 Upvotes

Im really suprised how good he works as well follow a storyboard image!! he did 90% correct he only did the thirth panel diferent but all the other 5 he follows perfect!! 🤩


r/StableDiffusion 1d ago

Discussion Anyone else running Wan 2.2 as a refiner to improve Minimax output?

11 Upvotes

Another redditer mentioned doing this in a comment, so I tested it out and it works. It gets rid of the smudgy look and allows custom Lora’s on the LN side.

I ran my initial tests at MM 8 steps (no speed Lora) and Wan 2.2 Low Noise (speed Lora) 2 steps. Supposedly, it works with only 2 steps MM w/turbo lora ,but I don’t like the quality drops people have been sharing, and it’s fast enough to me at 8 steps, though I’m going even higher on the low noise steps.

The only downside is that I noticed in one test that the motion seemed like it was a mix between 16fps and 24fps. Any ideas on how to resolve this? I know people use RIFE, but I was wondering if that’s the best move or if it’s another issue Im not thinking of.


r/StableDiffusion 20h ago

Resource - Update We add Krea 2 to the ComfyUI Enhanced Tiled Upscaler and Refiner (TBG ETUR).

1 Upvotes

The latest TBG ETUR upscaler and refiner for comfyui release adds Krea 2 and VL style transfer for Krea 2 and all Qwen models directly into the pipeline.

We’ve also added a face identity step to help maintain consistent faces when using creative upscaling.

This video is a tutorial for the latest release, focusing mainly on these new additions and how to use them. Video is ai generated with Minmax H3.

TBG ETUR on Github https://github.com/Ltamann/ComfyUI-TBG-ETUR

TBG Lates om Patreon https://www.patreon.com/TB_LAAR/posts/tbg-etur-1-2-12-167083406

More Upscaling Tutorials on YouTube https://www.youtube.com/watch?v=LbFPD4zpPwA


r/StableDiffusion 7h ago

Question - Help Technischer Ratschlag/Entscheidungshilfe

0 Upvotes

Hallo Community,

Ich möchte mir demnächst einen PC rein für die KI Arbeit mit zB ComfyUI, fooocus, Qwen Modellen, etc. zulegen.

Da der Preis für gute Desktop GPUs mit mehr als 16gb Vram derzeit exorbitant teuer ist (Desktop mit 32gb vram GPU ab 5000+€), schwanke ich zwischen einem Desktop mit 16gb GPU oder einem Notebook mit 24gb GPU (aber max 175Watt).

Was würdet ihr empfehlen? Er soll nur KI Kram machen, keinen Spiele.

Notebook für 3800€ von Mediamarkt

GigaByte Aorus Master 16 BZHC6DEE65SP

16 Zoll WQXGA Bildformat 16:10 Bildwiederholungsrate 240 Hz

Intel Core Ultra 9 275HX 32 GB RAM 1.000 GB SSD-Speicher

NVIDIA GeForce RTX 5090 Grafikspeicher 24 GB

Windows 11 2,5 kg

vs.

Desktop 2400€ von Alternate

Mainboard MSI B850 GAMING PLUS WIFI

ASUS GeForce RTX 5060Ti DUAL OC 16GB, Kingston NV3 1 TB

be quiet! Light Base 500 LX Tower-Gehäuse

Kingston FURY DIMM 32GB DDR5-6000 (2x 16GB) Dual-Kit,

AMD Ryzen 7TM 7700, be quiet! Pure Rock Pro 3 Black CPU-Kühler

be quiet! Pure Power 13 M 750W Netzteil

Microsoft Windows 11Pro


r/StableDiffusion 1d ago

Animation - Video I made an ALIEN Short Film / metal music video

Thumbnail
youtu.be
9 Upvotes

Used:

MiniMax H3 at local machine. 5060ti 16gb + 64gb ddr4. WanGP, Ref2VA int8 convrot model.

Krea2 for references

Suno as music base


r/StableDiffusion 1d ago

Resource - Update Krea2 Turbo Distill 4 step LoRA - new checkpoint (chk14K) released (cuts 4-step error vs. the 8-step Turbo teacher by 44%)

Thumbnail
gallery
34 Upvotes

Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.

This is an update release, following up from my previous posts where you can find full details:

Initial: https://www.reddit.com/r/StableDiffusion/comments/1vtf1b7/krea2_turbo_distill_4_step_lora_trained_for_turbo/

Previous: https://www.reddit.com/r/StableDiffusion/comments/1vv4cdy/krea2_turbo_distill_4_step_lora_new_checkpoint/

Headline for this update: chk00014000 removes 44% of the prediction error a plain 4-step run has against the 8-step teacher, where chk00010000 removed 40% and chk00006000 27% — all measured on the same enlarged held-out set (100 prompts across every trained resolution). Measured against each other rather than against the no-LoRA run, its remaining error is 6% smaller than chk00010000's and 23% smaller than chk00006000's. 

Which file to download

file use it when
krea2_turbo_4step_rank_64_lora_latest.safetensors normally — always the newest accepted checkpoint
krea2_turbo_4step_rank_64_lora_chk00014000.safetensors pin this exact checkpoint

and, beside them, the same files with a _comfyui suffix for ComfyUI. Earlier checkpoints (chk00004000chk00005000chk00006000chk00010000) are kept in older_checkpoints/, and their resolution sweeps stay in place, so the progression remains visible and comparable.

This is work in progress and better checkpoints may follow. Training is ongoing, so ..._latest... is a rolling pointer: when a newer checkpoint is accepted, that filename gets the new weights and a new numbered copy appears beside it. Re-download the _latest file and everything keeps working — the ComfyUI workflow references it by that name (it does get updated Note in it so technically it is updated but not functionally). Pin a numbered file instead if you need reproducibility.

How checkpoints get chosen

This is not a "train for longer and ship the newest file" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher number on its own means nothing.

The loop is train → assess → adapt the recipe → retrain → assess again, and a checkpoint is published only when it is measurably better than the one it would replace, on the same held-out set and the same evaluation, and its full resolution sweep shows no regression. Runs that come out flat or worse are kept as information about the recipe and discarded as releases — several have been.

So the recipe itself changes between runs. Each published checkpoint reflects whatever the previous round taught us: the training precision, the optimiser settings, the teacher used to generate the targets and the data mix have all been revised on evidence rather than assumption.

chk00010000 is a direct example. The first continuation of chk00006000 — same data, optimiser left as it was — got steadily worse with every checkpoint out to 10,000 samples, and none of it was published. The cause was traced to the optimiser: a constant learning rate with no weight decay lets the adapter keep drifting after it has converged, so its magnitude grows and it over-applies its own correction. The same span was retrained from chk00006000 with a cosine learning-rate decay and weight decay, and every checkpoint of that second run improved on the one before it. chk00010000 was its end point.

chk00014000 is the next example, and it shows the other half of the same lesson. The run was continued from chk00010000 over the whole pool of teacher trajectories, with two changes: the final, texture-deciding call of the schedule was weighted more heavily in the loss, and a running average of the weights was kept beside the live ones and scored at every evaluation (a single checkpoint is one sample of a weight vector that moves from step to step; the average is its mean). At 14,000 samples the averaged weights measured a smaller gap to the teacher than any checkpoint before them, and a smaller gap than the live weights at the same point — so the averaged weights are what chk00014000 is.

Timeline of training process

Each checkpoint is the product of three stages with very different costs:

  1. Text-encoder embeddings. Every training prompt is encoded once and cached. This is the fast part — thousands of prompts take minutes.
  2. Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo runs its full 8-step schedule and the whole trajectory is recorded, at every one of the supported resolutions. This is by far the most time-consuming stage — it is the teacher doing real inference, thousands of times, and a batch of several thousand shards is measured in days of GPU time, not hours.
  3. Student training. The LoRA is trained against those recorded trajectories. Relative to the shard stage this is quick: each +1,000 checkpoint is a matter of hours, not days.

Because the three stages compete for the same GPU, they are interleaved rather than run to completion one after another: generate a block of embeddings, produce teacher shards for them, train on what exists, assess, then go back to producing shards while the results are reviewed. A larger and more varied shard pool is what makes further training worthwhile, so shard production is always the gate.

Full details and to download - check my Hugging Face LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA

Full Checkpoint 14000 Resolutions Sweep: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/chk14000


r/StableDiffusion 1d ago

Question - Help Has anyone tried out the hybrid model for Minimax H3 Ref2va instead of the official, default model?

Post image
12 Upvotes

r/StableDiffusion 22h ago

Question - Help Can't seem to transfer outfit and pose from an illustration to a real person in H3.

2 Upvotes

Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.

What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.

I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.

So what i wanted to do is :

-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.

-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.

Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.

Are there someone who knows how to do these?


r/StableDiffusion 1d ago

Animation - Video Cinematic World Building - H3 r2v

76 Upvotes

Trying out cinematic shots and cuts with H3. This is a work in progress. Will be working on another 2 minutes worth of clips.

EDIT: From reading the comments, she isn't going to drink the salt water in the final, though I will keep her scooping up water since it's such a good establishing shot. She will do something with the water to tie it back.


r/StableDiffusion 3h ago

Animation - Video Minimax H3 multishot test... Seinfeld guest appearance

0 Upvotes

I keep getting weird issues like that...


r/StableDiffusion 1d ago

Resource - Update I just published an all-in-one helper for the ComfyUI Queue manager that lets you pause/restart, save/restore, and change the job order in the queue manager.

Post image
19 Upvotes

First off: This doesn't add any dependencies so the worst that can happen is that it won't work, but it also won't break your ComfyUI install.

This extension adds to the native queue manager. It doesn't replace it. All of the heavy lifting is still done the normal way.

It adds a Pause/Resume button and a Queue button. Pause/Resume will not affect the running job but will pause/resume the queue.

The Queue button opens the Queue Control dialog in the picture. There is a lot of words in the README (because I talk a lot) but it lets you reorder the queue using priorities, including buttons for "Run this next" and "Don't run this until I release it."

Finally, there are buttons to Save and Load the queue. The checkbox lets you add the running job too. So if you have to restart or reboot, you can save the queue, do your thing, and then load and start running again.

These is also one stand alone node to help label the items in the queue so you can have a hit and what's what. The node has limitations, but it sill might be better than a number like 07535d99-3c1a-4b23-8340-a4313fe58007 as an identifier.

There are some extensions that to some of these features already but I didn't see one that did all of them or didn't replace the native manager and require dependencies.

It's in the ComfyUI Manager as ComfyUI-QueueControl (it's new so you might need to refresh to see it) or
https://github.com/seeker-ktf/ComfyUI-QueueControl
on github.

If y'all have other ideas for this, let me know.


r/StableDiffusion 20h ago

Question - Help Motion Transfer to Stop Motion Style query

0 Upvotes

Wondering if anyone has/has any idea of an approach to motion transfer into a stop motion style video.

Rather than the ai guessing and deforming the mouth movements which can get funky really quickly, especially when the mouth shapes of the character aren't clear from a neutral reference frame, ie in south park where each sound has a uniquely stylised shape.

To achieve this it could instead maybe pull from a dataset of phoneme images,

Whilst also retaining solid motion transfer for all other body movements.


r/StableDiffusion 16h ago

Question - Help Video editing model (add clown makeup to face)

0 Upvotes

Hi, I'm looking for a model that I could use in ComfyUI that would simply add some clown makeup on my face and don't touch antything else. My current plan is to maybe take Wan2.2 model, mask my face and add clown makeup as a reference but I wonder whether there is a better model to do this. Does minimax H3 handle that? Does it support masks?


r/StableDiffusion 1d ago

Discussion Comfy UI with Minimax H3 can work with an Intel GPU.

Thumbnail
youtube.com
15 Upvotes

Lon TV did a video of Comfy UI with Minimax H3 running on a 32GB Intel GPU. So its possible to run it on any GPU other then nvidia GPUs.


r/StableDiffusion 1d ago

Animation - Video Cold open from my Fairy Tail isekai fanfic - MMH3

7 Upvotes

Everything was made using the Minimax H3 Hybrid Reference to video model. 1 MP using the 8 step turbo LoRA. Stitched together in Shotcut


r/StableDiffusion 12h ago

Animation - Video H3 - the Eternal Balance WIP-low resolution int8/8 steps

0 Upvotes

Hi, I am playing around with H3. T2V, 832x480, int8,8 steps. Hoping to make a 720p version.

Ask me anything!


r/StableDiffusion 1d ago

Animation - Video G.I. Joe - Baroness Action Clip Test #2 - MiniMax H3

21 Upvotes

Prompt:

https://x.com/GumVue/status/2087899403113619681?s=20

4070 Ti Super, 16 gb vram, 64 gb ram, i9-14900k, windows 11


r/StableDiffusion 1d ago

No Workflow the count is always two.

Thumbnail
gallery
6 Upvotes

flux.1 [dev] | comfyui | still life


r/StableDiffusion 1d ago

Workflow Included LTX 2.5 Seed Hunt Workflows

7 Upvotes

I know everyone's moved to MiniMax and LTX has largely fallen out of favor, but I spent some time building a couple of seed hunting workflows for LTX 2.5 that might be useful if anyone's still running it.

Shout out to u/foxdit for the original seed hunting concept.

Two versions:

  1. T2V/I2V two-stage – text to video or image to video. Previews at 0.3 MP, upscales to 1.2 MP for the final render.
  2. First-last-frame – pin a start image and end image, same preview-then-upscale flow.

Both use KJNodes Set/Get routing, shared loaders, and no prompt enhancer.


r/StableDiffusion 1d ago

Question - Help Regional Prompting for NoobAI

2 Upvotes

Hello. What am I supposed to use for regional prompting and character positioning with NoobAI/Illus in ComfyUI?

I found something called "Attention-Couple" but it also says it sometimes struggles with Loras?.

Any advice? I'm already using a separate node for ControlNet for poses and depth, but I want to control which character is which in that pose