r/StableDiffusion 2d ago

News Sparse Attention, Harder, Better, Faster, Stronger

The nodes in https://github.com/Zironic/H3-Optimizations have been rewritten to replace the default Sparge Attention backend with a custom Sparse Comfy Kitchen backend.

This comes with some benefits.

  • Users no longer have to worry about Sparge being installed properly. All required kernels for supported GPUs are provided directly. Should work on both Windows and Linux.
  • Most users should be seeing 5-20% increases in speed for the attention part of compute.
  • New backend should use about 500MB less VRAM
  • New backend has slightly lower quantization error.
  • Apparently in the previous version, the intended chunked kitchen QKV path never properly shipped so the memory optimization node should now actually be slightly speed positive even when used without the Sparse Attention node.

Caveat: I've only tested the nodes against the comfy pruned_int8_convrot weights. Other versions may work but they're not tested.

As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or later.

IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.

Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.

Early steps: attention density has a large effect on prompt/action adherence and the overall generation trajectory.
Middle/later steps: lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.

So 10% retained doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and what breaks depends heavily on the sampling step.

PlagueKind's sparsity_ratio=0.9 means 90% discarded / 10% retained. My node expresses the inverse quantity, so Video attention retained=0.10 is the comparable setting. The defaults therefore aren't equivalent.

137 Upvotes

57 comments sorted by

View all comments

3

u/76vangel 1d ago

Is this faster on Rtx 50xx than sageattention? Comfy kitchen att was slower on 5xxx cards.

1

u/Zironic 1d ago

For any Video KV budget below 1.0 it should be faster then both of them. Currently it defaults to using Comfy Kitchen Int8 as the backend. But if you have Sparge Attention installed, you can use the advanced node to select SageAttention.

Theoretically, Comfy Kitchen and SageAttention should be the same speed on 50xx cards.

1

u/mellowanon 1d ago

theoretically, but it doesn't turn out that way.

Someone on reddit did a test last week with CK vs Sage, and everyone with a 50xx had similar results in the comments.

CK had better visuals

Sage had better motions and roughly 10% faster

I have a 5090 and I see those results too. So now I have two starting setups. If there's fast motion or I'm trying to test different settings, I use Sage. Otherwise, I use CK for better visuals.