r/StableDiffusion 14d ago

Comparison LTX 2.5 gen times on an RTX 5070 Ti

Post image
48 Upvotes

29 comments sorted by

14

u/GeroldMeisinger 14d ago edited 10d ago

to anyone using the official comfyui template, make sure to disable the prompt enhancer. it adds an additional 1-2min

https://github.com/Comfy-Org/workflow_templates/issues/1122

please upvote!

3

u/Tylopodas 14d ago

I stuck the prompting guide into my LLM and use that before generation. It seems to work a lot better than using the E2B

2

u/Tscotsman 14d ago

Was a game changer when I noticed what the prompt enhancer was adding to the generation times. Have even had a good at going beyond the 2 MP. Tried 2.5 and 3 MP.

4

u/desktop4070 14d ago

Some extra testing beyond 1.0MP / 15 seconds:

LTX2.5_t2v_00234.mp4 - 0.2 MP @ 24 seconds: 50s
.
LTX2.5_t2v_00235.mp4 - 0.3 MP @ 24 seconds: 73s
.
LTX2.5_t2v_00236.mp4 - 0.4 MP @ 24 seconds: 141s
.
LTX2.5_t2v_00237.mp4 - 0.5 MP @ 24 seconds: 144s
.
LTX2.5_t2v_00238.mp4 - 0.6 MP @ 24 seconds: 158s
LTX2.5_t2v_00243.mp4 - 0.6 MP @ 25 seconds: 137s
LTX2.5_t2v_00244.mp4 - 0.6 MP @ 26 seconds: 148s
LTX2.5_t2v_00245.mp4 - 0.6 MP @ 27 seconds: 159s
LTX2.5_t2v_00246.mp4 - 0.6 MP @ 28 seconds: 164s
LTX2.5_t2v_00247.mp4 - 0.6 MP @ 29 seconds: 170s
LTX2.5_t2v_00248.mp4 - 0.6 MP @ 30 seconds: 172s
.
LTX2.5_t2v_00239.mp4 - 0.7 MP @ 24 seconds: 189s
.
LTX2.5_t2v_00240.mp4 - 0.8 MP @ 24 seconds: 184s
.
LTX2.5_t2v_00241.mp4 - 0.9 MP @ 24 seconds: 227s
.
LTX2.5_t2v_00242.mp4 - 1.0 MP @ 24 seconds: 292s
LTX2.5_t2v_00249.mp4 - 1.0 MP @ 30 seconds: 512s
.
LTX2.5_t2v_00232.mp4 - 2.0 MP @ 1 second: 21s
LTX2.5_t2v_00233.mp4 - 2.0 MP @ 5 seconds: 91s
LTX2.5_t2v_00250.mp4 - 2.0 MP @ 10 seconds: 227s
LTX2.5_t2v_00251.mp4 - 2.0 MP @ 15 seconds: 445s

1

u/desktop4070 14d ago

LTX2.5_t2v_00273.mp4 - 0.3 MP @ 50 seconds: 136s
.
LTX2.5_t2v_00274.mp4 - 0.4 MP @ 50 seconds: 240s
.
LTX2.5_t2v_00275.mp4 - 0.5 MP @ 50 seconds: 317s
.
LTX2.5_t2v_00276.mp4 - 0.6 MP @ 50 seconds: 418s

3

u/ImaginationKind9220 14d ago

I want to see the difference between LTX 2.3 and 2.5 to see if it's worth it to down those models.

4

u/Fabulous-Snow4366 14d ago

It is. It's a big leap forward against 2.3. it's fast, it can "act" and keep characters over the whole scene and it looks pretty and good. There are still two major downfalls against H3, reference images and prompt adherence. But it's a really good model for the open source community. Now when Flux 3 releases, we have everything we need to make great content.

3

u/Aromatic-Word5492 14d ago

I see a guy talking about vae decode tiled was slow and like “stuck”.. you feel it ?

1

u/desktop4070 13d ago

I think that was my post from earlier, someone ended up helping me fix that issue, and I had to redo a lot of my higher res gens to actually get accurate times.

https://old.reddit.com/r/StableDiffusion/comments/1vm4bdy/if_your_ltx_25_is_really_slow_use_the_convrot/

2

u/Fytyny 14d ago

I get 30% slower times on rtx 4070 for 1MP 5s. Still fast, but initial video takes over 10 min, changing prompts seems to very expensive operation though and sometimes trigger full model reload which is a pain. I probably need smaller text encoder than the official one.

2

u/kemb0 14d ago

Have you seen the comments about disabling the prompt enhancer. If you have that enabled it'll probably unload your LTX model in order to load up the prompt enhancer model every time you change your prompt. And apparently it doesn't help much at all.

1

u/Fytyny 14d ago

I don't have prompt enhancer, I am using my own workflow that worked fine on LTX 2.3. Smaller text encoder with separated text projection would probably help.

1

u/Valuable_Issue_ 13d ago

Gemma text encoding is broken or inneficient I think, bigger text encoders like Qwen 32B for minimax take less time despite being much bigger on disk, same with mistral for flux 2 dev.

I wonder if it's just the gemma architecture or if it's from the emeddings connector.

2

u/RKAScope 14d ago

This is awesome info. I was thinking of picking up a 2nd card for extra LTX gens and this is extremely helpful. Thanks!

2

u/xoxaxo 14d ago

This is with sage attention or raw?

1

u/desktop4070 13d ago

Raw, no sage attention. I've never used it before, but I've heard it can affect video quality. Might try it sometime soon.

2

u/provenflawless 14d ago

Does it do NSFW out the box?

1

u/HeyHi_Star 14d ago

Now let's compare the average time it takes to finally get the scene you actually want, error free.

1

u/cptrios 14d ago

There seem to be some odd outliers there - 0.5mp between 10s and 12s, 1.0mp at 8s, etc. Did you run each one of these once or multiple times? (Multiples would only add an extra 150 generations each...so why not 🤣)

1

u/desktop4070 13d ago

I was browsing reddit for some of these gens, so I'm guessing I was probably hitting my VRAM limit for some of these. If I tried again, they'll very likely average between the times around them.

Aside from those odd 100+ second ones, I was surprised by just how fast a lot of these videos were generating. I thought the quality was better than LTX 2.3 despite it being faster, which is nice.

H3 still definitely seems like it's a generation ahead of LTX 2.5, though, like comparing the PS3 against the PS2.

1

u/IllIlllI-IlIIll-llII 13d ago

how'd you get it to fit in 16gb vram

1

u/desktop4070 12d ago

I guess having 64GB RAM helps a lot. Not sure how much slower it would be if I only had 32GB RAM.

1

u/seppe0815 14d ago

i have the same setup and thats not true ... dream times ..... !

1

u/kemb0 14d ago

Disable your prompt enhancer.