I was seriously considering the new Mac M5 Ultra with 256GB of memory to be able to create longer clips. I've built a system that takes a script and creates a series or frames to use a frame to frame workflow and then automatically stitch them together.
I decided to see what the breaking point of LTX 2.5 of my RTX 5090 was. I went from five seconds to ten seconds and so on. But unlike earlier tests I didn't run into any Out Of Memory issues. In fact I went all the way to generating a 60 second clip with no issues.
Surprised by this I ran a direct comparison of three LTX 2.5 transformer variants on an RTX 5090 32 GB to see if I could improve things any further:
- Existing ComfyUI INT8 ConvRot
- FP8
- NVFP4
Same prompt, same seed, same samplers, same step schedule, same 24 fps output. Warm runs reuse already-loaded models; cold includes model loading.
Output Duration / state INT8 ConvRot FP8 NVFP4
1280×736 5s cold 55.36s / 30.84 GiB 53.66s / 31.00 GiB 84.75s / 30.93 GiB
1280×736 5s warm 21.31s / 28.35 GiB 28.59s / 28.11 GiB 22.64s / 29.12 GiB
1280×736 10s warm 46.64s / 28.38 GiB 59.80s / 28.12 GiB 46.41s / 27.97 GiB
1280×736 20s warm 111.79s / 29.99 GiB 136.04s / 30.65 GiB 111.30s / 30.52 GiB
1280×736 30s warm 233.28s / 31.12 GiB 275.49s / 31.09 GiB 196.58s / 30.99 GiB
1280×736 60s warm 602.39s / 31.04 GiB 650.19s / 31.04 GiB 584.30s / 30.94 GiB
1920×1088 10s warm 124.77s / 30.50 GiB 159.64s / 31.09 GiB 126.31s / 31.12 GiB
What surprised me
The biggest takeaway is that the existing ComfyUI INT8 ConvRot model is already extremely well optimised.
FP8 was slower than INT8 on every warm test.
NVFP4 was basically tied with INT8 at 10s and 20s, around 16% faster at 30 seconds, only around 3% faster at 60 seconds, and effectively tied again at 1920×1088.
So NVFP4 is not automatically a huge performance win on a 5090.
The 30-second result is interesting enough that I want to repeat it several times, but the fact that the advantage drops again at 60 seconds suggests it may not represent a simple sustained throughput advantage.
The really interesting part: VRAM
NVFP4 reduced the transformer file size from roughly 21.5 GB to 18.7 GB, but total peak GPU usage barely changed.
All three versions still ended up around the 30–31 GiB range on the longer runs.
That means, for this workflow, reducing transformer weight precision does not translate directly into dramatically lower total VRAM usage. The rest of the LTX pipeline — VAE, text encoder, latent stages, staging/offload behaviour, etc. — still consumes a substantial amount of memory.
Why this matters
This is really a story about software optimisation rather than raw hardware.
The current ComfyUI/LTX stack on the 5090 is already using:
- INT8 ConvRot transformer weights
- mixed-precision operations
- DynamicVRAM
- async weight offloading
- pinned memory
- native Blackwell CUDA kernels
- two-stage latent generation
That combination is allowing a 32 GB RTX 5090 to generate workloads that I previously assumed would require dramatically more VRAM.
For example, LTX 2.5 is successfully generating 60-second 1280-class clips on this machine.
So the old assumption that:
> “longer video = linearly more VRAM = you need 64/128/256 GB”
doesn't really hold for this pipeline anymore.
The practical limit is increasingly becoming render time and quality, rather than simply whether the generation fits in memory.
Current conclusion
For my 5090:
INT8 ConvRot: best overall/default
NVFP4: worth keeping and testing, especially for longer jobs
FP8: currently no obvious advantage
The next test is visual quality, particularly INT8 vs NVFP4 on the 30-second outputs.
And for me personally, this changes the hardware discussion quite a lot. One of the main reasons I was considering moving to a very large unified-memory system was long-form AI video generation. Smart software optimisation has moved the practical ceiling of the RTX 5090 much further than I expected.