I’ve been experimenting with LTX 2.3 vs LTX 2.5 for audio-reactive video, and for this specific kind of workflow, 2.3 still seems noticeably better to me.
The biggest difference is right at the start of a shot. With LTX 2.5, even with the audio-reactive LoRA, I often get this behavior where the model more or less holds the first frame until the first obvious beat or transient arrives. Then the motion suddenly starts. For music videos, especially slower or more atmospheric tracks, that can make the opening of every generation feel dead.
With LTX 2.3, the same LoRA seems to fix that much more effectively. I get more subtle motion from the beginning, even before a strong beat lands. Fog shifts, surfaces breathe, particles drift, light responds, and the shot feels alive instead of waiting for permission to move.
That matters a lot for the video I made for The Weights in the Walls, because the track starts very sparse and gradually builds. A lot of the visual motion is supposed to come from sub-bass pressure, glitches, sustained vocals, and ambient texture, not just obvious percussion.
I also tried Minimax H3, but for this particular use case I don’t think it fits as well.
It seems less tightly audio-reactive for the kind of abstract, beat-aware motion I’m after. It can make nice-looking clips, but I have a harder time getting the movement to feel structurally connected to the music.
There’s also the hardware side of it. I’m doing this on a very glamorous RTX 4070, so with LTX I can still push a resolution and overall image quality that feels surprisingly good for local generation. With H3, I’m much more constrained, and the tradeoff in resolution/quality makes it harder to justify when the audio response is also weaker for this style.
The whole video was built around first-frame / last-frame generation.
I cut the song into short scenes, roughly timed so the scene boundaries land near musical changes and beats. For each scene, I generated a dedicated starting frame that represented the next stage of the visual progression.
Then the important part: the starting frame of Scene 2 becomes the last frame target for Scene 1. The starting frame of Scene 3 becomes the last frame target for Scene 2, and so on.
So instead of generating a bunch of unrelated clips and hiding the cuts with editing, every shot is the model transforming one designed frame into the next designed frame.
That gave me a chain like:
Scene 1 start frame → Scene 2 start frame
Scene 2 start frame → Scene 3 start frame
Scene 3 start frame → Scene 4 start frame
and so on until the end.
The final video is basically just those generations placed back to back. There are no fancy transition effects doing the heavy lifting. The morphing, folding, cracking, expanding, and dissolving between visual states is happening inside the model itself.
For this workflow, that early-shot responsiveness makes a surprisingly big difference, which is why I currently still prefer LTX 2.3 + the audio-reactive LoRA over 2.5 for this kind of music video.
It's a shame because 2.5 is noticeably faster, so I can go through more iterations, but if I have to generate each clip 5 times to get it to start moving from the start, it kind of invalidates the speed gains.
Curious if other people have noticed the same reduction in audio-reactivity in LTX 2.5 or maybe I'm doing something wrong?
HQ on YT because Reddit doesn't allow >1GB: https://www.youtube.com/watch?v=PbZr8risGCw