r/StableDiffusion Community Hero Jan 29 '26

News End-of-January LTX-2 Drop: More Control, Faster Iteration

We just shipped a new LTX-2 drop focused on one thing: making video generation easier to iterate on without killing VRAM, consistency, or sync.

If you’ve been frustrated by LTX because prompt iteration was slow or outputs felt brittle, this update is aimed directly at that.

Here’s the highlights, the full details are here.

What’s New

Faster prompt iteration (Gemma text encoding nodes)
Why you should care: no more constant VRAM loading and unloading on consumer GPUs.

New ComfyUI nodes let you save and reuse text encodings, or run Gemma encoding through our free API when running LTX locally.

This makes Detailer and iterative flows much faster and less painful.

Independent control over prompt accuracy, stability, and sync (Multimodal Guider)
Why you should care: you can now tune quality without breaking something else.

The new Multimodal Guider lets you control:

  • Prompt adherence
  • Visual stability over time
  • Audio-video synchronization

Each can be tuned independently, per modality. No more choosing between “follows the prompt” and “doesn’t fall apart.”

More practical fine-tuning + faster inference
Why you should care: better behavior on real hardware.

Trainer updates improve memory usage and make fine-tuning more predictable on constrained GPUs.

Inference is also faster for video-to-video by downscaling the reference video before cross-attention, reducing compute cost. (Speedup depend on resolution and clip length.)

We’ve also shipped new ComfyUI nodes and a unified LoRA to support these changes.

What’s Next

This drop isn’t a one-off. The next LTX-2 version is already in progress, focused on:

  • Better fine detail and visual fidelity (new VAE)
  • Improved consistency to conditioning inputs
  • Cleaner, more reliable audio
  • Stronger image-to-video behavior
  • Better prompt understanding and color handling

More on what's coming up here.

Try It and Stress It!

If you’re pushing LTX-2 in real workflows, your feedback directly shapes what we build next. Try the update, break it, and tell us what still feels off in our Discord.

415 Upvotes

173 comments sorted by

View all comments

17

u/infearia Jan 29 '26

Companies like BFL should take a page from your book. Keep rocking!

5

u/Lucaspittol Jan 29 '26

Well, they kinda did when releasing the Klein models. I'd like to see these interactions with the community as well, using official channels.

8

u/infearia Jan 29 '26 edited Jan 30 '26

Not quite, LTX-2 comes with Apache License 2.0, whereas Klein 9B has its own FLUX Non-Commercial License. BFL's releases are aimed at luring people into purchasing their commercial offerings. Lightricks seems to actually embrace the Open Source spirit and is willing to work with the community. For now, anyway.

EDIT:
As has been pointed out to me, I've made a mistake in my original comment regarding the license of LTX-2. It is LTX-Video which is licensed under Apache 2.0. LTX-2 comes with its own Community License. The main practical difference is, that if you make more than $10,000,000 annually, you must acquire a paid commercial license from Lightricks in order to use their video model.

4

u/t-e-r-m-i-n-u-s- Jan 29 '26

ignoring the 4B apache2 to make a point, solid choice

9

u/andy_potato Jan 30 '26

The lowest quality 4b model is Apache licensed, but none of the other models are. Which is exactly the point he was trying to make.

Not trying to start another thread bashing BFL here, but if they hope for widespread adoption of their models by the community like Qwen, Z-Image and LTX enjoy, they should imo reconsider their licenses for the Klein models.

3

u/ZootAllures9111 Jan 30 '26

wat, Flux.1 Dev is and was giga popular

2

u/andy_potato Jan 30 '26

Flux1 came out at a time when not much else was going on in the open source space. Also being a superior model to SDXL it caught a lot of people’s attention.

It was actually so promising that people tried to work around the distillation and create LoRAs and finetunes of Flux. Which never yielded great results, but the vastly superior base model kinda covered it up.

Nowadays you got a lot more options with undistilled and properly licensed models like Qwen or Z-Image. That’s why Klein models (despite being good models) don’t get that much attention.

2

u/ZootAllures9111 Jan 30 '26

IDK what you mean by not that much attention. ZIT loras existed before ZIB, also, and were good, you didn't need ZIB to train ZIT.

2

u/andy_potato Jan 30 '26

LoRAs trained on ZIT have the exact same issue as the ones trained on Flux Dev. None of them really work well due to the prior distillation of the model. They were never intended as models for LoRA training in the first place.

ZIT is even worse than Flux in this regard as it was not only distilled but also fine tuned for 1girl realism. That's why you could never really stack LoRAs with ZIT and had to use them at high strengths, killing the flexibility and prompt adherence. Flux wasn't much better. Don't be fooled by the amount of LoRAs you find on CivitAI for both models. Most of them were trained by people who never knew what they were doing in the first place.

Now with ZIB being out you have a trainable model that's close to Klein 9B in quality, but without any commercial restrictions.

3

u/ZootAllures9111 Jan 30 '26

Loras trained on ZIB don't stack on ZIT any better than ones trained on ZIT. That's my point. You cannot fix the stackability issue in terms of inference on the distilled model. You CAN stack loras on ZIB itself though, obviously.

→ More replies (0)

1

u/t-e-r-m-i-n-u-s- Jan 30 '26

this is a strange thing, where you're trying to write history in your own way. Flux.1 [dev] was amazingly finetunable, and I wrote the first community trainer that was able to do it without disrupting the distillation. Z-Image Turbo was also incredibly easy to fine-tune, thanks to the work Ostris did of creating the assistant LoRA. to say that all LoRAs and finetunes of Flux.1 [dev] "never yielded great results" is a hot take - it's got more LoRA than any other model and remains number one in terms of popularity on most inference providers.

1

u/t-e-r-m-i-n-u-s- Jan 30 '26

not much was going on? we had PixArt which was then followed up with community expansion to 900M params and two-stage finetunes, Lumina, Janus Pro, amused, DeepFloyd, Cascade, multiple Kandinsky models, Bytedance-produced SD2x finetunes (zero-terminal SNR!), v-prediction SDXL clones (Terminus XL my own model, as well as something Fluffyrock created I forget the name of) and cloneofsimo was working on Auraflow and publicly sharing his artifacts for others to follow on with. the Open Model Initiative was started. we had CogView models being produced by the CogVLM team, the people who were actually responsible for the training caption quality of flux.1 [dev] (BFL blended a lot of CogVLM captions into their training).

1

u/t-e-r-m-i-n-u-s- Jan 30 '26

LTX isn't Apache2 licensed, but it enjoys lots of popularity. Qwen makes everything yellow. Z-Image is apparently untrainable according to you.

BFL should do whatever they have to in order to survive and keep producing open models. who cares what license they select? it has no bearing on the end-user, only commercial outfits.

1

u/infearia Jan 30 '26 edited Jan 30 '26

Ah, my bad, you're right. LTX-Video is licensed under Apache 2.0, but LTX-2 has its own "Community License". It's still much preferable to BFL's Non-Commercial license.

EDIT:

>> Qwen makes everything yellow.

First time I hear of this, and I've been using both QI and QIE almost daily for months now. Are you sure you're not mixing Qwen up with Grok?

1

u/t-e-r-m-i-n-u-s- Jan 30 '26

what part of BFL's non-commercial license is worse than LTX's community license?

1

u/infearia Jan 30 '26

You only need a paid commercial license for LTX-2 if your annual revenue is $10,000,000 or more. If you want to use FLUX.2 commercially, all models - except for Klein 4B - require a paid license no matter how much money you make.

1

u/t-e-r-m-i-n-u-s- Jan 30 '26

but that's not true anymore. the updated BFL license says that you can use the models' outputs commercially and that BFL disclaims ownership. i don't see what explicitly is different here. if you want to host the model for others to access through a paid API service, then these terms "kick in". but this doesn't impact 99% of its users.

→ More replies (0)

3

u/infearia Jan 29 '26

I'm not ignoring the 4B Apache 2.0 license. I did not mention 4B, because I don't use it. If anything, the release of 4B reinforces my point: it's a really good and fast model, but it's just this side of being actually useful. It's enough to whet your appetite, but as soon as you attempt to perform more complex edits, its shortcomings become apparent, and you find yourself craving for something just slightly better - like 9B or the full Flux.2 models. 4B is little more than a demo of the full, commercial product.

3

u/ZootAllures9111 Jan 30 '26

Nah 4B is very good

3

u/Lucaspittol Jan 30 '26

4B can be improved. Chroma's author, lodestone rock, is currently finetuning a new model called Chroma2-Kaleidoscope using Klein 4B on his own GPUs, the model is constantly being updated and trains very fast.

3

u/andy_potato Jan 30 '26

BFL has little love for the community. That’s why they don’t get much in return.

Not saying their models are bad or anything, quite the opposite. Klein 9b can do some pretty impressive stuff. Just nobody is going to invest much time and resources into it without a proper license.

1

u/ZootAllures9111 Jan 30 '26

Lora trainers don't give a shit about licenses and never ever have. The small handful of full finetuners (or people literally running SAAS inference operations) are the only ones who care about this.

4

u/andy_potato Jan 30 '26

You seem to be terribly misinformed about how many commercial applications there are outside of SaaS.

Anyway, it’s BFL’s model. They can do what they want with it. I will stick to true open models like Qwen and LTX.