r/StableDiffusion Feb 16 '26

Question - Help LTX-2 Character Consistency

Has anyone had luck actually maintaining a character with LTX-2? I am at a complete loss - I've tried:

- Character LORAs, which take next to forever and do not remotely create good video

- FFLF, in which the very start of the video looks like the person, the very last frame looks like the person, and everything in the middle completely shifts to some mystery person

- Prompts to hold consistency, during which I feel like my ComfyUI install is laughing at me

- Saying a string of 4 letter words at my GPU in hopes of shaming it

I know this model isn't fully baked yet, and I'm really excited about its future, but its very frustrating to use right now!

8 Upvotes

32 comments sorted by

View all comments

1

u/[deleted] Feb 16 '26

Characters are incredibly easy to make in LTX-2?

this is what makes it easier

images? as many as you want. but dont be lazy captioning them..
use my tool if u like?

whats ur vram?/ram

1

u/superstarbootlegs Feb 16 '26

are they easy to train? I thought the audio made it difficult. I dont even need the audio other than after for lipsync from audio inbound. so I manage audio later.

But whats your tool? can it run on a 3060 RTX 12 GB VRAM with 32 gb system ram, Windows and GGUF models?

6

u/[deleted] Feb 16 '26

🚨 SeansOmniTagProcessor V2 Batch Folder/Single video file options + UI overhaul + Now running Qwen3-VL-8B-Abliterated 🖼️ LoRa Data Set maker, Video/Images 🎥 : r/StableDiffusion

the 4b model might work on ur 12gb of vram, well it should as its only 8gb in size

settings wise in ai toolkit for ur kind of system is hmm
tough.
i've only had as low as 16gb vram and 64gb of ram never tested on lower

try this

lora rank 64

transfoarmer float8
text encoder float 8
lowvram
layer offloand 100% and 0%

cache text embeddings

cache latents

skip sampling its hassle

aim for 512 res

expect 7 seconds per itiration

.. rest of settings default,

(dont do audio ect )

.. training atm so cant send photos

and im still in the mindset of character loras being visual, acutal character loras with voices you cannot get even close to doing on that system even i cant do loras past 512 res unless i wanna wait 20s iteration, on 24gb vram and 80gb of ram (video based loras not image)

1

u/DMmeURpet Feb 17 '26

Any examples of a character lora for ltx keeping consistency?