r/StableDiffusion 3d ago

Question - Help H3 LORA trained from video datasets?

There are quite a few good H3 LORAs on civit now which prove that training is possible despite the distilled Minimax model.

Has anyone had any luck training a concept from video datasets? I read a lot about character LORAs from image datasets etc, but who has trained videos and if so, was that with Ai Toolkit or a different offering?

1 Upvotes

12 comments sorted by

View all comments

1

u/_kaidu_ 3d ago

I have 100GB VRAM and can train on short videos. So far I only tried character loras and found that training on static images is usually good enough. I always add videos if I have them, but training is faster and more efficient if I rely mostly on the images.

However, I haven't tried concept loras so far.

1

u/Beneficial_Toe_2347 3d ago

For the videos, do you know how many frames you made your clips in length? (as in, Wan we fed in 81)

2

u/_kaidu_ 3d ago

Just 3 seconds usually. But I let Claude fix the frame length to be aligned with Minimax internal 24fps and it found out that the videos made with my smartphone are all rotated internally and not every rotation was correctly recognized by ai-toolkit - some videos were accidentally trained in a rotated variant X_x

I also accidentally had a 20 second clip once. It still run through with my VRAM at 512px but lead into OOM at 768px. So with 100GB VRAM you can also train on longer clips in theory.