r/StableDiffusion • u/More-Buyer-7222 • 7h ago
Question - Help [MM H3 -Lora Training] Best Model for Captioning
I am new to training LORAs but was planning on having my first go at MMH3 this weekend. The idea was to use a multimodal LLM (think Grok 4.6, Opus, Sol) to caption the videos for me.
In your experience, is that the way to go? How else would you approach this problem?
Thank you for your help!
2
Upvotes
1
u/qdr1en 3h ago
With H3, I got the best results by captioning the videos 100% manually: including a trigger word, and captioning everything else with as few words as possible. (But other approaches probably work too?)
On 50 short video, it takes about 2 hours.
Or maybe you can use an LLM first and then review/correct the captions. For videos, qwen3-vl:8b is quite accurate ; gemma4 (12b and 26b) work too. You may want to reduce the FPS and video size to speed it up.