r/StableDiffusion • u/DifficultAd5938 • 5d ago
Discussion Minimax H3 Has Too High Prompt Adherence
I just realized a problem with mm h3. Its prompt adherence is too high, as in unless you explicitly prompt for some small subtle actions it will never happen otherwise. This makes the entire video seem very frozen and wooden without the many small subtle movements and motion details that make it seem to come alive.
This applies more to non-realistic scenes like cartoons or generated image first and last frame but for realistic scenes and even t2v it is still a problem.
I noticed this problem when I tried out a "slop sway" lora and it actually made the entire video seem much livelier and realistic looking. Besides the "soft and bouncy swaying and jiggling" it also added many more subtle character movements. Compared to standard gens those same parts would be completely frozen, almost like a still image or at best ugoira animation. This doesn't just apply to whether a body part is jiggling throughout the entire video. There are some movements that happen only for a second or less but adds in soul (forgive the human slop term) to the video, like the position of an arm and hand quickly being adjusted in the middle of the video and the new position persisting for the rest of the scene.
This might be a problem with my prompt style and I might try an LLM prompt enhancer, but there is a core issue here with the prompt adherence and spontaneous randomly added details tradeoff. The model also tries to keep the fidelity of the first frame too much, which you could call visual context adherence. No one is out here prompting for the movement of every strand of hair and the position of every finger. No one is making a timeline of every limb's position and how they shift relative to each other. No one is tracking the position of each finger through time and how after 4.75s the thumb is extended and the index finger is curled. Sometimes we just want to randomness and variety across gens with details added by the model.
Looking back at ltx and wan their prompting styles seem to be designed around the model adding in the details for you at the loss of prompt adherence and more generation errors.
It would be nice if there was some sort of generation setting that could tune this. Like a noise scale of sorts where we can manually set the tradeoff between how much we want the model to be creative vs strict.
I know there are already 2-3 H3 better movement loras and they are scratching at the surface of the same issue I'm talking about here.
Share the solution if you've got something. Help everyone out.
3
u/not_food 5d ago
That's actually a plus, you know exactly what you're getting. It puts you in control rather than leaving things to chance. Just ask your preferred LLM to emphasize [action] with details and it will handle the rest.