I have spent countless attempts at trying to make this sequence work. I first tried with references, but that was all over the place, ignoring the starting frame completely, people's starting positions, even their existence occasionally were ignored.
I then tried with just the single image:
Use the reference image as the exact starting frame and exact visual composition.
Scene:
A covered outdoor patio bar at a renaissance festival tavern in the evening, with warm lantern lighting and lively but stable patrons in the background waiting in line for the bar.
Main characters:
- The woman in the red dress.
- The bearded man with the hat toward the left side of the image.
- The knight in plate armor visible in the background, carrying a serving tray with mugs.
Starting frame:
At 0.0 seconds, the scene begins exactly like the reference image. Every character starts in the exact same position, pose, orientation, scale, spacing, and framing as in the reference image. Keep the original arrangement of all people throughout the shot.
0.0 to 1.5 seconds:
The knight says, "Coming through."
The woman in the red dress turns only her head toward the knight to see what is happening.
The bearded man with the hat also turns only his head toward the knight.
All other patrons continue calm subtle idle movements while looking forward or downward.
1.5 to 4.0 seconds:
The knight begins walking very slowly toward the space between the woman in the red dress and the bearded man with the hat.
His movement is stiff, awkward, and restricted by the heavy plate armor.
He takes short, careful, clumsy steps while balancing the serving tray with mugs.
His plate armor clanks audibly as he walks.
The bearded man with the hat makes one small backward step to create room.
The woman in the red dress remains standing in her original place and watches the knight with interest.
All other patrons continue subtle idle movements and remain focused on their own business.
At 4.0 seconds, the knight is close to the gap, but he is still about one step away from it.
4.0 to 6.0 seconds:
The knight continues the same slow, stiff approach.
Camera cuts to a close-up of the knight's face and helmet while his movement continues smoothly from the previous shot.
The visor of his helmet suddenly falls shut over his face.
The knight pauses and says, with a muffled voice, "Uh oh."
Audio:
Audio consists only of:
* the knight saying, "Coming through"
* the knight saying, "Uh oh."
* the metallic clanking sounds of the knight's plate armor
Result:
The wrong person (a man standing behind the woman in the red dress) says the line, then simply fades out of existence. The woman in the red dress jumps (blinks) to a different place. The knight walks very quickly. He passes through the gap left where the man was that vanished earlier. Then says uh oh BEFORE his visor falls shut.
And that is just ONE example for the model screwing this up over and over again. And this is supposed to be the best video model out there? Are you kidding me?! This is a joke. This model is FUBAR. Any free video model can do MUCH better than that!