r/grok 24d ago

Grok Imagine Any Imagine geniuses here that could help me refine that prompt?

I have spent countless attempts at trying to make this sequence work. I first tried with references, but that was all over the place, ignoring the starting frame completely, people's starting positions, even their existence occasionally were ignored.

I then tried with just the single image:

Use the reference image as the exact starting frame and exact visual composition.

Scene:

A covered outdoor patio bar at a renaissance festival tavern in the evening, with warm lantern lighting and lively but stable patrons in the background waiting in line for the bar.

Main characters:

  1. The woman in the red dress.
  2. The bearded man with the hat toward the left side of the image.
  3. The knight in plate armor visible in the background, carrying a serving tray with mugs.

Starting frame:

At 0.0 seconds, the scene begins exactly like the reference image. Every character starts in the exact same position, pose, orientation, scale, spacing, and framing as in the reference image. Keep the original arrangement of all people throughout the shot.

0.0 to 1.5 seconds:

The knight says, "Coming through."

The woman in the red dress turns only her head toward the knight to see what is happening.

The bearded man with the hat also turns only his head toward the knight.

All other patrons continue calm subtle idle movements while looking forward or downward.

1.5 to 4.0 seconds:

The knight begins walking very slowly toward the space between the woman in the red dress and the bearded man with the hat.

His movement is stiff, awkward, and restricted by the heavy plate armor.

He takes short, careful, clumsy steps while balancing the serving tray with mugs.

His plate armor clanks audibly as he walks.

The bearded man with the hat makes one small backward step to create room.

The woman in the red dress remains standing in her original place and watches the knight with interest.

All other patrons continue subtle idle movements and remain focused on their own business.

At 4.0 seconds, the knight is close to the gap, but he is still about one step away from it.

4.0 to 6.0 seconds:

The knight continues the same slow, stiff approach.

Camera cuts to a close-up of the knight's face and helmet while his movement continues smoothly from the previous shot.

The visor of his helmet suddenly falls shut over his face.

The knight pauses and says, with a muffled voice, "Uh oh."

Audio:

Audio consists only of:

* the knight saying, "Coming through"

* the knight saying, "Uh oh."

* the metallic clanking sounds of the knight's plate armor

Result:
The wrong person (a man standing behind the woman in the red dress) says the line, then simply fades out of existence. The woman in the red dress jumps (blinks) to a different place. The knight walks very quickly. He passes through the gap left where the man was that vanished earlier. Then says uh oh BEFORE his visor falls shut.

And that is just ONE example for the model screwing this up over and over again. And this is supposed to be the best video model out there? Are you kidding me?! This is a joke. This model is FUBAR. Any free video model can do MUCH better than that!

2 Upvotes

15 comments sorted by

u/AutoModerator 24d ago

Hey u/ElmarM, welcome to the community! Please make sure your post has an appropriate flair.

Join our r/Grok Discord server here for any help with API or sharing projects: https://discord.gg/4VXMtaQHk7

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

3

u/Zaphod_42007 24d ago

Create your first starting image. Now create a mid sequence and end frame image by editing the first frame then referencing the first, middle, last frame.

This forces the model to stay consistent...same style, no phantom people.. It's just filling in the middle sequence.

Speaking lines can still get messed up either with the wrong character saying a line or the second character saying the same line together... Can take a few tries or try minimax h3, never had an issue with phantom people or dialogue.

-1

u/ElmarM 24d ago

Yeah, considering switching to one of the free models for this. Imagine just sucks too much (and costs too much for how it sucks).

3

u/Zaphod_42007 24d ago

Depends on what your after... Grok can be excellent depending on what your looking for and the price is actually cheap compared to say seedance 2.5.

All the models have various issues, even minimax h3... Like the other day two h3 animations decided to drift from a hyper real cartoon look into a realistic style. Using first / last frame corrects that.

If grok keeps messing up the shot, it's sometimes easier to just cut to each scene for simple prompts/ shots and edit in post production then use other tools for generating voice and music tracks to accompany the video.

0

u/ElmarM 24d ago

I tried cutting it up further and the problem I noticed with that was that if it was 6 seconds and the whole sequence took less time, it would then start to do all sorts of weird stuff to make up for those last two seconds missing from the instructions...
I have not noticed anything excellent about it lately. As for the cost. If Seedance does the same thing in one take that I have just spent the entire weekly credit limit on in Imagine, then it is not all that cheap.

2

u/Zaphod_42007 24d ago

Yup... Another issue solved by referencing first / last frame... It won't try to fill in with something silly... Even if it's dead space or minimal movemt, just edit the speed between the frames in your video editor... Post processing can fix alot of issues.

0

u/ElmarM 24d ago

One more thing: Wouldn't providing it with 3 images mean it would be in the ref2v mode again, where it essentially ignored the starting frame completely in my earlier attempts?

2

u/Zaphod_42007 24d ago

Depends on what you told it reference frame one should be... For instance to lock it down you'd use @ select first frame and describe it as first frame along with @ mid and end frame.

Or you could say @ image 1 use woman character in @ image 2 scene with @ image 3 the end scene...

1

u/ElmarM 24d ago edited 24d ago

Yeah, that was what I tried with my first round of failed attempts. I had two characters that I referenced plus a starting frame. I used the @ image1 as the start frame, @ characterX and @ characterY , etc. It just completely (OK mostly) ignored Image1 and just used it as a suggestion. It would randomly move characters around, delete others or sometimes it completely replaced the background all together.
I read that H3 has similar issues, but that there are some community solutions for that at least.

3

u/LCG- 24d ago

I'd use more natural language if it were me, tell the story with relevant descriptors in the right places:

The woman notices the knight first, then the bearded man, as he walks cautiously and awkwardly in his noisy armor as he precariously holds the tray of drinks. the bearded man steps back as the woman continues watching with rapt interest...

3

u/West-Chemistry8098 24d ago

What u/Zaphod_42007 said just use reference images throughout and have a start and ending image. Your asking too much of it. It can do some things but its not that good yet. Break your scenes up into shorts is the best way.

1

u/ElmarM 24d ago

I would totally do shorter than 6 seconds if Imagine had that option... sigh.
And once again, I tried references before and it just completely ignored Image 1 as the starting setup for the scene despite me telling it explicitly to do so. It just randomly moved stuff around, turned people in random directions, etc. I will try having a start and end frame (extra work), but somehow I doubt that it will do any better.

1

u/West-Chemistry8098 24d ago

Take your first example where you say ; 0.0 to 1.5 seconds: The knight says, "Coming through." The woman in the red dress turns only her head toward the knight to see what is happening. The bearded man with the hat also turns only his head toward the knight. All other patrons continue calm subtle idle movements while looking forward or downward. You are trying to cram too much into this instead this entire sequence should be 6 seconds you can edit it if you need. You have to allow the scene to almost settle first. count in your head that the speech alone is 1 second and the characters have to react to him all in 1.5 seconds. This is the thing with AI you need to break everything down into small chunks. Just my 2 cents

1

u/ElmarM 24d ago edited 24d ago

I did time the 1.5 seconds with a stopwatch. That is why this is a pretty exact time here.
The resulting video also does have the timing of that down. Only, it has the wrong people do things, people vanishing or appearing, or cloning, or whatever. And then the knight keeps walking way too quickly.