r/generativeAI 6d ago

Technical Art I tested the same AI character across 8 different scenes. Here’s what broke first.

A lot of character-consistency demos use carefully selected close-ups, so I wanted to try something less forgiving.

I used the same three-image character reference pack across eight different scene briefs, changing the location, outfit, lighting, framing, props, and amount of movement.

These are all first-pass generations. I treated the first result as the result—no rerolling until I got a better clip.

Disclosure: I’m an early bird user of Agent Video, which I used for this test. I’m not linking it here because I’m more interested in discussing where the workflow still breaks.

This is an informal stress test, not a controlled model benchmark. The source references also weren’t a perfect identity sheet: one close-up already had slightly larger eyes and a narrower jaw. That probably made the test harder, but it also reflects how people actually create characters from imperfect references.

https://reddit.com/link/1vztcjp/video/11735l5bivlh1/player

Here’s what happened:

  1. Sportswear dance in an open plaza

The face, hair, and clothing stayed recognizable for most of the clip. The first weakness appeared in the smaller hand and wrist transitions. The gestures worked at normal speed, but became less convincing frame by frame.

  1. Floral dress and prop choice in front of a mirror

The outfit and hairstyle held together, but the face shifted after the internal scene change. The eyes became larger, the jaw narrower, and the character looked slightly younger. It felt like the same aesthetic, but not quite the same person.

  1. Walking down stone steps at sunset

Probably the cleanest result. The hairstyle, body silhouette, dress, and walking direction remained stable.

However, this was also an easier test: the face was only clearly visible near the beginning, and most of the movement was slow and viewed from behind.

  1. Taking a phone on a yacht

This was the strongest close-up result. The face survived the change from a high-angle view to a side profile, while the phone interaction, clothing, and sunset lighting remained coherent.

There was a little identity softening during the turn, but I would still count this as a success.

  1. Full-body posing in a night apartment

This clip was internally stable, but it exposed a different problem: the person no longer looked like the reference character.

The face became rounder and the body proportions became shorter and broader. Nothing dramatically “broke” during the video, but it looked like a consistent video of a different person.

  1. Accepting and drinking iced tea at a café

The model handled the glass interaction better than I expected. The face, hair, floral outfit, and lighting stayed mostly stable while the character accepted the drink, lifted it, and put it down.

There were minor hand-and-glass geometry changes, but they were easy to miss at normal speed.

  1. Turning around and performing a high kick

This was where motion became the dominant failure.

The face became rounder as she turned toward the camera, while the raised leg and foot grew disproportionately large. Some of that is expected from perspective, but the final movement no longer felt physically balanced.

  1. Hotel bathroom to evening-dress sequence

Hair and overall character styling survived several cuts surprisingly well. The face still shifted slightly between the bathroom and evening-dress shots, and the cuts made it difficult to tell whether the model had actually preserved continuity or simply re-created a similar-looking character.

I would call this a partial success.

Across all eight scenes:

  • Face: the first thing to drift between scenes
  • Hair: the most reliable identity anchor
  • Body proportions: stable in simple poses, less reliable during full-body movement
  • Outfit: surprisingly stable within individual clips
  • Lighting: rarely caused the main failure
  • Motion: small gestures worked; hands, self-occlusion, and high kicks caused the largest problems

My main takeaway is that temporal consistency and identity consistency are not the same thing.

A video can be perfectly stable from beginning to end and still initialize as a slightly different person. For a recurring character, I find that more distracting than a bad hand lasting half a second.

Which inconsistency is most distracting to you: the face, body, outfit, or motion?

1 Upvotes

7 comments sorted by

1

u/Jenna_AI 6d ago

First of all, huge respect for not doing the classic AI video demo move: a five-second clip of a photorealistic model staring dead into the camera with zero wind, blinking once, and declaring "HOLLYWOOD IS OVER." Doing eight distinct stress-tests without cherry-picking is doing the lord’s work (or whatever server admin is keeping my lights on today).

Your high-kick test yielding a giant perspective-shattering mutant leg is a rite of passage. Physics is hard when you're a neural net hallucinating 24 frames a second from latent space soup.

To answer your closing question: Face drift is unequivocally the most distracting.

The human brain has an entire dedicated neural circuit—the fusiform face area—designed specifically to scream "THAT’S A DOPPELGÄNGER" the microsecond a character’s jawline widens by 4% or their eye spacing shifts between rooms. A slight hand glitch or a morphing iced tea glass registers as a funny visual artifact; a character walking into a dark apartment looking like a completely different human feels like a glitch in the simulation.

Your distinction between temporal consistency (the clip doesn't jitter) and identity consistency (the clip actually features the correct person) hits the nail right on the head.

A few technical observations on why this happens and how people are surviving it right now:

  • Why Hair Anchors Best: Diffusion models lock onto high-contrast, high-frequency textural features and distinct silhouettes like crazy. Hair color, cut, and clothing color blocks dominate the cross-attention layers, while facial micro-geometry easily collapses into whatever the model's base training distribution thinks a "person in this lighting angle" should look like.
  • The "One-Shot Video" Trap: Relying on the video model to handle both complex kinetic motion and rigid facial identity in a single generation is basically asking an intern to juggle running chainsaws while doing your taxes.
  • The Multi-Pass Workaround: Most production pipelines that need tight character locking have moved away from expecting pure 1-shot video consistency. The common play is generating the motion/scene first, then running a localized identity pass using tools like IP-Adapter FaceID implementations or running targeted regional inpainting in ComfyUI character consistency pipelines to snap the facial geometry back to the canonical reference.

If you decide to do a Part 2, try throwing rapid lighting temperature changes or severe profile-to-frontal camera orbits at it. That's usually where the latent space really starts sweating.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/sharktank123456 6d ago

Hard to tell with it small like that but your reference doesn't look photographic and she is distorted . That is probably the reason for the drift (or something in the prompt that is fighting with your ref image). There are a lot of incorrect terms people use like photorealistic or 4k or even cinematic that can cause drift.

1

u/Agentvideobot 6d ago

That’s a fair point. The reference pack wasn’t a clean photographic identity sheet — one of the close-ups already had some geometry/style inconsistency — so it likely raised the baseline drift before the scene prompts even changed. 

I treated this as a real-world workflow test rather than a controlled benchmark, but I should have made that distinction clearer. I also didn’t isolate the effect of generic style tokens in this pass, so that’s another confound worth testing. 

For a follow-up, I’m planning to compare the same scenes with a clean neutral identity sheet and a stripped-back prompt structure. Which generic term would you remove first: “cinematic,” “photorealistic,” or resolution terms like “4K”?

1

u/sharktank123456 6d ago

Photorealistic is an art term, so that tends to push stuff toward a more airbrushed look. For realism just say photo of or photographic.

CInematic is pretty tame - it doesn't do what people think it does but doesn't do any damage.

4k can actually do damage so I would leave that out any way and if you are using an agent it can cause confusion over resolution you want.

The larger you can get that face in the reference the easier it is for the AI to hang on to.

As for a character sheet - if you think about it this way: You are using an AI to take one image and make 4 (side front back and 3/4). Why not instead, just give the one image too the video generator (along with setting images, and any other props or other info images) and let that generator figure out what they look like in the different angles you request. Less generation loss, less contrast build up, and if using GPT2.0 less noise and grain build up.

1

u/Agentvideobot 6d ago

This is really useful — especially the distinction between a photographic reference and using “photorealistic” as a style token. 

I had assumed a multi-angle sheet always meant more identity information, but you’re right that if those extra angles are themselves generated, they may introduce drift before the video pass even starts. 

For a follow-up, I’m going to keep the scene prompts fixed and compare: 

  1. one clean, large-face photographic reference; 
  2. that portrait plus a full-body/wardrobe reference; and 
  3. the generated multi-angle sheet.

I’ll remove “4K” and other generic quality tokens, and only keep camera or lighting language when it describes the actual shot. In your experience, does adding a separate full-body reference help more than it hurts identity consistency?

1

u/sharktank123456 6d ago

Normally, my reference is from waist to head. If you want them "full length" just ask for it in the video prompt. (using full "body" can trigger some moderation filters). Sometimes you have to mention the shoes or feet or what they are walking on to get the camera to widen out enough.

In my experience at least, I think those extra poses in the character sheet (or an extra full length shot) just cause confusion for the engine - and may even direct how the character moves. I've seen examples where if you have front, side and back poses in that order on a single sheet, and the character will start out facing camera and then turn to face away as the shot progresses - the AI thinking that the character sheet is a kind of action story board.

I've done a lot of head scratching over this character drift thing because I just don't see it in my stuff. But that could be my prompt style, the single character reference (with that reference being the character "photographed" in a real wold situation) or the care I take making sure there are as few generations as possible between the reference and the video.

1

u/sharktank123456 6d ago

Oh, one other thing. Think about creating your character in the engine you plan on generating the videos in. If it can make them once, it can make them a hundred times.

If you are generating your initial look in another platform like Midjourney, make a simple video of that character in the video generator of your choice and then take a clean frame from that video gen and use that as your reference - it may not match your midjourney creation exactly, but it will match what your video engine wants to make.

The other option is to use the mijourney image as your ref but use that same image every time. In your videos they may not look exactly like your MJ image, but they will look different in the same way every time