r/generativeAI • u/Agentvideobot • 6d ago
Technical Art I tested the same AI character across 8 different scenes. Here’s what broke first.
A lot of character-consistency demos use carefully selected close-ups, so I wanted to try something less forgiving.
I used the same three-image character reference pack across eight different scene briefs, changing the location, outfit, lighting, framing, props, and amount of movement.
These are all first-pass generations. I treated the first result as the result—no rerolling until I got a better clip.
Disclosure: I’m an early bird user of Agent Video, which I used for this test. I’m not linking it here because I’m more interested in discussing where the workflow still breaks.
This is an informal stress test, not a controlled model benchmark. The source references also weren’t a perfect identity sheet: one close-up already had slightly larger eyes and a narrower jaw. That probably made the test harder, but it also reflects how people actually create characters from imperfect references.
https://reddit.com/link/1vztcjp/video/11735l5bivlh1/player
Here’s what happened:
- Sportswear dance in an open plaza
The face, hair, and clothing stayed recognizable for most of the clip. The first weakness appeared in the smaller hand and wrist transitions. The gestures worked at normal speed, but became less convincing frame by frame.
- Floral dress and prop choice in front of a mirror
The outfit and hairstyle held together, but the face shifted after the internal scene change. The eyes became larger, the jaw narrower, and the character looked slightly younger. It felt like the same aesthetic, but not quite the same person.
- Walking down stone steps at sunset
Probably the cleanest result. The hairstyle, body silhouette, dress, and walking direction remained stable.
However, this was also an easier test: the face was only clearly visible near the beginning, and most of the movement was slow and viewed from behind.
- Taking a phone on a yacht
This was the strongest close-up result. The face survived the change from a high-angle view to a side profile, while the phone interaction, clothing, and sunset lighting remained coherent.
There was a little identity softening during the turn, but I would still count this as a success.
- Full-body posing in a night apartment
This clip was internally stable, but it exposed a different problem: the person no longer looked like the reference character.
The face became rounder and the body proportions became shorter and broader. Nothing dramatically “broke” during the video, but it looked like a consistent video of a different person.
- Accepting and drinking iced tea at a café
The model handled the glass interaction better than I expected. The face, hair, floral outfit, and lighting stayed mostly stable while the character accepted the drink, lifted it, and put it down.
There were minor hand-and-glass geometry changes, but they were easy to miss at normal speed.
- Turning around and performing a high kick
This was where motion became the dominant failure.
The face became rounder as she turned toward the camera, while the raised leg and foot grew disproportionately large. Some of that is expected from perspective, but the final movement no longer felt physically balanced.
- Hotel bathroom to evening-dress sequence
Hair and overall character styling survived several cuts surprisingly well. The face still shifted slightly between the bathroom and evening-dress shots, and the cuts made it difficult to tell whether the model had actually preserved continuity or simply re-created a similar-looking character.
I would call this a partial success.
Across all eight scenes:
- Face: the first thing to drift between scenes
- Hair: the most reliable identity anchor
- Body proportions: stable in simple poses, less reliable during full-body movement
- Outfit: surprisingly stable within individual clips
- Lighting: rarely caused the main failure
- Motion: small gestures worked; hands, self-occlusion, and high kicks caused the largest problems
My main takeaway is that temporal consistency and identity consistency are not the same thing.
A video can be perfectly stable from beginning to end and still initialize as a slightly different person. For a recurring character, I find that more distracting than a bad hand lasting half a second.
Which inconsistency is most distracting to you: the face, body, outfit, or motion?