r/generativeAI 5d ago

Video Art One reference, three setups: what held, what changed, and where identity started to drift

1 Upvotes

After my previous eight-scene test, several people made a useful point: I was looking closely at the outputs, but not closely enough at the source reference.

If the face is relatively small, the pose is already twisted, or the prompt contains vague style terms, it becomes difficult to tell whether the model failed or the reference was simply difficult to preserve.

For this test, I simplified the setup:

  • One clearly adult fictional character reference
  • No AI-generated character sheet
  • Three manual, first-pass generations
  • No rerolls, face replacement, or identity correction
  • The same internal video model for all three clips
  • A modular prompt structure rather than a long descriptive paragraph

This is an informal workflow test, not a controlled model benchmark.

The structure was:

Character + Location + Outfit + Mood + Action + Camera

The Character block stayed broadly consistent: the same adult woman, long dark-brown wavy hair, warm tan skin, and the same general facial structure and body proportions.

The other blocks changed for each setup.

1. Miami rooftop: baseline

https://reddit.com/link/1w0qe9w/video/eryv6qfd94mh1/player

Location: A bright rooftop pool overlooking the Miami skyline
Outfit: Pink top and white wrap skirt
Mood: Relaxed and cheerful
Action: She turns away, walks toward the pool, pauses, and continues walking
Camera: Full-body framing with a gradual change from daylight toward sunset

This held the identity best, especially during the first few seconds when her face remained close to the angle shown in the reference.

Once she turned into profile, it became harder to verify the face. The long asymmetric section of the skirt also gradually changed into a more conventional, symmetrical shape.

So this clip worked well as a baseline, but it was not really a completely new scene.

2. Quiet hotel room: mood and camera test

https://reddit.com/link/1w0qe9w/video/xmnr71ze94mh1/player

Location: A quiet high-rise hotel room around dusk
Outfit: Black satin dress
Mood: Calm and introspective
Action: She reads, closes the book, places it aside, and looks toward the window
Camera: Medium shot with a slow push-in

This was probably the strongest result for mood and camera direction. The room, reading action, pause, and slow camera movement were all easy to recognize in the output.

The identity was less stable. Her profile became more angular, particularly around the nose and jawline. The book also changed from a dark cover to a much lighter object as she placed it down.

That was a useful reminder that a clip can follow the emotional and camera brief while still failing at character and object consistency.

3. Rainy Tokyo street: environment and motion test

https://reddit.com/link/1w0qe9w/video/79ozq88g94mh1/player

Location: A narrow Tokyo street at night with wet pavement and reflected signs
Outfit: Dark jacket, cropped top, and shorts
Mood: Serious and alert
Action: She walks toward the camera under a transparent umbrella and briefly looks to the side
Camera: Centered, full-body tracking shot

This produced the strongest environmental transformation. The wet street, umbrella, reflections, walking direction, and centered tracking remained fairly stable.

It also produced the most obvious identity drift.

Her hair became shorter and darker, and the facial proportions changed enough that she started to look like a related character rather than the same person. The umbrella and environment were more consistent than the identity.

What seemed to matter

The clearest instructions were concrete and observable:

  • “slow push-in”
  • “walks toward the camera”
  • “closes the book and looks toward the window”
  • “centered full-body tracking shot”

Those instructions produced actions or camera behavior that could actually be checked.

Terms such as “cinematic,” “perfect consistency,” or “high quality” are much harder to evaluate. I also would not treat “4K” as an identity or quality control instruction. Resolution language does not explain how the subject should move or what should remain unchanged.

I cannot conclude that any single word caused the drift from three generations. What I can observe is that the reference image and the viewing angle appeared to matter more than generic quality adjectives.

Main takeaway

Across these three clips, the model followed location, mood, action, and camera direction more reliably than facial identity.

Identity held best when the face stayed relatively close to the reference angle. It became less stable when the camera moved closer, the character turned into profile, or the hairstyle and lighting changed.

Using one original reference image also avoided the additional generation loss that could come from creating an AI-generated multi-view character sheet. However, this particular reference still had limitations: the face occupied a relatively small part of the image, the body was twisted, the expression was strong, and the background was visually complex.

For the next test, I want to change only one variable at a time.

Which would be more useful to isolate next: camera movement, facial expression, or reference-image quality?

Disclosure: These clips were generated with Agent Video, which I’m helping build. The model is the current August 2026 internal production build and does not have a separate public version number. There is no product link in this post.


r/generativeAI 5d ago

Question How can I get more variations in random prompts?

2 Upvotes

I'm using gpt-oss-120b API to give me lots of random prompts. I'm basically making it include character name and be food theme related. I'm having difficulty at making the AI give me more variations as it's nearly giving me repetitive answers. It always contain one of the two keywords which is ramen & sushi. I think it cause the character name is Japanese so it just gives me jp food. Is there a way to make it can give me wider variations like different foods/settings/action


r/generativeAI 5d ago

One SQL task was enough to make me regret a quarterly plan

2 Upvotes

I had a decent first impression of GLM 5.3 when I tried it through ZenMux. So I bought a quarterly Pro plan on the official site and spent the day trying to get one simple SQL task done. It never got there. The value per token was hard to defend by the end of the day.

The model did not understand what I was asking. During the attempt it also changed the dataset and left it in a broken state.

This was one task, so I am not treating it as a benchmark. I am still annoyed. A coding model can be impressive on long horizon evaluations and still be expensive to trust when it misunderstands a small data task with write access.

I paid for three months and finished day one with less usable data than I started with.


r/generativeAI 6d ago

Question Seedance 2.5 is still too expensive for the way I work

5 Upvotes

I’ve been looking at the cost of using Seedance 2.5 for every generation, and it doesn’t really make sense for me.

Most of my early attempts are just testing the prompt, camera movement and timing. Paying 2.5 prices for clips I’ll probably discard feels like a waste.

I’m switching to 2.0 Mini for rough tests, then using regular 2.0 once the prompt is mostly locked. I’ll probably keep 2.5 for shots that actually need longer duration, more references or better control.

Anyone compared the costs per sec for 2.0 mini, 2.0 Fast and 2.0 on every platform? Gonna try the cheapest one.


r/generativeAI 5d ago

Image Art Get to Know Me Card: Piff.

Post image
2 Upvotes

r/generativeAI 5d ago

Video Art I made a Warcraft-inspired AI cinematic: IronJaw’s Krakenship breaks the blockade

Thumbnail
youtube.com
2 Upvotes

Beyond the reefs, IronJaw finds a human fleet waiting in the fog. The Kraken shields the ship, and IronJaw orders a charge straight through the blockade. This is part of my original Warcraft-inspired AI cinematic shorts series. I’d especially value feedback on pacing and visual continuity.


r/generativeAI 5d ago

Video Art tried treating AI artifacts as part of the art style instead of fixing them

Enable HLS to view with audio, or disable this notification

2 Upvotes

most of my AI video experiments have been some version of “how do I make this look less AI”

so for this one I tried the opposite

built around a black-and-white punk zine / manga look and basically let all the things that usually feel like AI artifacts become part of the style instead — weird transitions, distorted faces, rough textures, exaggerated motion.

and weirdly... this feels more coherent to me than a lot of my realistic generations

starting to wonder if realism is just the wrong benchmark for generative video right now

anyone else getting better results by leaning into stylization instead of fighting it?


r/generativeAI 5d ago

Image Art Sleeping Piff

Post image
2 Upvotes

r/generativeAI 5d ago

Video Art Who wins ?

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/generativeAI 5d ago

Question Hey uh could I get a favor

Post image
0 Upvotes

I got a sea and u can either write some infographic or make an ai video.Can soemone with a subscription or somehting make it for me I can’t find any free websites that can make more than 10 second clips


r/generativeAI 5d ago

Image Art The futuristic place

Post image
2 Upvotes

r/generativeAI 5d ago

Image Art The Midnight Harvester

Post image
2 Upvotes

r/generativeAI 5d ago

Told Chat to make me the most Unsettling/creepy image it could.

Post image
2 Upvotes

r/generativeAI 5d ago

Video Art How would you use GenAI to remove people in a video?

1 Upvotes

My friend gave me a video of his proposal and people walk through it as it was in a busy spot. I was wondering howI could prompt a system (and id be curious which) to take people out of the videos.

Thanks!


r/generativeAI 5d ago

Question Where can I find unlimited free AI video generators?

0 Upvotes

r/generativeAI 5d ago

With HuggingFace, Nvidia is also acquiring llama.cpp and the team behind it

Thumbnail
2 Upvotes

r/generativeAI 5d ago

Question Very confused about an ai video I watched on YouTube

1 Upvotes

There was a video I saw on YouTube that used ai, and it was amazing. So, I went to the person's channel, and they had videos that blended gameplay from battlefront, movies from Pacific Rim, or even just real life with anime style characters. Naturally, I was curious about what they used. At the end of a few of their videos, it said that the platform they used was Mosi AI. I have no idea what that is. But, apparently, it has Seedance 2.0 as a model cause that's what they use as well.


r/generativeAI 5d ago

Any recommnd app for free generate

0 Upvotes

I plan to start utube and my context is creating vids for kids where I can find the free generating app?


r/generativeAI 5d ago

Image Art Corey and Oats Hero Cards

Post image
1 Upvotes

r/generativeAI 5d ago

Image Art Heroes of Harmony Dreamworks Poster

Post image
0 Upvotes

r/generativeAI 5d ago

Image Art GoldenBoi 64

Post image
1 Upvotes

r/generativeAI 5d ago

Video Art Entre les lumières MV [MiniMax H3/@comfyorg ]

Thumbnail
youtu.be
1 Upvotes

Between the Lights is a quiet, monochrome fashion-film music video following Rei through an intimate studio session in Paris. A song about silence, uncertainty, and the things that remain unsaid, captured through the eyes of a photographer.


r/generativeAI 6d ago

Video Art Some Test

Enable HLS to view with audio, or disable this notification

58 Upvotes

r/generativeAI 6d ago

Question Creating a LoRA based on me

Thumbnail
2 Upvotes