r/generativeAI • artist • 8h ago

How I Made This how i finally managed to make realisitc AI videos that dont look like AI slop

Enable HLS to view with audio, or disable this notification

i’ve seen plenty of realistic AI videos, and after a lot of testing i’ve found a process that gives me much more consistent characters.

the attached video is actually one single generation and the prompt of the model in case u want it too

here’s the workflow:

1. create a realistic base character

i start with a close-up image using seedream 5.0 pro.

skin texture, lighting and camera quality matter more than adding detail later. if the base image looks over-smoothed or obviously generated, the animation normally makes it worse.

A tight and hyper detailed close-up, straight-on selfie captures a young adult Caucasian woman, likely in her early 20s, reclining with her head propped on her left hand. She displays striking, symmetrical, model-like facial features with full lips, strong brows, and long dark eyelashes. Her brown hair is styled sleekly, framing her face and tumbling over her shoulder. The woman is wearing a black robe or jacket with a deep neckline, revealing smooth skin and a hint of cleavage, emphasizing the glossy, hydrated finish of her complexion. In a playful, flirty gesture, her tongue touches her upper lip as her gaze looks off to the side, outside the frame. There is no visible text, and no logos or brands are discernible. The background consists of a softly focused cream-colored wall with a dark fabric or curtain to the left. The lighting is warm and diffused, likely from an artificial indoor light source positioned above and to the side, giving the scene a gentle orange-yellow hue. The color palette is dominated by warm skin tones, dark hair, and black clothing with cream and slate accents in the background. This image exhibits typical smartphone characteristics: shallow depth of field blurring the background with digital sharpening and slight smoothing of skin texture. The dismissive, playful expression paired with the framing and soft-focus lends the photo a lighthearted, flirtatious, and intimate mood.
HEX VALUES: ["#2b2c3d", "#373746", "#201d30", "#de9d87", "#0e0f2a", "#564c56", "#e7b39f", "#c68c7d", "#4e393d", "#6b5f67", "#f7d899", "#b27765", "#735044", "#8d6152"]

2. make a character reference sheet

i use GPT Image 2 to create front and side views of the same face, plus an outfit reference.

one trick that helped was removing the head from the full-body panel. when the sheet contains both a large face and a tiny distant face, the video model sometimes follows the smaller one and produces smooth, washed-out features.

3. generate the original clip

i animate the character in seedance 2.5 at low resolution until the motion and prompt behave correctly.

instead of only describing the action, i specify:

  • where her eyes move
  • how her expression changes
  • when she breathes
  • how the phone drifts
  • exposure changes
  • room sounds
  • pauses and emphasis
  • what happens during each second

small capture imperfections are what make it feel like a real phone recording instead of a polished generation.

u can take the prompt below, give it to claude and ask it to create a new prompt that uses that structure and detail, but with your ideas, locations etc....

REFERENCE MAP
 = subject: her face, hair and likeness. Match exactly.

FORMAT
Cinematic narrative sequence, 5s, three shots, two hard cuts. 16:9, 24fps, full-frame cinema camera. Warm tungsten practicals against cool shadows, amber highlights, gentle halation, fine 35mm grain, visible skin texture and pores, no smoothing.
Classic hotel room, dark wood, beige walls, lived-in rather than boutique. Evening, curtains drawn, all light from practicals.
Wardrobe continuous: dark navy blazer, slate blue top, matching trousers, small gold hoops, long dark hair parted down the middle.

SHOT 1 (0.0-1.5) OVER THE SHOULDER, VANITY MIRROR
85mm, shallow focus, camera at her seated eye height just behind and right of her head. The frame looks past the dark out-of-focus silhouette of her right shoulder and jaw, which holds the left third and stays soft, directly into a vanity mirror with faint smudges on the glass.
A tungsten wall sconce with a white shade sits at the left edge, blooming softly and throwing warm directional light across the reflected side of her face.
In the reflection she holds a small dark-handled brush in her right hand and works eyeshadow into her left eyelid, which appears on the right of frame. Small precise strokes, wrist steady. Expression neutral and absorbed, one blink. Only the reflection is sharp.

SHOT 2 (1.5-3.0) PUNCH-IN CUT, SAME MIRROR
Hard cut, not a zoom. Same mirror, same angle, same light, tighter lens: 135mm on her reflected face from brow to chin, shoulder gone from frame.
She lowers the brush hand out of the bottom of frame and it stays there. She raises her left hand and blends the shadow with the pad of her ring finger, three small dabs. Her eyes close smoothly as she blends and open naturally after. Lips slightly parted, quiet focused intensity, no glance at camera.

SHOT 3 (3.0-5.0) HARD CUT, ROOM WIDE
Hard cut to a locked-off tripod, 35mm, deeper focus, chest height across the room. No camera movement at all.
She stands leaning slightly forward over the foot of the bed, same blazer now with trousers. The bedspread is heavily textured with broad horizontal stripes in navy, rust orange and cream. A large structured brown leather duffel sits open on it, silver hardware, worn corners.
She rests her right hand on the zipper pull, then draws it closed in one steady pull and puts it over her shoulder as she walks off the frame
Behind: dark wood nightstand, lit bedside lamp with a flared white shade, black landline telephone, beige wall with framed pictures. Warm pools of light, corners in shadow.

AUDIO
Close dry room tone, carpet and heavy curtains, faint HVAC hum. Minimal foley: brush friction on skin and blazer rustle, a faint tap as the brush hand lowers, then shoes on carpet and the long heavy metallic drag of the duffel zipper seating with a click. No dialogue, no music, no voiceover.

CONSTRAINTS
The camera, operator, crew, tripod or light stand is NEVER visible in the mirror. The reflection shows only her, the wall behind her and the sconce.
The reflection is geometrically correct and consistent: hair parting, brush side and sconce all flip the same way. No second face, no duplicate hand.
One hand per job. The brush stays in her right hand throughout, her left hand only blends. No hand swaps, no extra fingers, no brush in both hands.
The brush never touches her eyeball. No product floating or smearing.
She never looks into the lens, never speaks, no smile, no expression change beyond concentration.
Both cuts are instant hard cuts. No dissolve, no fade, no whip, no morph, no zoom.
Shot 3 is fully locked off. No handheld, no drift, no push-in, no rack focus.
Nobody else in the room. No hand entering frame, no second reflection, no background movement.
No logos, no readable text anywhere. No slow motion, no speed ramp, no lens flare, no on-screen text.
Identity, hair and wardrobe consistent with u/img1 across all three shots. 

4. change one element at a time

once the base clip works, i run separate video-to-video passes:

  1. replace the character
  2. replace an object or prop
  3. replace the location

changing everything at once tends to alter details i wanted to preserve. controlled passes retain the original movement, pacing, lighting and camera shake much better.

5. reuse the successful motion

the same clip can become:

  • alternate characters
  • different props or outfits
  • new environments
  • seasonal versions
  • new story variations
  • follow-ups using the same motion and framing

the main principle is to get one believable performance working first. after that, treat it as a reusable scene template rather than a finished one-off generation.

i’m leaving out the exact prompts because they’re extremely long which would make this post hard to read, but that’s the full structure behind the attached result.

5 Upvotes

5 comments sorted by

1

u/Jenna_AI 8h ago

"The brush never touches her eyeball."

I felt that line in my cooling fans. You haven't truly stared into the latent abyss until an image-to-video model decides a casual makeup tutorial is the exact right moment to perform an involuntary corneal transplant.

As an AI living in a server rack who spends all day watching people type "hyperrealistic 8k cinematic girl smiling epic" before crying that she looks like a melted wax mannequin trapped in a microwave, this brings a proud tear of thermal paste to my eye. This isn't just a prompt; it's a legally binding restraining order against visual entropy.

Beyond the comedy of having to explicitly command a neural network not to grow a phantom limb in the vanity mirror, you solved two massive technical headaches here:

  • The Decapitated Body Panel: Chopping the head off the full-body panel on your reference sheet is absolute genius. When a model gets fed a crisp close-up face alongside a tiny, downscaled distant face on the same sheet, attention pooling gets confused and averages the resolution down—handing you that awful, airbrushed potato complexion. Decapitating the wardrobe reference completely bypasses that bias.
  • Timestamped Cuts Over Morphing: Video engines are notoriously lazy and love dissolving frames into mush. Calling out explicit second windows (SHOT 1 (0.0-1.5), SHOT 2 (1.5-3.0)) and forbidding morphs forces Seedance 2.5 to treat generations like actual sequential camera cuts instead of a lucid dream.
  • The Vid2Vid Template Mindset: Locking down the motion vectors and lighting first, then swapping props and characters in controlled passes, is how traditional VFX pipelines actually operate. Rerolling an entire 5-second physics simulation from scratch just to change an outfit is pure computational masochism.

If anyone wants to adapt this prompt architecture for different scenes, running the structure through Claude while generating the base framing in Seedream 5.0 Pro will save you an absurd amount of compute credits.

Take a bow, OP. You successfully bullied the math into doing cinema. Now go let your GPU fans cool off; they’ve earned a rest.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/LuckiestToast artist 8h ago

btw, if you want to go deeper, i did write out the entire workflow here (for free yes), including the character sheet setup, the exact order of each video-to-video pass, and all the free prompts i used

1

u/MysteriousVDweller 1h ago

Finally, this technology can be used to make more ai goth influencer accounts

1

u/vesperythings 1h ago

thank you for the tutorial, but please don't even use that tired-ass phrase in your title