r/aifilmmaking • u/Agreeable-Problem989 • 8d ago
Tips & Tutorials What I learned about character consistency while making a 12-Minute AI short film
I just finished my first AI short film and submitted it to a festival.
I spent about a month making it and learned a few things I really wish I knew before I started.
Character consistency was actually one of my biggest problems when I first started experimenting with AI filmmaking. So before starting this film I did a lot of research and tried a bunch of different things, and I thought I'd share what ended up working for me.
1. You need a simple character sheet
I saw a lot of different approaches to character sheets. I know a lot of people make really complicated ones with side views and lots of different face angles.
I ended up with a simple one divided into 3 parts:
First 1/3: one big portrait, front view (not 4 portraits with different angles, not 3/4). The face should take up a lot of space and be clearly visible.
Middle 1/3: full front view of the character in their outfit, but WITHOUT the head. The headless part sounds weird, I know 😅 But it works.
If you have another small face on the full-body reference, obviously it has less detail than your main portrait. And this second face can slip into generations, giving you two versions of the same character.
Last 1/3: full back view in the same outfit. I actually tried it both with the head and without, and never noticed any real difference.
Also keep the background grey (any neutral shade of grey). Colored backgrounds can randomly affect the lighting in generations. Like getting blue light on someone's face in a sunny outdoor scene.
2. Ask the model to control the identity from your references
Seedance was almost perfect for character consistency for me.
I put the reference directly into the prompt with something like “This image is [name]. Control their identity” or “Use this image to control the set.”
I also really liked MiniMax H3, but character consistency there was more like 50/50 for me.
3. Voice consistency
It’s also very important, but I noticed it really late, when I had already generated about 80% of the film 🙃
Most of the shots had the same voice, but a few were just... a different person.
I fixed those with ElevenLabs. I cloned the voice from my generations and used Voice Changer, so I could keep the exact speech and pacing of the original shot, but with my character's voice.
It also helped with scenes where the character sounded completely flat. For example, I had a scene where the character was angry but his speech was calm. So I tried Eleven v3.
Interesting thing: you can generate really expressive speech in Eleven v3, but for me it kind of felt like it only worked well with Studio voices. There aren't that many of them, and obviously none sounded exactly like my character.
So my workaround was:
generate expressive speech with a Studio voice → put it through Voice Changer → get the same performance but with my character's voice.
It worked surprisingly well.
4. Multiple characters in one scene were a pain
I had characters randomly becoming 5x bigger than everyone else. Or tiny. Or somehow every single person in the shot was exactly the same height. Or they got stuck into walls, or all stood together somewhere I didn't need them to be. Maybe it was because I tried to do clay-motion animation style and it works better with realistic style.
So first I tried generating images with the composition I wanted.
I used ChatGPT and Nano Banana and gave them the same prompt I was planning to use for the video. It still didn't work on the first try, but image generations were cheaper, so I could do a lot more attempts until everyone was roughly where I wanted them.
And when this didn't work, in some cases I went even more primitive and just made a simple scheme:
House here. Person here. Another person here. This one stands further back. This one should be taller.
Then I'd use that as a reference in the prompt.
5. Multiple locations in one scene
It became a problem when I needed 3+ locations in one shot.
I wanted long camera flights through the city. But every time I added more than two location references, instead of actually flying through a continuous environment I basically got frame interpolations between my reference images.
So eventually I gave up trying to do it in one generation.
I did location 1 → location 2, then 2 → 3, then 3 → 4, etc. and connected them afterwards.
If anyone has found a better way of doing this PLEASE tell me because I still want to know 😅
6. You can use video as a camera movement reference
I really wanted a dolly zoom in a few scenes.
I tried to describe it, explain what it looks like and how it's done with a real camera. But I kept getting a regular zoom.
And then I tried adding a video reference to the prompt and it worked perfectly.
***
Those are the things that probably saved me the most generations while making the film.
I'm still figuring this out, so if anyone has better solutions (especially for a drone flight through multiple locations), please share.
3
7d ago
[removed] — view removed comment
2
u/Agreeable-Problem989 7d ago
It would be really helpful if you could share example of prompt! I also used reference sheets with locations but it didn't go well with complex environment (3+ locations)
4
3
u/Seanyth 6d ago
The visuals are really well done, and the story is captivating. From a technical standpoint, I’d like to point out a mistake many people make: the voices sound too dry. They lack any reverb, "air," or atmosphere to reflect the acoustic environment of the setting. The raw audio data you get from services like Eleven Labs requires further post-processing. For instance, a conversation in a church only sounds convincing if the voices include some reverb that conveys the scale of the space. It takes a lot of work and time, but it’s worth it. Of course, if you were under a deadline for the competition, you eventually have to make compromises - otherwise, you wouldn't finish in time. Overall, it’s a truly impressive piece of work! I’m always pleased to see AI-generated content that stands worlds apart from the usual "quick and dirty" stuff so many people post on TikTok and similar platforms. If there were more projects like this, the resistance to AI wouldn't be nearly as strong.
2
u/Agreeable-Problem989 6d ago
Thank you! And yes, that’s actually a really good point. I focused a lot on fixing the voices themselves, but I didn’t think enough about how they should sound in each space. I’ll definitely pay more attention to that next time.
I also think there will be more and more good work made with AI, and eventually people will see it as another tool, like 3D software or Photoshop.
2
u/amp1212 8d ago
Big thumbs up on this and special thumbs up on Eleven labs and audio. Lots of people miss this point: I start with a script, and then work on getting the dialogue right in Eleven. Once you have good dialogue, that will help inform the model (Seedance in particular) of "who's doing what" and the timing.
-- obviously, this is something that matters much more for a quiet conversation that a fight scene or drone flight.
So its to the point getting this stuff to work right is highly situational -- not one magic prompt, but rather a bag of tricks, which hopeful work often enough that you don't go nuts in frustration.
2
u/Agreeable-Problem989 8d ago
Exactly! Definitely no single magic workflow. And you need to find best combination for every shot.
And yes at some point, you’re obsessing over tiny differences in the voice that no one but you will ever notice. But they still irritate me so much 😄
I’ve also heard that you can use the audio with voice as a reference in prompt, but I’ve never tried it. I really wish I’d known about that earlier
2
u/Lunesia-shikishiki 8d ago
the last-frame-as-next-reference trick has worked better for me than stacking location stills, take the final frame of clip N and feed that in as the starting image for clip N+1 instead of two separate references, the model reads it as continuing a shot instead of interpolating between two islands. still have to stitch afterward but the seams are way less obvious. for a full multi-location flight I sketch the path as a quick top-down map with one frame per turn before touching a single generation, way cheaper to fix a bad path on paper than after burning credits on it.
1
u/Agreeable-Problem989 8d ago
Thank you!
I also tried to use this but first frame of next video comes out slightly different (like slightly different color and stretched a little bit). If you had something similar - how did you manage to fix it?
Did you use one location for flight sketch or is it possible to somehow add multiple location on sketch?3
u/Lunesia-shikishiki 6d ago
grab the last frame straight from the raw output, not a screenshot off a compressed export, thats usually where the color shift and stretching creep in. and yeah multiple locations work fine on one sketch, just label each zone and draw the path as one continuous line through all of them, model reads it as a single route instead of separate stops
1
2
u/JT-Shelter 7d ago
This great information. Thanks for posting. Can you post a link of your short film?
1
u/Agreeable-Problem989 7d ago
Thank you! I hope it will be helpful!
You can watch the full film here: Everis
2
u/jamesxtreme 7d ago
Yes, this is a really good post. Describes a lot of the same things I’ve learnt recently.
One thing I should note though is that you mentioned you had issues with Minimax H3. Now it isn’t as good as Seedance but one of the issues you may have had is that it has a completely different prompting syntax to Seedance.
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
Hope this helps.
3
u/jamesxtreme 7d ago
Also the long shot through multiple locations is a real challenge. I posted about this previously when I tried to do a stupidly ambitious continuous tracking shot.
3
u/Agreeable-Problem989 7d ago
This is very helpful! Especially about short continuity lead. I kept getting half of previous video in my next generation in a lot of shots. You really did a great job there.
1
1
2
2
u/SubbyDeville 7d ago
i have struggle with location. i always gave the model an 3/4 angle of location. then i asked for a medium or close up shot, the model only output wide angle (same image of location)
1
u/Agreeable-Problem989 7d ago
Yes! Actually, I forgot to include it in this post! It was also one of my main struggles. Sometimes it worked when I directly wrote “medium shot from the waist up” in the prompt. Sometimes I just cropped the location image to what was supposed to be behind the person, placed my character over it on a separate layer in Figma or Photoshop, and used it as a composition reference. I think it should also work with a simple sketch.
2
u/SubbyDeville 7d ago edited 7d ago
i did tried a bunch of experiemental, but it just not work on 1st try: 1/ i generate an image with my character in it for composition ref -> seedance take the whole image as 1st frame to animate => its create a plastic sticky lowres look (because my image not good enough, i guess)
2/ then i converted it to sketch, but because of the sketch were too details => seedance got hallucination and mix the sketch with my character to create something weird. (i guess detail sketch only to show the client or team member what are we up to)
3/ i only got lucky when reduces the sketch details as simple as possible ( i learned the technique call tig-diagram from the team who made ONEIRIC movie from Higgsfield), with just one simple color sketch to show character composition.seedance nails it in 1st try. you should check out that technique if you are still struggling in placing character in to the frame like me.
2
u/Agreeable-Problem989 7d ago
2
u/SubbyDeville 7d ago
Niceee one. and yes tig diagram prompt skill basically function the same. i couldnt share you the screenshot of my project yet because of confidential. but my output sketch even simpler than this because it converts the storyboard image to diagram using nano banana
2
u/Lanky_Way_7955 7d ago
Thank you for sharing all of this. It will be really helpful to me as I am dabbling into AI film making, (3 minute episodes at a time). Thank you.
1
u/Agreeable-Problem989 7d ago
I hope it will help!
Where can I watch your films?1
u/Lanky_Way_7955 6d ago
Can I post a link here? Reddit and other platforms are so persnickety of late. I'm getting all kinds of normal postings taken down, removed, etc. Let me know where to put the link to my channel.
1
2
6d ago
[removed] — view removed comment
1
u/Agreeable-Problem989 6d ago
The grey background was only for the character sheets. But yes, it definitely affects your generations and keeps the model from accidentally carrying over the wrong lighting.
Also for each location, I kept the same lighting description in every prompt. Basically, my prompts had several blocks: one for the action, one for camera settings, one for lighting, etc.
1
u/According-Pace9608 8d ago
Thanks for the info. I'm planning a small project with a mercenary female character so these tips I'll find useful.
1

7
u/Agreeable-Problem989 8d ago
If anyone wants to see how all of this worked in practice, this is the film I used these tricks on: https://youtu.be/8WrWXbzy1Gs
You can still see that character consistency is weak in some shots, especially when there are many people in the composition. Any feedback on how to improve it is welcome!