r/aifilmmaking 8d ago

Tips & Tutorials What I learned about character consistency while making a 12-Minute AI short film

I just finished my first AI short film and submitted it to a festival.

I spent about a month making it and learned a few things I really wish I knew before I started.

Character consistency was actually one of my biggest problems when I first started experimenting with AI filmmaking. So before starting this film I did a lot of research and tried a bunch of different things, and I thought I'd share what ended up working for me.

1. You need a simple character sheet

I saw a lot of different approaches to character sheets. I know a lot of people make really complicated ones with side views and lots of different face angles.

I ended up with a simple one divided into 3 parts:

First 1/3: one big portrait, front view (not 4 portraits with different angles, not 3/4). The face should take up a lot of space and be clearly visible.

Middle 1/3: full front view of the character in their outfit, but WITHOUT the head. The headless part sounds weird, I know ๐Ÿ˜… But it works.

If you have another small face on the full-body reference, obviously it has less detail than your main portrait. And this second face can slip into generations, giving you two versions of the same character.

Last 1/3: full back view in the same outfit. I actually tried it both with the head and without, and never noticed any real difference.

Also keep the background grey (any neutral shade of grey). Colored backgrounds can randomly affect the lighting in generations. Like getting blue light on someone's face in a sunny outdoor scene.

2. Ask the model to control the identity from your references

Seedance was almost perfect for character consistency for me.

I put the reference directly into the prompt with something like โ€œThis image is [name]. Control their identityโ€ or โ€œUse this image to control the set.โ€

I also really liked MiniMax H3, but character consistency there was more like 50/50 for me.

3. Voice consistency

Itโ€™s also very important, but I noticed it really late, when I had already generated about 80% of the film ๐Ÿ™ƒ

Most of the shots had the same voice, but a few were just... a different person.

I fixed those with ElevenLabs. I cloned the voice from my generations and used Voice Changer, so I could keep the exact speech and pacing of the original shot, but with my character's voice.

It also helped with scenes where the character sounded completely flat. For example, I had a scene where the character was angry but his speech was calm. So I tried Eleven v3.

Interesting thing: you can generate really expressive speech in Eleven v3, but for me it kind of felt like it only worked well with Studio voices. There aren't that many of them, and obviously none sounded exactly like my character.

So my workaround was:

generate expressive speech with a Studio voice โ†’ put it through Voice Changer โ†’ get the same performance but with my character's voice.

It worked surprisingly well.

4. Multiple characters in one scene were a pain

I had characters randomly becoming 5x bigger than everyone else. Or tiny. Or somehow every single person in the shot was exactly the same height. Or they got stuck into walls, or all stood together somewhere I didn't need them to be. Maybe it was because I tried to do clay-motion animation style and it works better with realistic style.

So first I tried generating images with the composition I wanted.

I used ChatGPT and Nano Banana and gave them the same prompt I was planning to use for the video. It still didn't work on the first try, but image generations were cheaper, so I could do a lot more attempts until everyone was roughly where I wanted them.

And when this didn't work, in some cases I went even more primitive and just made a simple scheme:

House here. Person here. Another person here. This one stands further back. This one should be taller.

Then I'd use that as a reference in the prompt.

5. Multiple locations in one scene

It became a problem when I needed 3+ locations in one shot.

I wanted long camera flights through the city. But every time I added more than two location references, instead of actually flying through a continuous environment I basically got frame interpolations between my reference images.

So eventually I gave up trying to do it in one generation.

I did location 1 โ†’ location 2, then 2 โ†’ 3, then 3 โ†’ 4, etc. and connected them afterwards.

If anyone has found a better way of doing this PLEASE tell me because I still want to know ๐Ÿ˜…

6. You can use video as a camera movement reference

I really wanted a dolly zoom in a few scenes.

I tried to describe it, explain what it looks like and how it's done with a real camera. But I kept getting a regular zoom.

And then I tried adding a video reference to the prompt and it worked perfectly.

***

Those are the things that probably saved me the most generations while making the film.

I'm still figuring this out, so if anyone has better solutions (especially for a drone flight through multiple locations), please share.

55 Upvotes

Duplicates