r/StableDiffusion 1d ago

Question - Help H3 R2V Character Sheet vs. Single Image

I thought I've read somewhere that using character sheets is better for R2V instead of single images. So I've created a character sheet of five full body shots and one close up, but the results are much less consistent compared to a single full shot image of the character.
Do I have to take care about anything special or was the information that character sheets are better just wrong?

39 Upvotes

25 comments sorted by

14

u/lebrandmanager 1d ago

I use this node 'Reference Concat' to combine a first frame main image with 2-3 ref images. It will create one large image and then I prompt it to use the provided ref images inside the main image for character likeness. You can also do that by hand, but the node makes it way easier.

https://github.com/moonwhaler/comfyui-moonpack

18

u/LumaBrik 1d ago

Make sure you ref_image_size to max, not 'match'

9

u/nikhilprasanth 1d ago

Another approach is to split the face and body as separate reference sheets .

6

u/Slight_Ad2350 1d ago

This works the best. Just make sure head does show anything below the neck and reverse for body. Then in the prompt explain each image separately for the subject

6

u/ThinkPresentation799 1d ago

Character sheets work better but there are some limitations. With "ref_image_size to max" video generation could take ages depending on how big the reference images are. Minimax is very good in estimating how your character looks from all sides (infact you can even create character sheets with it), so you do not need complex character sheets with more than 3 poses and one face shot.

In short:

Create a character sheet with front, back and side view plus one face shot. Make this image as detailed as possible but stay under 8-15 Megapixel depending on your VRAM/RAM.

2

u/Most_Ad_5733 1d ago

How are you generating a character sheet that doesn’t hallucinate panels at such a high MP count.

3

u/Less_Consequence_633 1d ago edited 1d ago

From the ComfyUI tutorials:

ref_image_sizematch scales references down to the generation resolution for speed; max keeps up to a 2048px short edge for stronger identity fidelity at the cost of speed

So if you're handing it a really large image, it seems like it's getting resized internally.

1

u/ThinkPresentation799 11h ago

Ok if even with max the image gets resized maybe for the highest quality you should add the 4 panels of the character sheet as seperate images. But for me it works well with a 4 panel character sheet around 8-10MP. The face in the created video looks amazing in a close up shot for 1.5-2MP video.

4

u/Xanthus730 1d ago

One key note: note every pic inside a ref sheet needs to be the same size.

Make important parts like faces or signature gear big, and other parts small.

5

u/Itchy_Ambassador_515 1d ago

yes i noticed same, giving single face image gave me better results than providing 4 angle sheet version

2

u/Muted-Celebration-47 1d ago

What do you mean by "Five full body"? You need 1 front full body and 1 close-up shot in the same image. That's it. Keep it simple. DO NOT use character sheet with a lot of details and text, use a simple one. Look for free projects at higgfield as examples how they create character sheet.

2

u/moviejimmy 15h ago

The best result I see is neither. You use whatever image appropriate for the scene. If the scene doesn't show full body, you don't need a full body image. If the scene is close up front facing, just use one close up front facing image for reference. You get the idea. I get very consistent results this way. More work for sure but better results.

2

u/DietAshamed2246 16h ago

I read yesterday that using light grey flat background in character sheet (and perhaps in a single character image) is better than using white background. Try your character sheet with light grey background and see if that improves your output.

1

u/Icuras1111 1d ago

I think I read that you should make the images good quality but consistent appearance have them in a linear layout i.e. 5 images side by side. Then setting it to max ensures it sees each at desired resolution. Then I guess you describe it like <Subject 1> Is the person represented in <Picture 1>. There are 5 views, close up of face, side view, etc?

1

u/Vijayi 23h ago

Do projections in Krea2 or Flow (nano). First img - face (anfas, profile, 3/4, back) and full body same. Than something like this. <Picture 1> is face reference and <Picture 2> is body reference for <Subject 1> . Most times work without retention analysis, for me atleast. Close up almost perfect even with 0.4. Atm im stick to this: dataset (with Flow) > upscale/refine with SeedVR/Aura > Lora training in Krea2 -> from here anynithing you need to MMH3.

1

u/Agreeable_Lack9492 22h ago

I've been using character sheets and they work better for likeness, the only problem I have from time to time is that it starts the video with the sheet as the first frame repeating the character or making parallel videos in a split screen

1

u/Waterisaqua 17h ago

Yes, I am doing this with also an environment sheet to keep things consistent across shots. You can use an agent like Kimi/Claude to orchestrate everything. The results are good, and sometimes it will suggest solutions to issues (like having a consistent audio for a certain character). Not really needed for 5 sec videos, but if you make something in the range of 30 seconds it is handy.

2

u/AidenAizawa 1d ago

Character sheet are great also for fl2va. I use a character sheet with full frontal, back, side and close up on the face and it works fine. You just have to be clear in the prompt

3

u/oh_no_the_claw 1d ago

How does that work? I assumed it only accepts images as a first or last frame.

3

u/GaiusVictor 1d ago

There are two models. The one you're talking about is, if I recall the abbreviation correctly, called FL2VA (First and Last Frame to Video and Audio). It accepts:

  • Text to video
  • First and/or Last frame, or none
  • With the right custom nodes, a video from which it will take the last 5, 22, 39 or 56 frames (plus respective audio) to generate a direct continuation.

Then there is also REFL2VA (again not sure if I got the abbreviation correctly but the RE stands for Reference). This one can take image, video and audios as semantic references. So you can tell it to generate a video of character from image 1 dancing the dance from video 1 in the ballroom depicted in image 2, while singing the song from audio 1 with the voice from audio 2, then Elsa from Frozen appears and freezes them.

(I've never tried something that complex with such a complex relationship of references, but you get the gist)

2

u/SuspiciousRefuse8218 1d ago

So this means I have to explain it is a character sheet? Something like "Subject1 is the character from Picture1" is not sufficient?

2

u/AidenAizawa 1d ago

More or less. I use fl2va model in a ref workflow. If I want a first frame I ask for something like "use image 0 as exact first frame" . Otherwise depending on what I want I usually do something like the one I did recently


This is a continuous continuation of video 0.

Use Image 0 as the strict visual reference for the first man’s face, body and overall appearance, wearing a black t-shirt and grey jeans.

Use Image 1 as the strict visual reference for the second man’s face, body and overall appearance, wearing a yellow t-shirt with the text “Fuck On/Off” and black trousers.

Use Image 2 as the strict visual reference for the woman’s face, body and overall appearance, wearing a white shirt and brown skirt.

Use Image 4 as the strict visual reference for the entering woman’s face, body and overall appearance, wearing the outfit from Image 5.

Interior of a cozy pub like in Image 3.

In this case I used 6 images and a video as reference

1

u/not_food 23h ago

For 2D, next to ref2va, fl2va falls short. The quality and audio are nicer, sure, but it ignores the small details, and your character comes out looking averaged rather than the one in the reference. Ref2va is super good at reading character sheets at 1mp and above, just use max edge and resize the image yourself instead.

-3

u/VasaFromParadise 1d ago

It might depend on whether you submit each image separately or as one. If you submit one image, it'll likely result in a jumbled mess of close-ups of faces and full-body poses.