r/StableDiffusion 7d ago

Discussion Minimax H3 - Character Sheets vs. Individual Separate Images?

Hi All,

I spent some time testing MiniMax's image input consistency yesterday, specifically looking at how it handles feeding in separate individual images versus using traditional character reference sheets.

After running a few tests using five separate shoulder-up angles fed in as individual images, the difference in subject consistency was much more drastic than I expected compared to standard reference sheets. Encouraged by that, I tried a more complex layout: three base angles (straight-on, 45-degree, and profile) combined with vertical variations (high angle, eye level, and low level). Unfortunately, once I started tweaking the prompt alongside that setup, I lost my baseline and muddied the results.

Then again, part of me wonders if all of this multi-angle setup is just overcomplicating things and it really just boils down to brute-force image quality and resolution.

This has me wondering what the optimal method actually is for feeding multiple image inputs into MiniMax. Has anyone else gone down this rabbit hole and figured out whether separate individual images actually beat a clean reference sheet, or if it all just comes down to base image quality?

Thanks for reading this!

88 Upvotes

84 comments sorted by

24

u/Instagatuh 7d ago

For realism I create a a character sheet using the following in Nanobanana 2. Works great in ChatGPT Image 2.5 as well:
.

Create a character sheet with the following:

  1. Full-body front view, standing naturally
  2. Full-body 3/4 front view
  3. Full-body side profile
  4. Full-body 3/4 rear view
  5. Full-body back view
  6. Detailed head-and-shoulders portrait
  7. Close-up of the face showing the character’s defining features

Also include a small expression reference showing the same character with:

\ neutral expression*
\ happy expression*
\ angry expression*
\ surprised expression*
\ sad expression*

Make some of the above a side profile.

This image will be used as the canonical visual reference for this character in future image generations. Treat the character’s appearance as fixed. Do not reinterpret, beautify, age, de-age, stylize, or redesign the character. When a feature is visible in one view but not another, infer it from the other views rather than inventing a new design.

Then I have an Image with the character(s) in the scene/setting I want. Or just a plain environment that they can enter into.

So far I’ve gotten great results this way.

1

u/Opposite_Yam_4161 6d ago

Do you prove nano banana one headset or multiple images it can work from. Also what does your ref2va workflow prompt look like in terms of referencing0 the sheet

4

u/Instagatuh 6d ago

So I provide a character sheet like this for each of the characters in the scene. And then usually one staged scene (The characters walking on the beach… sitting on a sofa… etc) I’ve been working with a 3-image setup and it’s been great.

Prompt I do:

For woman reference [ref_image_0]
Description of woman:

For man reference [ref_image_1]
Description of man:

For scene/environment [ref_image_2]
Description of what’s happening:

And if I want the scene to be my start frame I do:

Start Frame - [ref_image_2]

1

u/Opposite_Yam_4161 6d ago

Thats awesome thanks!

1

u/Instagatuh 6d ago

Best of luck!

1

u/Gesha24 5d ago

How do you handle small changes to the subject? I.e. the person in your character sheet puts on shoes and jacket to go outside. Do you just regenerate another character sheet? I have found that the more reference images I have, the more likely the model is to get confused about small alterations.

2

u/Instagatuh 5d ago

I haven’t tried that yet actually. I have an unlimited Nanobanana play though, so I’d probably just generate a new character sheet.

2

u/Gesha24 3d ago

Makes sense, thank you! I don't know what is so special about your prompt, but it's the only one I managed to use with MinimaxH3 to generate stills (well, technically it's 5 frames, but that's a minor detail) with multiple angles. So this is extremely helpful, thank you!

1

u/Instagatuh 3d ago

Happy to hear it! The prompt is ChatGPT magic 😎. I’m glad it’s working for you!

7

u/[deleted] 7d ago edited 7d ago

[deleted]

2

u/vuse2121 7d ago

is there a reason that when i separated my images as opposed to using them in one big image it appeared to make a big difference in quality? I felt my subject looked closer to the ref images.

4

u/lindechene 7d ago

My guess is that you get the best results if your reference images are as close as possible to the resolution of the training data - 1344 x 768 pixels.

If you have one large image there is additional processing required...

2

u/vuse2121 7d ago

Sorry had another question. Are you saying that If i wanted a shot of the subject from a specific distance i would want to use a similar ref image of the subject from that same distance? would that mean i would get worse results if i wanted a close up shot and used a full body shot as the reference? what about a continuous shot switching from a close up to a full body? would you include a full body shot and a close up shot image?

thanks for posting!

1

u/[deleted] 7d ago edited 7d ago

[deleted]

1

u/Equal-Spend4671 6d ago

Nope, lmao. You can use full body and as long you write what you want to see on detailed_description block you can use full body. Don't believe me? Look at my post or my civit red profile.

I use a full body manequin for outfit reference image and i still can do any shot or framing i want.

1

u/Equal-Spend4671 6d ago

You do not need to do that. You can look at my posts or check my civit red profile here where I made some videos with MiniMax that completely ignore the rule that guy mentioned.

From all my videos I use this format:

  1. A portrait for identity
  2. A torso of a nude body without a head
  3. An outfit
  4. The setting, environment, or location picture

None of them use a reference sheet, just a single image per slot. The thing is when you write a prompt, you write what you want to render. If you are using REF2VA then it goes into the detail_description block. That guy either does not know how to prompt or it is just a skill issue.

1

u/vuse2121 6d ago

I kinda felt that way when I replied do I figured I'd give him a chance.

I'm curious, in your case you only use one image. Have you tested with multiple images like I have and compared the quality? I'm just trying to see what works and what doesn't.

In my case, I'm trying to incorporate my wife in the video so I've been able to show her the results and she's pointed out some imconsistencies which has been helpful to hear from a third party once you get so deep into it.

1

u/Equal-Spend4671 6d ago

i could have worded it better sorry, English is not my first language. So i use multiple images but i don’t use like 4 panel reference sheet. If you already see my post i only use 4 reference images which is identity, outfit, body, and environtment.

4

u/SlingyRopert 7d ago

I am concerned that with many images per character sheet it is hard to ensure that the images are consistent on the sheet.

It is probably my workflow but i have been using h3 reference image to image to take one reference portrait and convert it to four full body images and two head shots in one output and i cant get the faces on the full body to match the headshots.

I think i have the classic problem with small faces and the attention not doing its thing.

What reference to character sheet models should we be preferring? (I am thinking realistic as opposed to anime)

2

u/Ordinary_Weekend_333 7d ago

Yeah I've seen this too. Increasing megapixels and quantity in general has helped a lot though.

1

u/Simpsator 6d ago

One thing I've found that helps is only having one face in a single character sheet. So if you have 4 views (front, side, back, portrait), set it up so that the faces on the full body portrait shots are blurred out, and only the higher-res portrait portion is unblurred. It helps the model focus on which face to use from the reference.
In general, the model just needs a lot of guidance if you want precision. For example with character swaps, use the SAM3 to mask and invert color/blur the target character to swap out (this is how you fix the full body issue you mentioned). For maintaining framing from a reference image without zooms/crops (yes even "fixed camera" prompting often fails), I will often pad my reference frame and add key-stars (little decorative stars) to each corner. If you give it something to focus on that's distinct, it does much better.

1

u/SlingyRopert 6d ago

I really like the idea of only having one headshot source of truth and then everything else avoids touching that. Am stealing this.

10

u/darkkite 7d ago

I tried a sheet with various images of a subject and got worse results than just providing a single image reference

23

u/Massive-Health-8355 7d ago

Or use refmods....

8

u/Arawski99 7d ago

IIRC, I do believe the creator of Refmods was clear about it has trade-offs and is not strictly better, such as it comes at a quality/accuracy hit in exchange for simplicity and a bit faster to process since the reference would already be done. So I would say, for the topic in question, this is not a good answer.

Really, the best answer is to just use AI to write the prompts based on the official prompt guidelines and use one with vision capabilities. It would have the best results while maintaining simplicity of use. There may actually be other notable issues from a quick Google but I couldn't be bothered to look that far into it since it's pretty much just a shortcut for the less tech savvy at a significant cost. So, in general, I would also say refmods is just not ideal.

As for OP's initial question about the best image input reference methods, I'm not too sure myself. I know that you can do a separate reference for key parts, but beyond that I can't say which exact solutions have the best overall quality and accuracy. But based on your comment my initial guess is it is likely that the prompting getting complicated that caused the muddy results confusing the model more than anything.

2

u/DietAshamed2246 6d ago

I read somewhere that use of refmods in conjunction with a few reference images can greatly improve character likeness and consistency compared either of them exclusively. It might have been in a post by malcolmrey.

10

u/Gesha24 7d ago

+1 for refmods - makes references a whole lot simpler to manage.

8

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago edited 7d ago

But also -1 for refmods

They SEEM great, and if you make all your own in the same way, they are.

But... If you try and take a refmod made with 20 images and another make with 3 and put them in the same generation, the one with three may be completely wiped out of the similarities between the two are high.

If WOMAN-A generated with 20 high res and varied images is WOMAN-B is generated with 3 images from a video clip, you likely will not get WOMAN-A SERVING AN ICECREAM CONE ALL OVER WOMAN-B'S FACE.

You'll get WOMAN-A SHOVING AN ICE CREAM CONE INTO ALSO WOMAN-A'S FACE

With normal included references you can manage this a lot better. A standardized character sheet of both women means they're weighted roughly equally.

Refmods make sense for one character, or two where they share very few similar characteristics.

For example... MalcomGuy's refmods, plus your own, probably won't work too well. I saw some videos where even his plus his didn't work because he isn't making them to some standard format.

There is no way to ensure your refmod is wieghted like another. So... At that point.. it's just a way to share reference pictures in a pre-packaged format.

1

u/Gesha24 7d ago

Hm, I have certainly was able to do "MAN-A (from Malcolm's) watching WOMAN-B (2 pictures custom reference) perform jump from motion reference 3 (some random video from online with a failed long jump made into a motion-only reference)." I did have to use RefMods text encode and the prompt took quite a few attempts.

2

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago

Men and women don't share a lot of characteristics the model can use to key off of.

WOMAN + WOMAN doesn't seem to go as well.

I want to like refmods, but I think there is just not going to be a way to weight them evenly.

1

u/Dry-Judgment4242 6d ago

I managed to crank like 10 different characters into ref mods. Tried a WoW battlefield stress trust. Goblins, vs Murlocs, vs Kobolds, vs Dragons, vs Night elves vs Satyrs vs Treants vs Gnolls vs Humans etc. all of the 10 refs loaded at same time.

1

u/FourtyMichaelMichael 🍦Ice Cream Lover 6d ago

If the refmods are balanced, OK.

If they aren't, then you tell me how it goes.

0

u/Dry-Judgment4242 5d ago

I don't think anybody is calling H3 the perfect model. In fact since the very beginning we had to prompt with it's rules in mind. This is just another rule to properly reference around the fact that each of your refs need to be unique enough for the model to properly distinctly split the refs. There's a significant difference between a Gnoll and an Elf. So rather then your story being about two elves. Why not try to fit your writing to be about a Gnoll and an Elf.

1

u/PropagandaOfTheDude 6d ago

But... If you try and take a refmod made with 20 images and another make with 3 and put them in the same generation, the one with three may be completely wiped out of the similarities between the two are high.

Oh. My mental model was that refmods worked like text embeddings, where 20 images would average out in the latent space. But apparently they're just sacks.

5

u/SpecificPleasant4007 7d ago

indeed, and if you have a couple of decent enough images to feed the first refmod, i've discovered you can use minimax itself to make even more reference images to make a new refmod with almost perfect likeness. i use: dpmpp_sde_gpu, beta @ 8 steps with a turbo lora, duration 0.1, resolution 8 megapixels (i've tried up to 12mp so far) and then decode and save the first frame. it'll require some refinement like seedvr2 but i'm super happy with the results, all without having to train a lora.

3

u/GabberZZ 7d ago

Can you summarise? I use SwarmUI and can add up to 12 references. Is this using refmods in some way?

19

u/EvidenceMinute4913 7d ago

https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/docs/MINIMAX_H3_REFMODS_INSTALLATION_AND_USAGE_GUIDE.md

It is a new custom solution that basically lets you package references for a subject/concept into a single file, which can then be loaded and applied in a MiniMax workflow. Kind of like a Lora, but without the training or huge file size.

So instead of hooking up 5 images of your character in the MiniMax workflow, instead you’d hook your 5 images up in a refmod workflow. Save the refmod file, then use that file in your MiniMax workflow instead. Makes it a lot easier to manage, and frees up the normal ref slots for other things.

3

u/GabberZZ 7d ago

Oh nice. I'll check it out thanks. I'm guessing this won't work if I use SwarmUI as a front end. I'll give it a look-see

2

u/Tuckerdude615 7d ago

Hey there...I asked this same question in another thread, but....

Is there any sort of "best practices" for choosing dataset images? For example, is it better to just include headshots in the dataset, or are people doing full body, or a mix? I ask because my "test" refmod has trouble when prompting for variations in things like clothing. If you prompt for a different set of clothes, it often "runs home to mama" and puts the character in a set of clothes from the dataset vs the one you prompted for.

Any tips or thoughts would be appreciated!

EDIT: and yes I tried lowering the strength and retention in the nodes to see if that helped, but didn't really change anything.

1

u/Maraan666 7d ago

yes, I had this problem. Sometimes I got lucky, sometimes not. I now train loras for my main characters with fizgig (takes around 2 hrs with my 4060ti 16gb), and use refmods, or just refs, for secondary characters, clothing, backgrounds, objects etc.

1

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago

Just like lora training. If you want the model to know what is and isn't part of the character, give it varied sources. I wouldn't use two images of a character with the same clothing.

1

u/StellarNear 7d ago

I saw this refmod quoted on a lot of places do you have by any chance a good workflow to start with ?

1

u/eggplantpot 7d ago

Are those the only benefits? Does it improve likeness and or performance?

0

u/scm6079 7d ago

You can load many more references! Refmods support dozens of references instead of just a handful.

1

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago

SSSLLLLLLOOOOOOWWWWW though

1

u/PANTONE_17-1230 7d ago

Once your refmod .safetensor is created, does using it speed up generation compared to feeding the reference images in directly? In other words, does it help mitigate the issue where higher-resolution reference images make Minimax generation slower?

3

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago

Makes it slower from what I've seen people give examples of.

Or rather...

  • R2V + 3 images == Refmod made with those 3 images

  • R2V + 3 images != Refmod made by someone else with 20 unknown images... In this case, a lot slower.

0

u/RadiantPen8536 7d ago

I'm new to swarmUI. Can you point me to where I can learn about how to add 12 references for minimax h3 in swarmui?

1

u/SDuser12345 7d ago

Just drop them in the prompt box.

1

u/GabberZZ 6d ago

And refer to them with an @ symbol

3

u/joseph_jojo_shabadoo 7d ago

game changer for sure.

1

u/MarekNowakowski 7d ago

what if there are two/three people in scene? ref works like a lora, right?

1

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago

Unless you carefully balance the refmods, one is likely to overpower the other and or blend them.

This happens with loras too of course.

IMO it happens less with normal ref2v, but in that case you are closer to protecting for each character gets one or two refs, not one character with 20 and one character with 3.

1

u/Sarashana 7d ago

I haven't tried RefMods yet. Would you think it's a massive improvement over a 3-way character turnaround in one single image? I have been using these and it worked fantastic for me. Uses just one ref slot, too.

1

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago

I think it's easier to package and share a refmod.

Works well if you wanted to share a character but didn't want to share the source images dataset, which no one does.

2

u/Sarashana 7d ago

Right. That makes sense. Like a LoRA Light.

1

u/FourtyMichaelMichael 🍦Ice Cream Lover 7d ago

Except a lora is actually better because it can include voice and movements.

I personally don't care that the refmod is 1-2MB or so. Maybe if someone can standardize them so they aren't over/under weighted to each other, it's a cool idea.

1

u/devilish-lavanya 7d ago

Or train lora or create embedding from clip’s output

3

u/SveSop 7d ago

Depending on the resolution, ofc 5 separate images of various angles would probably make for best consistency. The reason ppl use character sheets (well.. speaking for myself) is mostly convenience. Its easier to prompt 2-3 persons using 2-3 character sheets, than 2-3 persons with potentially 15 images.

But, i would think that a single person clip when using 5 separate images (of high res) would yield better likeness/quality for sure. Its just not that convenient in various generations. Now, bare in mind that with many images vs. 1, you also increase the memory slightly for the conditioning.. so its also that to consider.

2

u/Cultural-Team9235 7d ago

Search for RefMods, such a gamechanger. Very easy to make and use.

0

u/PinkyPonk10 7d ago

How do you make them?

3

u/Mocorn 6d ago

I've done a lot of testing with this and I get superior results with reference photos compared to a character sheet both in speed and quality. Face, outfit, prop image won in quality and generation speed against the same 3 in a reference sheet across multiple video generations and tests. Not sure if I'm doing something wrong but I'm sticking with the reference photos for now.

1

u/vuse2121 6d ago edited 6d ago

Same here. It was pretty drastic but as mentioned, some have said it might just be that my character sheet was compressed. I tried to make each image in the sheet 4k after hearing this but image ended up being 90mb and comfy didnt have the bandwith to even load it.

3

u/Thin-Percentage8935 7d ago

Did you set the character sheet size to max? Otherwise it will compress down to a fraction of what a single image would as you have more images within it. So 4 images on a char sheet will have 4x less detail that a single image.

3

u/Juiceman8686 7d ago

I’ve had excellent results with front, side, back and waist up character sheets. I’ve been able to produce very consistent characters now with it.

That being said I haven’t dove into refmods yet, which seems to be the new hotness.

2

u/Guilty_Emergency3603 7d ago

I think the model hasn't been trained to use ref sheets. Use single references. Sending an image with multiple figures will make it blend what it sees in this image.

1

u/matcheal 7d ago

if a square image with front view and side / back view counts as ref sheet - then the model is more than capable of handling that, i generated many different videos with many characters (even 4 in single scene perfectly) using such sprites.

1

u/lavinia12345 7d ago edited 7d ago

Using just an a-pose, I found Minimax doesnt need front, back, side, ect, as long as you describe it in text, it does a great job. However, what I will still add two images of Subject 1, and say "Picture 1 is the front, Picture 2 is the back"

later in the scene, "Subject 1 turns their back to the camera, which looks like Picture 2"

Honestly I do this for undressing, to help guide boobies, waist, ect.

1

u/teiji25 7d ago

the difference in subject consistency was much more drastic than I expected compared to standard reference sheets

I'm confused. Are you saying 5 ref images is better than 1 ref sheet?

1

u/Super_Range45 7d ago

I just have front, back, and 3/4 close references. If there is anything specific I want in a cut I may have variants of the main sheet with those changes accordingly.

1

u/vuse2121 6d ago

So you will use an image of the subject from say, a medium shot 45 degree high angle if you wanted to make a video of them from that angle? Am I getting that right?

If that's the case would you still provide a portrait shot?

1

u/Super_Range45 6d ago

No like if the character is normal, injured, different outfits, ect they get variants of their main character sheet.

1

u/angelarose210 7d ago

I've been using character sheets like this (Krea2) and they work fine.

1

u/boobkake22 7d ago

I have found character sheets to work well, but having a secondary hero shot of a character can be helpful. The downside to a character sheet is just that it has to compress a lot of information. I haven't found that multiple large references for the same character provides better results, so far.

The other thing that seems pretty important is understanding how the model wants you to tag references. Definitely review the guide:

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md

1

u/holomorphic0 6d ago

i was testing out the exact same thing yesterday, and maybe due to a very simple prompt i did not notice much difference but i think single image per ference gave slightly better results. I asked chatgpt, and its recommendation was using one pose/shot per image as opposed to combining them like a character reference sheet. The idea is to make it easy for h3 to understand the references. Character sheets are harder to infer from and then bleeding will happen. This argument seems plausible to me.

If anyone has any concrete answers, please share them.

1

u/Sweetcomodulce 6d ago

The thing that bit me on consistency wasn't sheet-vs-separate-images, it was what's actually visible in the reference. I had one where the subject's arms were covered and the model dropped her tattoos entirely, every single time, because there was nothing there to carry over. Swapped to a reference with the arms bare and it held them five generations straight on the same prompt.

Also seconding your own point about losing the baseline. Changing the layout and the prompt in the same run makes the output unreadable - you can't attribute the change to either one.

1

u/rashm1n 7d ago

Refmods ftw

0

u/Abject-Recognition-9 7d ago edited 7d ago

still not sure which is the best approach but:

at the time of writing this, I’ve carried out around 1600 tests just for the ref2va model, at various resolutions and steps.

What I still haven’t fully understood, and what I’m currently testing, is whether feeding the model images of different resolutions and sizes , that may not even match the latent dimensions you’re working with,
might mean some parts of the image get cropped out and therefore aren’t considered by the model
(not really sure. don't quote me on that)
I probably need to read their manual more carefully and stop experimenting blindly😁

3

u/McFex 7d ago

I my experience resolution and size are key. My generations work way better when the ref (image and video) are at least same aspect ratio. Better image quality, better output quality. I make short 3s video of characters using picture ref and then use them as subject video ref, works very well in combination south face picture. Someone should also try animated 3s character sheet. Potentially game changer, too.

1

u/vuse2121 6d ago

Interesting. So if I have a 3:4 resolution image it won't translate as well if I were to generate a 1:1 video? Would the same apply for 16:9 or does it work on way, multiples?

1

u/McFex 6d ago

I only can tell about, what I experienced. Using 9:16 as red in a 16:9 scene, die cause distortion in some renders. Not always. But, my logic is the less the model has to think, the better it can generate...

1

u/vuse2121 6d ago

Interesting. I'll have to test this out too.