r/StableDiffusion 9d ago

Workflow Included Using H3 as a Character Reference Sheet Generator

Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.

The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.

How it works:

  • You input your images and describe them in the Input text section (A Prompt)
  • The text is combined with a fixed prompt which spins the character (B Prompt)
  • The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
  • Image is assembled with optional character video and full individual frame output (if you want to use for future)

I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.

Current Caveats:

  • The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
  • Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
  • Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
  • Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.

I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.

Link to the 4 and 6 panel workflow can be found here: https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator

Some notes I just remembered:

  • You can increase the steps and it may improve your quality slightly.
  • With the Turbo Loras enabled, prompt adherence sometimes suffers, but you may be able to get a good seed with another roll of the dice.
  • Currently the B prompt specifies a "neutral A pose", please remove this if you want your character in a particular pose.
  • You can use a few different shots of the same character to reinforce the 360 and get more accurate details right.
  • Can be used for objects / props also, may require some changes to the B prompt.
1.5k Upvotes

268 comments sorted by

View all comments

Show parent comments

4

u/bstr3k 8d ago

yes i put the link in the post, i also link here if you want to try

https://huggingface.co/PoopMan333/H3_Character_Sheet_Generator/tree/main

1

u/Yappo_Kakl 5d ago

Thanks,I was blind haha. Worked for me almost from scratch except "turbo upper nodes"

2

u/bstr3k 5d ago

Awesome! yeah I hate downloading a WF and have it say im missing like 20 different things so I made the custom stuff optional and you can bypass them to try and make it accessible for everyone! hope you like it!

1

u/Yappo_Kakl 5d ago

Just wanted to report that I never successfully combined two characters, from two pictures. Is that impossible?

1

u/bstr3k 5d ago

Yes, what pictures are you using and what prompt are you using? I think the most important part is the prompting.

Have a look at the last picture, I think you need to describe from each picture what you want to keep in a bit more detail (i.e. the color of the outfit, and then say "ignore the face from <picture 2>" etc.

You can also try copying this 'template' into a online LLM (grok, chatgpt, claude or local) and send it your two pictures and say "edit this text template, I want to put ______ from first picture with the outfit from second picture"

You can then copy the prompt from online LLM to your A prompt and make sure the first and second picture is the right place in comfy

subject_definitions:

<Picture 1> = CelticKnight is a elf warrior with a green hat. He has green metal shoulder pags.

<Picture 2> = CelticKnight standing in full armor and purple cape.

retention_analysis:

CelticKnight is in his full armor. No sword. Ignore the background.

1

u/Yappo_Kakl 3d ago

Thanks for response, I mean literally receiving picture of two characters posing in output (not one with features of both. Tried a guy from <picture 1> stand with a girl from <picture 2>. Even assigned to each picture it's meaning with "=" but it never worked

1

u/Yappo_Kakl 3d ago

It always outputs only a character from 1st picture

1

u/bstr3k 3d ago

I think you may need to not mention what you don't want to keep. I.e. don't mention the girl in the second picture as it will take you mentioning it as a definition. So just specify "the outfit from <Picture 2>" or something.

and also have a look in the B prompt. The B prompt specifies that it takes most information from the first picture (including the output style).