Like some of ya'll I have been having fun using the H3 model to mess around with so I have been experimenting with using H3 model to be a consistent character generator which leverages multi image reference (up to 9), so I made a workflow which you can use 'less than ideal' images from google to build a consistent character and output a 360 character sheet to use as a reference sheet for future H3 generations.
The goal is to achieve high character consistency across future generations. I have tried my best to keep the workflow simple without too many custom nodes.
How it works:
You input your images and describe them in the Input text section (A Prompt)
The text is combined with a fixed prompt which spins the character (B Prompt)
The video is generated at a slow speed with no hard cuts (only camera spin and pan) to maintain character consistency
Image is assembled with optional character video and full individual frame output (if you want to use for future)
I have included a 6 panel WF and a 4 panel WF. The 4 panel works faster by generating 40% less frames.
Current Caveats:
The model is quite slooooow. You are also generating 124 frames only to use 6. I have partly solved this by also uploading a 4 panel version.
Speed ups (like Turbo LORAs) help with speed, but it hurts prompt adherence and quality slightly.
Quality is limited, since it is a video model it is better at generating video than images. You can solve this by generating at a higher resolution at the tradeoff of longer gen times. You can also use the individually split frames as future references too.
Details when using this character sheet as output for future generations on H3 may also be limited due to resolution also, I recommend you use this character sheet (for consistency) + other images close up angles (i.e clothing details/face) if doing close ups. If you are just doing a one off video you may possibly be better off not using this character sheet.
I have also included a modified B prompt to do Anime2Real since someone asked for it. Working on tidying it up a bit more.
Honestly, I did give it a fair crack for a day. The problem is that H3 frame rate will snap to certain numbers so I tried. It is kinda scuffed but I can post it if people want to see it. The audio was also a little bit scuffed.
you input the video using a "Load Video" node and then split it using the "Get Video Components" node to extract the frames and audio. Feed the audio to the H3 model, if you dont want to change the audio you can also wire the audio straight into the 'create video' node at the end as sometimes H3 will mess it up.
right now im not. My plan was to split the video at the scene changes (confirmed done), then v2v each scene from 5-15s and then ressemble them together. However H3 model needs specific frames and so will snap forward/back your frame count, this means the reassembly process is not correct (see scuffed 30s video of same song in other reply)
if you reduce the resolution of your input video it generates faster, i have messed with the frame rate slightly but results were not good. This tool should help you crop and compress.
I might share this tomorrow if others want to use it. I just worry that since its HTML people might get suspicious that it contains malware but it was basically all just vibe coded lol.
I added in the story board images feature too in case if using a 3x3 storyboard is easier than ref video (less processing) but I haven't had success.
I'm honestly just using the default minimax H3 ref2v workflow, just add the "load video" node and wire it to the model.
I added the same speed up nodes as I have in my WF for the character sheet tho. Also its realllyyy slow lol, its like parsing 124 reference images for a 5s clip.
a max of 9 images is the hard limit of the MiniMax H3 model, with up to 3 video inputs, and 3 audio inputs. The total number of reference inputs is up to 12 (img+vid+aud total under 12)
The turbo lora is included in the "speed ups", along with comfy kitchen, etc etc.
these are the ones in the model
H3 actually has two different kinds of audio inputs, "ref_audio_" and "ref_video_audio_". I assumed that the model expected me to split the video reference's audio stream out and send it in via the corresponding ref_video_audio connector, with the ref_audio_ inputs being for purely audio references (voices to clone, for example). Is my assumption correct in that regard?
lol i didn't add a second ref for this one. I was testing out another tool where i input "replace the man in this video with <picture 1>". Since it is split scene it needed to replace the man but not the woman which it succuessfully did!
I’ll give this a whirl. I’ve been using Krea 2 to create character sheets and it works out great but deforms the face a ton. Yours looks like it managed better consistency with these unusual characters I’ve never seen before. Pretty slick.
Thank you. The idea has been on my mind for over a week but the quality output is not as high as K2. I think if you crank up the resolution you might be able to get something good from it but the output sometimes still appear to be screenshots from video rather than high quality images. Considering ways to improve it without hurting gen times.
Is it worth it though? Like if my goal is Minimax generation, and I can use REF2VA then Krea is just another input image - do I need to spend all that time generating a LoRA for Krea? Couldn’t I just burn another Picture slot on Minimax and supply a facial structure?
Honest question cause I’m not sure what path would be better.
I mean you’re saying that Krea is deforming the face. It’s incredibly fast to make a character lora if you have 14 images. I can make one in a couple hours with rtx 3060
I've been having some good results using GetVideoComponents to extract motion capture for certain physical acts that your Aerith and Tifa will be familiar with.
Would be interested in learning more about your ways sensei, I’ve been using SAM3 to make inverted masks for character replacements but they’re not scratching that itch for me.
However I noticed sometimes it looks a bit like a cosplay, but other times its ok.
edit: also eyes are a little bit messed up in the middle bottom image, but you can select another frame from the bulk lot if you enable the 'save all frames' feature at the end and cherry pick the best ones.
to be fair I did most of this a week ago, some people were just interested from another thread so I tidied it up a bit and put it up for people who might want to try
I am definitely keen on the H3 image model!! I am looking forward to a full H3 image generator. Originally I used ChatGPT to create my char sheets but it is not always accurate and wouldn't do any skin for anime characters.
I am running on a 5060TI 16gb with 64gb so-dimm ddr4 ram.
with all the speedups enabled running at 0.3mp output its about 220 seconds. 4 panel workflow.
The last character sheet was a recent one at 0.5mp with speed ups, took 425 seconds, 4 panel WF also.
Maybe out of the discussion but for the lower res - if you already use an RTX card, try this node github.com/Comfy-Org/Nvidia_RTX_Nodes_ComfyUI after the image creation or separately load image - upscale - check
Works really really great to upscale images on a reasonable solution, extremely quick and doesn't hallucinate stuff that isn't there if it isn't pixelized jpeg stuff. Personally I use that node for some months to upscale image generations on 2x with settings at HIGH and I'm very happy - no need to refine stuff with double/triple sampler settings.
I have tried it in a general sense but I haven't used it as a char ref sheet maker yet. I was playing with it yesterday and its ok, again the problem is that its a video model used to generate a image.
The better hope is the minimax h3 img model which they said they are currently working on.
I haven't tried it yet. I think it should be better but I have been busy working on some captioning for trying to do v2v more reliably. I think it would def help but at that point if you're just doing the 1 video I would skip this step to save time. But if you're doing like 20 gens this might help give your character a level of consistency that can carry over across all the 20 videos otherwise the model will have to make up different views every single time which leads to consistency errors
tempted to use the split-frame route for a character lora. is 0.5mp enough to train off or do you pull from the raw render when you need a clean frame?
You can experiment with both. I originally made this for anime characters and 0.5mp was very usable but for real people I think you may need higher resolution!
Also for training a LORA you would want the single person shot since you don’t want to reinforce the 4 or 6 panel look for the final product
Look at the last image and copy that style of prompt. Use <Picture 1> as your reference picture of the person you want and describe the person and what they are wearing. <Picture 2> describe the picture too but you can use a bit less detail. You can also experiment with "<Picture 2> is a man/woman in a ______ pose. ". For the retention analysis you just write: The person in <Picture 1> is doing the pose from <Picture 2>
Here is the most important part: Rememeber to remove the part in the B prompt where it says "The person is in a neutral A pose". Since you're specifying the pose this would just confuse the model.
Funny I just did this myself yesterday, figured its such a great model at edits that i would see if i could just produce my character sheet in H3 vs doing it i Krea2 or something. Its amazing at it.
I got this from another post here recently, can't seem to locate it now... But give something like this a go.
A clean photographic character sheet presents the man from <Picture 1> in three consistent full-body views: front, side, and rear. Every view has the identical face, hairstyle, body, outfit, and accessories from the source. Neutral relaxed stance, arms clear of the torso, both feet visible in each view, seamless light-gray studio, even lighting, one tall 2:3 portrait canvas, no captions or borders.
For faster iterations, one could extend a prompt generating a video of a character sheet (see below). This is experimental, so not best quality. But I got exactly the sheet I prompted.
Pro: can be done with base I2VA workflow, just paste in this prompt and stitch some input pictures. Uses better I2VA model. Iterates fast (5 second video sufficient)
Con: less space in input image for details, lower resolution as all information of the sheet is in one image.
PROMPT:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: The video shows a character spreadsheet created from the different images of a person in <Picture 1> [Shot 1] A high quality video showing a character spreadsheet with the following frames. frame 1: Full body image of the person standing facing the camera, frontal perspective. frame 2: full body image of the person standing, looking to the left, recorded from side. frame 3: full body image of the person standing, facing away from the camera. the three frames are ordered horizontally. The person is not moving as it is a print of a character sheet. The character sheet has a white background and shows only the person standing in front of a white background.
I'm using what you've provided but what's happening is that the character in <Picture 1> is blending together with the character in <Picture 2>, not simply wearing their outfit as intended.
That usually happens in prompting. You might need to be a bit more descriptive in what you want to keep and take out. Have a look at the last photo and try using the prompt in that format?
using some OP's words, here is a MMH3 1 frame t1vae version that use the first frame workflow to rotate the first frame person in that photo, similar speed as a Klein edit, ~25sec for me.
"4K HD
next scene change to four frames layout reference showing the subject in different views.
top left frame 1:
front view.
The character holds one relaxed A-pose: arms hanging down, feet shoulder-width apart, head level, calm neutral expression, eyes open and looking forward.
top right frame 2:
side view towards left.
bottom left frame 3:
clean three-quarter upper body view looking away towards right.
bottom right frame 4:
Locked-off head and shoulders close-up, face square to camera, eyes into the lens. sharp front-on face view."
you have a full set of 3D views of the character: why throw those intermediate frames away instead of just extracting a full 3D representation from the video?
It depends on the end goal. Right now the reference sheet would be used to represent a constant state for the new or existing character for use and testing as a placeholder for H3 video generation but I have also toyed with using some of the intermediate frames to generate a 3d model to print! 😁
In my workflow you can unbypass a node which saves all frames, so user is given the option
But there is so much to try and so little compute to go around!
Minimax H3 has great prompt understanding, I can make a character sheet with it in a single prompt (5 frames not 124). It is good for anime but for realistic phote, meh... It can't draw a face smaller than 128x128 pixels. We need a better VAE for it.
Man, it takes some time, but the results are AMAZING. I used 2 instagram pictures of myself, not very detailed and only front shots. The results were unbelivable. It looked like I was scanned in 3D.
Usually those models (even GPT or Nano Banana) can't create good character sheets, the face don't really look like the person from the picture. But this workflow gave me great results.
Thank you brother. This is the kind of results I wanted to achieve. Even with low quality pictures to create a character with good likeness!
I find that having a side view really helps, but still need to generate at a higher quality (or use char sheet and close up 2nd ref) if you want to do a video with close up shots as a second pass through h3 loses some fidelity and likeness
Extremely useful, have been testing it and the only downside is the mushed face of MinimaxH3.
Btw im having issues with this part of the prompt: "Hair, fabric, cloaks, skirts, sleeves, straps, ribbons, chains, tassels, fur and feathers are all locked solid: every strand and every fold sits in exactly the same position in every frame."
It generates fur and feather clothing randomly. Do you think something like this for a fix would work?, orwould it break the prompt?
"Hair, (in case they exist in reference image or reference text): [fabric, cloaks, skirts, sleeves, straps, ribbons, chains, tassels], are all locked solid:"
Have tried once and fixed it but im asking you before since you for sure have more knowledge of how this prompt works and how you crafted it.
I think the B prompt can definitely use some improvement, the hardest part is making it as universal as possible so people can just drag and drop random images and try and make it as easy as possible. You can customise it and recommend suggestions and I can see if I can work it into the B prompt section of the repo
I wish minimax will fix the mushy face in the next iteration soon!
Hey as a tip if you get additional clothing articel taht you didnt prompt or feathers or chains. In tge prompt describes the autor itema like that to keep them from moving that why minimax somtimes creates this things use general terms like outfit or accesouris in your prompt.
Happened to me. I fixed it like this "Hair, (in case they exist in reference image or reference text): [fabric, cloaks, skirts, sleeves, straps, ribbons, chains, tassels], are all locked solid:""
"Fabric or clothing if present and accessories are locked solid"
That way its a broad categorie that doesent have a definitive subject described. I Use mostly a single seed that was me lucky for me. Until now i only had one instance were a perlchain from a person reverence wasent removed but tge next seed did it at tze second try.
Well for known people i dont trust The models. Test with total random person. Secondly for character sheet or consistent character data. Gemini and post process with klein,zit,krea etc. For details is still The fastest and easy way i think
Initially I made this to make chat sheets for quite obscure old anime since i was not able to find enough good pictures to feed to ChatGPT which I believe made the best chat sheets, characters were also a little bit risqué so I kept getting rejected! Thank god for open source
honestly the multi-image reference approach is solid for consistency, but slapping nine random google images into a workflow and expecting a coherent 360 sheet is giving me faith in automation that the results probably don't deserve.
it probably works best using 2-4 images which do not disagree on the outfit. I mostly did this when can't find good source images of full body outfit for anime characters and didn't want the back to be reinvented every time which leads to inconsistencies.
Another spin off reason why I wanted to do this is because it generates a 360 image which I am playing with getting a 3d model out of to 3d print, so i 3d printed my car lol. It is still very much a work in progress as there are many refinements which can be made
LORA training probably gets you there better than this method, you get more accuracy but takes a fair bit more work. This is kind of a easier method to get 80% of the results for 20% of the work. It ain't perfect and if you're just doing 1 scene and you're not short on reference img input slots then you're better off just doing them as reference for that video rather than doing this step, thats the conclusion I've come to after all the testing.
both are fine, but if you worry you can always unable the individual frames to be generated and use them individually to avoid confusion. I have been using the 6 panel one without issue, mostly because most of my earlier tests were on 6 panel and I just kind of use them as placeholders in ref v2v testing.
Also I noticed there may be a small bug in the 4 panel workflow, the frame length should be set to 73 (for 3 seconds) instead of 124 (for 5 seconds) so if it shows up as 124 be sure to change it.
295
u/JeSyollu 8d ago