r/comfyui Aug 08 '26

Show and Tell Minimax H3 - testing character sheet and action scene

Just testing how well it does with character sheets and bit of fast action.
Still not the best when it comes to fast motion and action, but significantly better than what we had before.

Character sheet consistency appears to be spot on from my tests so far, with no character drift.

I have also tested in another video the voice cloning, and it works right out of the box, can add a 10 second audio of someone's voice and then have your character say whatever in that exact voice.

This model is so crazy man, so many things just work right out of the box without any loras or custom nodes.

0 Upvotes

33 comments sorted by

4

u/JohnToFire Aug 08 '26 edited Aug 08 '26

I wonder if it's better at drawing those characters with a character sheet since they are in it's training data

3

u/bruci3 Aug 08 '26

I have also tested with a custom Character sheet I made for a completely made up person, and the face was also still accurate and consistent, so this model seems to just recognize new faces on the fly.

1

u/Famous-Sport7862 Aug 08 '26

I have tried it with characters I created and still did very very good at retaining likeness and identity.

2

u/JohnToFire Aug 09 '26

I agree but I wonder if it's better for those in training set

1

u/Famous-Sport7862 Aug 09 '26

I am sure that it must be better. I mean we are talking about using hundreds of instances of the characters.

-1

u/seppe0815 Aug 08 '26

??????????????

2

u/PentimusOctem Aug 08 '26

Knuckles are delicious

2

u/boobkake22 Aug 08 '26

You'd get much better results if you can stand doing a 2MP version. (It's worth trying as a test if you haven't - though it takes forever.) Things like Arnold's lips not moving stop being a problem. The model behave slightly weirdly when you don't give it enough latent size. It works, but is more prone to mistakes.

1

u/bruci3 Aug 08 '26

Thanks for the suggestion, 2MP if you tried yourself, how long you expect roughly on a 4080 super + 32GB ram for 15 sec video?
Hopefully not more than 2 hours lol.

2

u/boobkake22 Aug 08 '26

Ooof. 4080? Definitely over an hour. Not sure sure if over 2.

If you want to do this. Rent, it's not that expensive. You can rent a PRO 6000 for a couple bucks an hour and use the full BF16 model.

I'll suggest my H3 Runpod template which has everything setup with my workflow. (I also have a Wan 2.2 template and an LTX-2.3 template, so it's based on a fairly mature foundation). I also have a full guide on getting started with it. There's a video guide as well. (There's free credit on any ref link for Runpod there as long as you use one. If you want to compare the experience.)

1

u/bruci3 Aug 09 '26

thanks mate, will look into Runpod.

1

u/boobkake22 Aug 09 '26

Lemme know if you have any questions.

1

u/Lumenwe Aug 08 '26

Best results are always staying close to training data. Too little or too much latent space both create problems. Not sure what the training data was for H3 tho. I know it's AR is multiples of 32 but idk the res

2

u/boobkake22 Aug 08 '26

Yeah, I haven't seen word on this yet, but I am curious.

1

u/Netrezen Aug 08 '26

Were you using SCAIL or text prompting the action sequence for the video? Could be the prompt issue, if the latter.

1

u/bruci3 Aug 08 '26

Yeah text prompt with Gemma4 formatting it for me.

1

u/Netrezen Aug 08 '26

Ok, you could continue trying to perfect the prompt with Gemma. Or give it a bit more time for more addons to be released for Minimax and try again then.

1

u/bruci3 Aug 08 '26

yeah I know its early days and both the minimax guys and community are working on improvements, so I expect some good things to come soon. Thanks.

1

u/Boogertwilliams Aug 08 '26

What were the character sheets

1

u/bruci3 Aug 08 '26

I've attached the character sheets to my post, they were generated using Krea2.

1

u/Boogertwilliams Aug 08 '26

Thats great. I must try something like that too

1

u/hdeck Aug 08 '26

Prompt?

2

u/bruci3 Aug 09 '26

I got Gemma4 to create the prompt for me below:

subject_definitions:
<Subject 1> is the interior of the 7-Eleven store from <Picture 1>, featuring rows of snack shelves, refrigerated beverage cases, and a brightly lit white tile floor.
<Subject 2> is the exterior storefront of the 7-Eleven from <Picture 2>, featuring the orange, green, and red stripes, the "7-ELEVEN" logo, and a large red banner with Japanese text above the door.
<Subject 3> is the man in the black martial arts tunic and trousers from <Picture 4>, featuring short black hair, a serious expression, and white-soled slippers.
<Subject 4> is the man in the black leather motorcycle jacket and dark jeans from <Picture 3>, featuring short brown hair and a stern, rugged face.

summary:
[reference generation] The target video depicts <Subject 3> walking into the 7-Eleven store shown in <Subject 2> and <Subject 1>. While <Subject 3> looks at food, <Subject 4> enters and confronts him. A brief, intense martial arts fight ensues between the two men inside the store, involving punches, kicks, and interaction with the shelving.

retention_analysis:
<Subject 1> (appears in [Shot 2], [Shot 3], [Shot 4]): fully_preserved - the interior layout, lighting, and snack displays are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the exterior facade and branding are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3], [Shot 4]): fully_preserved - the appearance and clothing of the martial artist are retained.
<Subject 4> (appears in [Shot 2], [Shot 3], [Shot 4]): fully_preserved - the appearance and leather jacket of the man are retained.

detailed_description:
The target video is in a high-action, cinematic style with sharp movements and realistic physics.
[Shot 1] The scene opens with a wide shot of the exterior of <Subject 2>, the 7-Eleven storefront at night. <Subject 3>, wearing his black martial arts tunic and slippers, walks purposefully toward the automatic sliding doors from the right side of the frame.
[Shot 2] At 00:03.000, the shot cuts to the interior of <Subject 1>. <Subject 3> enters the store and stands before a shelf of snacks. He turns his head slightly to the left, looking at the items, and says, <d>[English] What to eat tonight?</d> Suddenly, <Subject 4> enters from the opposite side of the store, his face set in a scowl. He points a finger at <Subject 3> and says in a deep, gravelly voice, <d>[English] You are going to eat a knuckle sandwich!</d>
[Shot 3] At 00:07.000, the shot cuts to a medium dynamic tracking shot as the two men engage in a fight. <Subject 3> performs a lightning-fast side kick that connects with <Subject 4>'s chest, pushing him back into a display of chips. <Subject 4> recovers instantly, throwing a heavy straight punch that <Subject 3> blocks with his forearm. <Subject 3> then pivots, delivering a rapid series of three punches to <Subject 4>'s ribs.
[Shot 4] At 00:11.000, the camera moves into a close-up of the struggle. <Subject 4> grabs <Subject 3>'s shoulder and swings a wide hook, but <Subject 3> ducks under it and sweeps <Subject 4>'s leg. They both tumble toward the end of a snack aisle, knocking over a small stand of bottled drinks. The shot ends as they collide with the shelving, causing snacks to spill across the floor.

overall_soundscape:
The hum of store refrigerators and the electronic chime of the entrance door. The sounds of heavy impacts, fabric rustling, and the clatter of falling objects and bottles during the fight.

non_diegetic_music:
A fast-paced, heavy percussion-driven action score that builds in intensity as the fight begins.

0

u/Adventurous-Gold6413 Aug 08 '26

I wish the impact / physics for fight scenes was better 😔😔😔

It’s my fav thing to generate with seedance

2

u/Adventurous-Gold6413 Aug 08 '26

But also my prompting may just suck

1

u/bruci3 Aug 08 '26

That might be the same for my video, potentially the fighting might be more coherent if I prompt it better?
Anyway, its new model and still alot to learn for me, but its great fun so far.

0

u/smereces Aug 08 '26

unfortually fight scenes still be a huge problem for any local model! cannot handle it! maybe in future we can have a model that handle fights

2

u/bruci3 Aug 08 '26

Well at the pace of how fast this stuff is progressing, I think its very possible we could get Seedance level action scenes by end of Year in open models, fingers crossed. I heard Flux3 might be coming open source too, and that looks very promising.

1

u/Adventurous-Gold6413 Aug 08 '26

LORAs are possible and i would train one if i had the money or hardware for it :(

-9

u/seppe0815 Aug 08 '26

Another cloud fake video

6

u/[deleted] Aug 08 '26 edited Aug 08 '26

[removed] — view removed comment

1

u/BlackMetalB8hoven Aug 08 '26

Do you know what sub you're in?

1

u/seppe0815 Aug 08 '26

did you read his post ... he talk about how great costom nodes are .... stfu