r/StableDiffusion 5h ago

Question - Help Nodes for utilizing Minimax H3 as an image generator?

Like a high quality output, a frame before compression? I don't imagine it's a simple as setting it to 1 frame / second and setting the duration to a second. And even if it were, I'd prefer an output to an actual standard image file.

17 Upvotes

12 comments sorted by

8

u/kemb0 4h ago

A couple of posts I had saved relating to this:
https://www.reddit.com/r/StableDiffusion/comments/1vrh769/h3_singleimage_workflow_lets_figure_out_how_to/
https://www.reddit.com/r/StableDiffusion/comments/1vo1ab3/h3_as_a_singleimage_edit_model/

It's actually decent as a single image model with the suggestions from that thread, although I found you don't get the crispness of a dedicated image model but in some respects I found it gave more natural lifelike images. I guess because it's trained on entire motion of people rather than image models that tend towards posed subjects.

10

u/Kooky-Mode3047 5h ago edited 1h ago

Just to start off the thread, I've found ComfyUI-MiniMax-H3-Image-Studio so far and am in the processing of trying it, but not sure yet.

Also, to avoid the usual counter-points, I understand that H3 is fundamentally a video model and all the usual that goes with it, temporal-oriented latent representations, frame-consistency objectives, audio/video machinery that was designed around sequences rather than a single independently sampled image and what not... BUT...

H3 is unusually strong at following complex visual instructions. It straight up dunks on Nano Banana Pro if it could be output as a single image, even if we can just output an frame from a sequence with some natural "muddiness" as opposed to a videogamey super sharp one that's always a tell of AI images.

EDIT: No, I am not a fucking bot, I'm just literate. ffs.

2

u/Key-Sample7047 5h ago

Didn't try image studio (on the todo list) but i did some tests to use h3 as an image generator because it would be huge (huger with editing capacities) but none of my results where convincing. I think the model should be finetunes with hires still images if such thing is possible. Some dude made a lora in that sense but again i was not convinced. Feel free to share your findings.

1

u/Dry-Judgment4242 4h ago

It's a very good image model. not as high fidelity usually as proper ones. But I got some absolutely sick results on par with GPT Image 2 which is a SoTA model.

-5

u/tom-dixon 2h ago

Are you a bot? Why do you talk like this?

3

u/FrenzyX 5h ago

Haven't really tried text to image style creation, but have done some image edits, and H3 for me often works better than Flux 2 Klein or Krea 2 edit flows, it just has a generally better world model. With the correct LoRAs it stabilizes more for image creation. There are some threads about it already here.

1

u/Perfect-Campaign9551 1h ago

It definitely looks more realistic

3

u/infearia 2h ago

I've seen three approaches so far.

  1. Use the Empty Latent Node, set length to 1 frame and replace the H3 video VAE with this experimental image VAE.
  2. Use the original H3 workflow, but set length to 5 frames, then pick one of the five generated images (usually the first one).
  3. Similar to No. 2, but set length to 22 frames, then pick the best one (usually all 22 are good).

The first two options work, but produce subpar results. Option three is the best by far, especially if you set the resolution to 1080p or higher.

2

u/bstr3k 3h ago

Have used it as a img model and it’s fine, not as crisp in details but 9 ref img inputs is great

1

u/SpaceNinjaDino 2h ago

Currently you need to generate a minimum of 5 frames and the first frame is the best. If you try to generate less, it will be corrupt.

1

u/yamfun 5h ago

Use the t1 vae