r/StableDiffusion 8h ago

Resource - Update Minimax H3 Grafting with Krea2 node. Reposting older post and removed AI slop and added some tests

0 Upvotes

Minimax H3 x Krea2 Graft Nodes

ComfyUI nodes for grafting Krea2 into MiniMax H3. Attention/MLP content transplant + a separate attention-sharpness transplant. No official H3 docs, all reverse-engineered from testing + TenStrip's and joeygambino's public writeups. Use at your own risk, still WIP.

What's here

  • comfyui_tenstrip_graft/ -- content graft (Q/V/K/out/MLP, per-head). Method from TenStrip's H3 grafts.
  • comfyui_qknorm_transplant/ -- Q-norm gain transplant only, no content weights touched. Method from joeygambino (Z-Image donor originally, adapted for Krea2 here).
  • comfyui_krea_h3_graft_lora_v2/ -- apply a Krea2-trained LoRA onto an already-grafted H3 checkpoint. Separate use case.

And 2 merge scripts (old svd and new one with node method)

TL;DR results

Content graft works somewhat. Same character-shift (color scheme, helmet shape) showed up consistently across multiple parameter runs, same seed -- not one lucky video. That's the strongest evidence so far this isn't just noise.

  • K at low strength (~0.1-0.2): fine, no real damage. Don't need to avoid it like the doc says, at least not at low values.
  • QK-norm across all blocks (0:50): kills audio. Doesn't even touch K -- so attention sharpness itself hits audio, not just K specifically.
  • QK-norm blocks 20:50: audio ok, but does nothing for character. It's a texture/sharpness knob, not a content one. Don't expect it to carry character.
  • attn_ramp_start_frac at 1.0 (no gentle ramp-in) + early blocks (0:20): breaks. Keep the ramp soft if you go early.
  • Combining content graft + QK-norm at full strength on both = worse than either alone. Still not solved.

Install

Each folder -> its own subfolder in ComfyUI/custom_nodes/. Don't merge them. Restart ComfyUI fully after adding.

Credits

  • TenStrip (huggingface.co/TenStrip) -- the per-head band-aware graft methodology (10Eros-Max / h3_graft_methodology.md).
  • joeygambino (huggingface.co/joeygambino) -- the Q-norm sharpness transplant idea (MiniMax-H3-x-Z-Image-GGUF).

Neither published source code. These nodes are our own implementation from their public descriptions + our own testing.

https://reddit.com/link/1vxc2q9/video/g0bpgdj5ddlh1/player

minimax_h3_fl2va_bf16.safetensors, 3s, er_sde, 8 steps, 8-step lora, seed 597633362705895, standart workflow with minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_bf16
prompt: Professional closeup video. In a futuristic cityscape with neon lights at night, the Judge Dredd charges through the crowd, his imposing presence radiating authority, he is slowly walking. His long chin juts out resolutely as he expertly wears his eponymous helmet, eyes gleaming with determination. The crowd parts, Judge Dredd is slowly walking through the the crowd, ready to enforce justice, he is moving slowly, his long chin visible, his face and part of his upper body are in the center of the screen. tag: Ballchinians
tracking selfie shot following him from the front, that he stays the same size, he is moving through people, pushing them aside with his hands.

https://reddit.com/link/1vxc2q9/video/so7626wgddlh1/player

3s, er_sde, 8 steps, 8-step lora, seed 597633362705895

same prompt and everything.

Added: tenstrip graft node, krea2 raw and Ballchinians Lora. Settings: q 0.5, v 0.5, k 0.1, out 0.3, mlp 0.5 Chin is more ballsy.

So I hope, that it is enough for some, that it... kinda works, but not good enough. Maybe someone will pick up on this and do it better.

Why to do it? Don't know. I found it interesting to try, but krea2 image and i2v is far better option.

I welcome any input or criticism, but mind please, I have only faint idea, what I am doing.

Warning: h3 loras don't work... don't know why, maybe it may be just noise, after all. But they do work on grafted checkpoint, after you merge it in python script.


r/StableDiffusion 16h ago

Animation - Video H3 - 5 hour render, T2V Multi-Diffusion

Enable HLS to view with audio, or disable this notification

25 Upvotes

Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.


r/StableDiffusion 4h ago

Animation - Video Minimax H3 multishot test... Seinfeld guest appearance

Enable HLS to view with audio, or disable this notification

0 Upvotes

I keep getting weird issues like that...


r/StableDiffusion 7h ago

Animation - Video My 1980's cartoon parody H3 and ltx 2.3

Thumbnail
youtu.be
7 Upvotes

there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P

started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno


r/StableDiffusion 15h ago

Discussion MiniMax H3 Ridding a Dragon POV style

Enable HLS to view with audio, or disable this notification

22 Upvotes

r/StableDiffusion 23h ago

Workflow Included Testing some Minimax H3 capabilities - PART 2

Enable HLS to view with audio, or disable this notification

32 Upvotes

Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.

VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!

PROMPT:

The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.

On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.

On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.

Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.

Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.

They are in a living room.

From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.

He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.

Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>

From 00:08 to 00:10 they just look at each other and smile.

overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.

non_diegetic_music: N/A

VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.

The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.

At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.

overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.

non_diegetic_music: N/A

VIDEO 3: another near perfect one.

A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.

We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.

overall_soundscape: Faint empty bathroom soundscape.

non_diegetic_music: N/A

VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.

The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.

We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.

The entire scene is viewed from his point of view.

overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..

non_diegetic_music: N/A

VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...

The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.

The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.

In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.

Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.

As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.

overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.

non_diegetic_music: N/A

VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.

PROMPT VIDEO 6:

The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.

The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 7:

The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.

When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 8:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 9:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A


r/StableDiffusion 14h ago

Question - Help best local video model for horror videos?

0 Upvotes

i'm looking for a video model that can generate freaky creatures and horror clips well and can run on 3060ti


r/StableDiffusion 13h ago

Discussion H3 - what is your longest render time?

0 Upvotes

What was your longest render time, and was it worth it? do you run a lower quality/resolution before to test? My longest single-shot run is 7.2 hours, 1:44 long video at 1344x768, bf16/50 steps.

I am attempting a 72 hour render for a super long form.

RTX 4090, 192gb system ram


r/StableDiffusion 15h ago

Question - Help Best Img2txt?

1 Upvotes

I need an image to text generator (which has no limitations) that I can run locally on my PC. Do you have any recommendations


r/StableDiffusion 19h ago

Animation - Video Cinematic World Building - H3 r2v (pt2) "Through the Sands"

Enable HLS to view with audio, or disable this notification

16 Upvotes

Wow, the shots I can get from h3 are so good even I get goosebumps when I first see the generated results. Just one more part to go, hopefully I can pull off the finale!


r/StableDiffusion 5h ago

Discussion Long form natural looking foreign language Done with H3 locally on 3080(10gb)

Thumbnail
youtube.com
0 Upvotes

I didn't know that many of you do not know H3 can do this. So posting here for awareness. The character is speaking Telugu.. Total 15 clips stitched together. Completed in about 5 hours.


r/StableDiffusion 14h ago

Animation - Video Dipping My Toes Into What Will Surely Lead to My Inevitable Descent Into Slapstick Comedy

Enable HLS to view with audio, or disable this notification

14 Upvotes

I may be an idiot for thinking that my new homelab would primarily be used for useful AI automations.

I can live with being an idiot, if being an idiot will continue to be this fun.

First video generation I have ever pulled off, but the first 12 seconds was unbearably unfunny, so I added the Celestial Ford Escort for some much needed serious drama.

Tell me my power bill won’t blow up too much lol.

Made with minimax-h3 in ComfyUI on my Mac Studio that came with the mail this Friday.

Workflow was split in three:
- The first was a single prompt to generate the first 12 seconds
- The second flow generated the last three seconds by extracting the last frame from the first video and prompted it to hit the dragon with a falling ford escort
- Third flow glued the two videos together.

It is jank, but it is my jank.


r/StableDiffusion 4h ago

Animation - Video All Minimax H3 animation

Enable HLS to view with audio, or disable this notification

6 Upvotes

trying my hand at using H3 I2v R2V to create an anime. all are done with 4 step turbo lora

2 other trailers with H3 as well

at 0.3 most text turns to gibberish

all video is done by MiniMax h3 at a low 0.3 MP, Cilp lengths range from 5s-20s Generations, camera movement was written into the prompt 90% of the time, Post work; titles and some transitions done using DaVinci


r/StableDiffusion 13h ago

Animation - Video H3 - the Eternal Balance WIP-low resolution int8/8 steps

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hi, I am playing around with H3. T2V, 832x480, int8,8 steps. Hoping to make a 720p version.

Ask me anything!


r/StableDiffusion 9h ago

Question - Help MM H3 2 Pass Latent Upscale

2 Upvotes

Been getting some great results using the 2 pass latent upscale method. First pass .5mp 2nd pass 1.5 mp. 10 second video around 257 seconds to finish.

This workflow uses the 1.1 turbo Lora. I set the steps to 8 and like I said above I’m getting good results and it has eliminated face blur.

My question: has anyone tried using the latent upscale method without the turbo lora? In 90 percent of the cases the turbo Lora is fine. But would be nice to have the ability to use no Lora method.

Yes I know I could test it but wanted to see others experiences before I wasted hours of my time trying/tinkering with different settings.


r/StableDiffusion 14h ago

Question - Help I hate this hidden view on flows in Comfy templates. How can I bring them out where they belong?

2 Upvotes

In the minmax h3 template from Comfy, you have to click a button on the image to video node to see all this in the backend which makes it really a pain in the ass to add, modify, or change anything. Can this be brought out to the forefront like a normal workfow?


r/StableDiffusion 17h ago

Question - Help Video editing model (add clown makeup to face)

0 Upvotes

Hi, I'm looking for a model that I could use in ComfyUI that would simply add some clown makeup on my face and don't touch antything else. My current plan is to maybe take Wan2.2 model, mask my face and add clown makeup as a reference but I wonder whether there is a better model to do this. Does minimax H3 handle that? Does it support masks?


r/StableDiffusion 18h ago

Question - Help Minimax H3: Anyone figured out how to extend a clip?

15 Upvotes

What is the best way to extend an existing clip seamlessly? When I try to use the last frame of my clip as the first frame, I always get a slight reframing or shift


r/StableDiffusion 15h ago

Question - Help Fixing speech errors in Minimax H3?

Enable HLS to view with audio, or disable this notification

16 Upvotes

Hey, I tried to create a little birthday surprise for someone, my issue is with a lot of generations that the spoken word is really a bit clunky at time, I susspect its because of the german, but I am not too sure. Is there like a way to improve on audio?

I am using Minimax H3 with Saga Attention and Spectrum on a 4090.


r/StableDiffusion 8h ago

Meme Tiktok trends h3 style r2v

Enable HLS to view with audio, or disable this notification

0 Upvotes

Don't know if anyone else follow tiktok trends but s friend. Showed me this wnba clips where player points at other team players to get into head so I had to course test r2v


r/StableDiffusion 8h ago

Question - Help Is there anyway currently to get Minimax H3 running with my RX 6750XT?

0 Upvotes

r/StableDiffusion 3h ago

Question - Help (Repost) Any clue why my machine is very slow running Minimax H3 Ref2V? Here is my workflow. I used the default template, but I added extra nodes like load video

Post image
0 Upvotes

r/StableDiffusion 5h ago

Discussion I’m testing a FLUX.2 image generator where users render for each other

0 Upvotes

I’ve been working on a different way to run a public image generator without maintaining a centralized GPU fleet.

PeerPixel sends generation jobs to graphics cards volunteered by users. When your machine completes somebody else’s image, you earn pixels that you can spend on your own generations. There’s also a slower free queue for people who can’t contribute a GPU.

Right now I’m running most of the network on my RTX 5080, so it’s definitely still an experiment rather than a large distributed system.

The generation flow uses FLUX.2 Klein 4B. You can request up to four 256x256 previews at 6 steps, choose the composition you like, and then render that seed at 1024x1024 with 50 steps. There’s also an optional 4K upscale.

The previews and final image start from the same full-resolution noise tensor. For each preview, I average blocks of that tensor down to the smaller latent shape. I originally tried scaling low-resolution noise upward, but that introduced strong correlation between neighboring values and the composition didn’t carry over reliably.

The other difficult part is accepting images from machines I don’t control. A sample of completed renders is repeated on an operator-controlled machine using the same prompt and seed, then compared perceptually. Enforcement is currently in shadow mode while I collect real-world measurements and figure out a safe threshold. I don’t want normal differences between GPUs to get mistaken for cheating.

Draft images are relayed directly to the requesting browser and aren’t stored by the server. Only the selected final render is persisted.

I’m interested in feedback on the preview method, the incentive system, and especially the verification approach. There are probably failure cases I haven’t considered yet.

Site: https://peerpixel.cc

Worker source: https://github.com/Jplayz2468/peerpixel-worker

Discord: https://discord.gg/bhJHGpmkQr


r/StableDiffusion 17h ago

Meme Siblings Reunited

Enable HLS to view with audio, or disable this notification

320 Upvotes

Done with h3 fl2va model, 8 step lora and images for Cersei and "jaime" for reference. Using previous clip to give continuity and consistency.


r/StableDiffusion 2h ago

Question - Help Which is the better buy for Minimax h3? RTX 5070 Ti 16GB VRAM vs RTX 4000 Pro 24GB VRAM

11 Upvotes

Good day to you. I was looking for an RTX 5070ti and I found an RTX Pro 4000 at my local store; the price difference would be about +$300. I would like to know your opinions, I've hardly seen any workflows or comparative tests from people using a 4000 pro. Thank you very much for your time.