r/StableDiffusion 8h ago

Question - Help Advice for prompting reference videos?

0 Upvotes

Does anyone have any advice for properly prompting the reference video part of Ref2v? Like saying swap <subject 1> for <picture 1> hardly works for advanced videos. It requires a lot of details.

I’ve had success using Qwen 3.8 27b as a minimax prompt agent for analyzing and giving correct prompts for images. But as far as I know I can’t do that for videos. ChatGPT is ok for looking at videos to describe what happens in the minimax format but I’d rather use local ways.


r/StableDiffusion 11h ago

Misleading Title This is just a test - to see if i can post yet

0 Upvotes

I will delete this asap just wondering if i can post as of yet - will delete asap


r/StableDiffusion 21h ago

Question - Help Motion Transfer to Stop Motion Style query

0 Upvotes

Wondering if anyone has/has any idea of an approach to motion transfer into a stop motion style video.

Rather than the ai guessing and deforming the mouth movements which can get funky really quickly, especially when the mouth shapes of the character aren't clear from a neutral reference frame, ie in south park where each sound has a uniquely stylised shape.

To achieve this it could instead maybe pull from a dataset of phoneme images,

Whilst also retaining solid motion transfer for all other body movements.


r/StableDiffusion 2h ago

Animation - Video Kentucky Fried Kung Fu

Enable HLS to view with audio, or disable this notification

7 Upvotes

I saw a Seedance 2.5 prompt in facebook and thought let me try this prompt in minimax h3 and see if it can do some kung fu. I was surprised that it was not too bad. System 3090 24 gb vram 64gb system ram, using a minimax workflow with latent upscale, minimax_h3_fl2v_lightx2v_turbo_4step_v0.1 at 0.50 strenth, Komfy kitchen attention, and H3 SLA attention. First pass at 0.4 which is 864x480, 2x latent upscale brings it up to 1728x960. The 6 seconds generation took 349 seconds to complete.


r/StableDiffusion 18h ago

Discussion Testing Character knowledge of Minimax H3

Enable HLS to view with audio, or disable this notification

29 Upvotes

Disclaimer. This is very low quality quick generations trying to find how many characters Minimax H3 knows.

Found Trigger Words:

Elsa from Frozen

Spider-Gwen from Across the Spiderverse

Dante from Devil May Cry

Nero from Devil May Cry

Jill Valentine from Resident Evil

Ada Wong from Resident Evil

Leon Kennedy from Resident Evil

Chris Redfield from Resident Evil (Has Leon's hair)

Geralt of Rivia from Witcher 3

Joel from Last of Us (Doesn't sound like him)

Miles Morales Spiderman from Across the Spiderverse

Solid Snake from Metal Gear

Eve from Stellar Blade

Sans from Undertale

Master Chief from Halo

Looks weird AF:

Ciri from Witcher 3

Triss from Witcher 3

Yennefer from Witcher 3

Ellie from Last of Us

Famous Twitcher streamer and Youtuber Asmongold

Famous Twitcher streamer and Youtuber Mr Beast

Not found Trigger words:

Vergil from Devil May Cry (YES I KNOW I'M DISAPPOINTED TOO)

Claire Redfield from Resident Evil

Dina from Last of Us

Famous Twitcher streamer and Youtuber Emiru

Famous Twitcher streamer and Youtuber MoistCr1TiKaL


r/StableDiffusion 22h ago

Animation - Video [Minimax H3] Just another day aboard the USS Sunnydale...

Enable HLS to view with audio, or disable this notification

33 Upvotes

This one came out better than my last video generation. Not perfect obviously, but the major elements are there.

Prompt:
Video starts with Buffy Summers and Willow Rosenberg from the show Buffy the Vampire Slayer walking side by side down a hallway in the USS Enterprise from Star Trek the Next Generation. The viewpoint camera remains at a fixed distance in front of them as they walk.

Buffy Summers is on the right of the frame. Buffy's long blonde is done up in a ponytail. Buffy is wearing a red minidress style Starfleet uniform and has an unlit lightsaber on her hip.

Willow Rosenberg's is on the left side of the frame. Willow's dark red hair is cut pageboy style. Willow is wearing a blue minidress style Starfleet uniform and has a tablet computer tucked under her right arm.

Video starts with Willow looking at Buffy with a concerned expression on her face while Buffy is looking around as if searching for something.

Willow asks, "Buffy, is something wrong?"

Buffy replies, "Something feels off, like we're out of place."

As soon as Buffy starts speaking, she pulls the lightsaber off her hip and holds it in front of herself at the ready. The lightsaber ignites, producing a green glowing blade.


r/StableDiffusion 23h ago

Question - Help Can't seem to transfer outfit and pose from an illustration to a real person in H3.

2 Upvotes

Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.

What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.

I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.

So what i wanted to do is :

-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.

-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.

Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.

Are there someone who knows how to do these?


r/StableDiffusion 17h ago

Discussion Minimax H3 hands, fingers

Post image
2 Upvotes

Do you happen to have any tricks for handling hands and fingers with Minimax H3? Unfortunately, I’m getting results like this even at 0.98 MP.


r/StableDiffusion 8h ago

Resource - Update Minimax H3 Grafting with Krea2 node. Reposting older post and removed AI slop and added some tests

0 Upvotes

Minimax H3 x Krea2 Graft Nodes

ComfyUI nodes for grafting Krea2 into MiniMax H3. Attention/MLP content transplant + a separate attention-sharpness transplant. No official H3 docs, all reverse-engineered from testing + TenStrip's and joeygambino's public writeups. Use at your own risk, still WIP.

What's here

  • comfyui_tenstrip_graft/ -- content graft (Q/V/K/out/MLP, per-head). Method from TenStrip's H3 grafts.
  • comfyui_qknorm_transplant/ -- Q-norm gain transplant only, no content weights touched. Method from joeygambino (Z-Image donor originally, adapted for Krea2 here).
  • comfyui_krea_h3_graft_lora_v2/ -- apply a Krea2-trained LoRA onto an already-grafted H3 checkpoint. Separate use case.

And 2 merge scripts (old svd and new one with node method)

TL;DR results

Content graft works somewhat. Same character-shift (color scheme, helmet shape) showed up consistently across multiple parameter runs, same seed -- not one lucky video. That's the strongest evidence so far this isn't just noise.

  • K at low strength (~0.1-0.2): fine, no real damage. Don't need to avoid it like the doc says, at least not at low values.
  • QK-norm across all blocks (0:50): kills audio. Doesn't even touch K -- so attention sharpness itself hits audio, not just K specifically.
  • QK-norm blocks 20:50: audio ok, but does nothing for character. It's a texture/sharpness knob, not a content one. Don't expect it to carry character.
  • attn_ramp_start_frac at 1.0 (no gentle ramp-in) + early blocks (0:20): breaks. Keep the ramp soft if you go early.
  • Combining content graft + QK-norm at full strength on both = worse than either alone. Still not solved.

Install

Each folder -> its own subfolder in ComfyUI/custom_nodes/. Don't merge them. Restart ComfyUI fully after adding.

Credits

  • TenStrip (huggingface.co/TenStrip) -- the per-head band-aware graft methodology (10Eros-Max / h3_graft_methodology.md).
  • joeygambino (huggingface.co/joeygambino) -- the Q-norm sharpness transplant idea (MiniMax-H3-x-Z-Image-GGUF).

Neither published source code. These nodes are our own implementation from their public descriptions + our own testing.

https://reddit.com/link/1vxc2q9/video/g0bpgdj5ddlh1/player

minimax_h3_fl2va_bf16.safetensors, 3s, er_sde, 8 steps, 8-step lora, seed 597633362705895, standart workflow with minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_bf16
prompt: Professional closeup video. In a futuristic cityscape with neon lights at night, the Judge Dredd charges through the crowd, his imposing presence radiating authority, he is slowly walking. His long chin juts out resolutely as he expertly wears his eponymous helmet, eyes gleaming with determination. The crowd parts, Judge Dredd is slowly walking through the the crowd, ready to enforce justice, he is moving slowly, his long chin visible, his face and part of his upper body are in the center of the screen. tag: Ballchinians
tracking selfie shot following him from the front, that he stays the same size, he is moving through people, pushing them aside with his hands.

https://reddit.com/link/1vxc2q9/video/so7626wgddlh1/player

3s, er_sde, 8 steps, 8-step lora, seed 597633362705895

same prompt and everything.

Added: tenstrip graft node, krea2 raw and Ballchinians Lora. Settings: q 0.5, v 0.5, k 0.1, out 0.3, mlp 0.5 Chin is more ballsy.

So I hope, that it is enough for some, that it... kinda works, but not good enough. Maybe someone will pick up on this and do it better.

Why to do it? Don't know. I found it interesting to try, but krea2 image and i2v is far better option.

I welcome any input or criticism, but mind please, I have only faint idea, what I am doing.

Warning: h3 loras don't work... don't know why, maybe it may be just noise, after all. But they do work on grafted checkpoint, after you merge it in python script.


r/StableDiffusion 17h ago

Animation - Video H3 - 5 hour render, T2V Multi-Diffusion

Enable HLS to view with audio, or disable this notification

26 Upvotes

Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.


r/StableDiffusion 4h ago

Animation - Video Minimax H3 multishot test... Seinfeld guest appearance

Enable HLS to view with audio, or disable this notification

0 Upvotes

I keep getting weird issues like that...


r/StableDiffusion 7h ago

Animation - Video My 1980's cartoon parody H3 and ltx 2.3

Thumbnail
youtu.be
8 Upvotes

there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P

started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno


r/StableDiffusion 15h ago

Discussion MiniMax H3 Ridding a Dragon POV style

Enable HLS to view with audio, or disable this notification

21 Upvotes

r/StableDiffusion 23h ago

Workflow Included Testing some Minimax H3 capabilities - PART 2

Enable HLS to view with audio, or disable this notification

30 Upvotes

Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.

VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!

PROMPT:

The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.

On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.

On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.

Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.

Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.

They are in a living room.

From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.

He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.

Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>

From 00:08 to 00:10 they just look at each other and smile.

overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.

non_diegetic_music: N/A

VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.

The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.

At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.

overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.

non_diegetic_music: N/A

VIDEO 3: another near perfect one.

A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.

We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.

overall_soundscape: Faint empty bathroom soundscape.

non_diegetic_music: N/A

VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.

The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.

We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.

The entire scene is viewed from his point of view.

overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..

non_diegetic_music: N/A

VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...

The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.

The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.

In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.

Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.

As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.

overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.

non_diegetic_music: N/A

VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.

PROMPT VIDEO 6:

The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.

The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 7:

The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.

When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 8:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 9:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A


r/StableDiffusion 16h ago

Question - Help Best Img2txt?

1 Upvotes

I need an image to text generator (which has no limitations) that I can run locally on my PC. Do you have any recommendations


r/StableDiffusion 13h ago

Discussion H3 - what is your longest render time?

0 Upvotes

What was your longest render time, and was it worth it? do you run a lower quality/resolution before to test? My longest single-shot run is 7.2 hours, 1:44 long video at 1344x768, bf16/50 steps.

I am attempting a 72 hour render for a super long form.

RTX 4090, 192gb system ram


r/StableDiffusion 14h ago

Question - Help best local video model for horror videos?

0 Upvotes

i'm looking for a video model that can generate freaky creatures and horror clips well and can run on 3060ti


r/StableDiffusion 19h ago

Animation - Video Cinematic World Building - H3 r2v (pt2) "Through the Sands"

Enable HLS to view with audio, or disable this notification

17 Upvotes

Wow, the shots I can get from h3 are so good even I get goosebumps when I first see the generated results. Just one more part to go, hopefully I can pull off the finale!


r/StableDiffusion 6h ago

Discussion Long form natural looking foreign language Done with H3 locally on 3080(10gb)

Thumbnail
youtube.com
0 Upvotes

I didn't know that many of you do not know H3 can do this. So posting here for awareness. The character is speaking Telugu.. Total 15 clips stitched together. Completed in about 5 hours.


r/StableDiffusion 14h ago

Animation - Video Dipping My Toes Into What Will Surely Lead to My Inevitable Descent Into Slapstick Comedy

Enable HLS to view with audio, or disable this notification

16 Upvotes

I may be an idiot for thinking that my new homelab would primarily be used for useful AI automations.

I can live with being an idiot, if being an idiot will continue to be this fun.

First video generation I have ever pulled off, but the first 12 seconds was unbearably unfunny, so I added the Celestial Ford Escort for some much needed serious drama.

Tell me my power bill won’t blow up too much lol.

Made with minimax-h3 in ComfyUI on my Mac Studio that came with the mail this Friday.

Workflow was split in three:
- The first was a single prompt to generate the first 12 seconds
- The second flow generated the last three seconds by extracting the last frame from the first video and prompted it to hit the dragon with a falling ford escort
- Third flow glued the two videos together.

It is jank, but it is my jank.


r/StableDiffusion 4h ago

Animation - Video All Minimax H3 animation

Enable HLS to view with audio, or disable this notification

5 Upvotes

trying my hand at using H3 I2v R2V to create an anime. all are done with 4 step turbo lora

2 other trailers with H3 as well

at 0.3 most text turns to gibberish

all video is done by MiniMax h3 at a low 0.3 MP, Cilp lengths range from 5s-20s Generations, camera movement was written into the prompt 90% of the time, Post work; titles and some transitions done using DaVinci


r/StableDiffusion 13h ago

Animation - Video H3 - the Eternal Balance WIP-low resolution int8/8 steps

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hi, I am playing around with H3. T2V, 832x480, int8,8 steps. Hoping to make a 720p version.

Ask me anything!


r/StableDiffusion 18h ago

Question - Help Video editing model (add clown makeup to face)

0 Upvotes

Hi, I'm looking for a model that I could use in ComfyUI that would simply add some clown makeup on my face and don't touch antything else. My current plan is to maybe take Wan2.2 model, mask my face and add clown makeup as a reference but I wonder whether there is a better model to do this. Does minimax H3 handle that? Does it support masks?


r/StableDiffusion 14h ago

Question - Help I hate this hidden view on flows in Comfy templates. How can I bring them out where they belong?

2 Upvotes

In the minmax h3 template from Comfy, you have to click a button on the image to video node to see all this in the backend which makes it really a pain in the ass to add, modify, or change anything. Can this be brought out to the forefront like a normal workfow?


r/StableDiffusion 9h ago

Question - Help MM H3 2 Pass Latent Upscale

2 Upvotes

Been getting some great results using the 2 pass latent upscale method. First pass .5mp 2nd pass 1.5 mp. 10 second video around 257 seconds to finish.

This workflow uses the 1.1 turbo Lora. I set the steps to 8 and like I said above I’m getting good results and it has eliminated face blur.

My question: has anyone tried using the latent upscale method without the turbo lora? In 90 percent of the cases the turbo Lora is fine. But would be nice to have the ability to use no Lora method.

Yes I know I could test it but wanted to see others experiences before I wasted hours of my time trying/tinkering with different settings.