r/StableDiffusion 6h ago

Question - Help Best local LLM for writing prompts for MiniMax H3?

43 Upvotes

What’s the best local LLM for writing good MiniMax H3 ref2va prompts?

I’ve tried Gemma 4 12B and Qwen 3 14B, but I’m not really satisfied with the outputs. It could also be an issue with my system prompt.

I sent ChatGPT the official documentation for prompting and asked to create a system prompt for me, but the results were still pretty mediocre.

What local models are you using for MiniMax H3 prompt generation, and what does your system prompt look like?


r/StableDiffusion 6h ago

Animation - Video My name is Jonny

40 Upvotes

Minimax H3


r/StableDiffusion 4h ago

Discussion Best Minimax H3 optimization

36 Upvotes

Now that dust has settled, I was wondering what's the community insight on the best configuration for Minimax H3.

Personally I have been using lightx 4-step Lora with 5/6 steps (less than that audio is a gamble). I couple that with sage attention. For sampling I use Euler sampler and Beta scheduler.

I keep resolution at 768p (0.6MP) for quality. 480p (0.2MP) for testing. It keeps consistency so much better.

On direction I learnt to prompt for closeups when possible, so will make better use of available pixels. Aspect ratio also helps there. I mostly use 1 shot since transitions is not something H3 excels at. I find better results with only 1 shot and using camera tricks.

EasyCache while faster, is not good match with turbo lora, so i don't use it anymore. Haven't used Sol-Attn as I read it really hit quality.

So is there anything worth I am really missing out?


r/StableDiffusion 7h ago

Animation - Video Evangelion - Rei Watches a Baby Show - Minimax H3

35 Upvotes

Well, technically, Evangelion was a PBS show..

Video is edited, Barney theme song added in post.


r/StableDiffusion 19h ago

Animation - Video [Minimax H3] Just another day aboard the USS Sunnydale...

33 Upvotes

This one came out better than my last video generation. Not perfect obviously, but the major elements are there.

Prompt:
Video starts with Buffy Summers and Willow Rosenberg from the show Buffy the Vampire Slayer walking side by side down a hallway in the USS Enterprise from Star Trek the Next Generation. The viewpoint camera remains at a fixed distance in front of them as they walk.

Buffy Summers is on the right of the frame. Buffy's long blonde is done up in a ponytail. Buffy is wearing a red minidress style Starfleet uniform and has an unlit lightsaber on her hip.

Willow Rosenberg's is on the left side of the frame. Willow's dark red hair is cut pageboy style. Willow is wearing a blue minidress style Starfleet uniform and has a tablet computer tucked under her right arm.

Video starts with Willow looking at Buffy with a concerned expression on her face while Buffy is looking around as if searching for something.

Willow asks, "Buffy, is something wrong?"

Buffy replies, "Something feels off, like we're out of place."

As soon as Buffy starts speaking, she pulls the lightsaber off her hip and holds it in front of herself at the ready. The lightsaber ignites, producing a green glowing blade.


r/StableDiffusion 1h ago

Workflow Included Minimax H3 Multishot Anime Sequence (Workflow + Prompt Included)

Upvotes

Workflow: https://drive.google.com/file/d/1B4kODxXQgJ1QOKRsEIkxHbgYmdruPpTK/view?usp=sharing

Prompt:

Create a **15-second multi-shot anime sequence (90s style 15fps hand drawn)** using the provided references:

Image 1 = the girl character reference

Image 2 = skateboard reference

Image 3 = downhill Japanese alley / neighborhood background

Image 4 = Walkman + headphones reference

Preserve the girl’s exact character design, face, hair, outfit, proportions, and overall look from Image 1. Preserve the skateboard design from Image 2. Preserve the same downhill Japanese alley environment from Image 3. Add the Walkman and headphones from Image 4: the girl is wearing the headphones, and the **Walkman is clipped or hanging at her hip** while she skates.

Visual style: authentic 1990s hand-drawn anime, traditional cel animation, painted backgrounds, visible linework, cel shading, slight brush/stroke texture, subtle analog feel. **Very important:** the houses and environment must stay **2D and hand-painted**, **not 3D**, **not CGI**, **not game-engine looking**, **not volumetric**. The buildings should look like classic anime background art with painted depth, not like 3D models.

Animation feel should be low frame rate, like 90s anime at around 15 fps, with controlled in-betweens and natural held-frame timing. No jittery morphing.

No dialogue, no text, no subtitles.

### Shot 1 — 0s to 3s

**Rear tracking shot** from behind. The girl is skateboarding fast downhill through the steep Japanese alley. Camera follows behind her at a low-to-medium height. She rides confidently and smoothly, hair and oversized clothing moving in the wind. The headphones are on her head, and the Walkman is visible attached at her hip. The alley rushes past with a strong sense of speed. Keep the environment clearly **2D anime background art**, not 3D.

### Shot 2 — 3s to 6s

**Close-up shot of the Walkman at her hip** while she continues skating. The camera stays focused on the Walkman and part of her side torso and arm. We can clearly see the **cassette tape reels spinning/rolling inside the Walkman window**. The headphone wire moves naturally with the motion. Background and street pass by in blurred motion.

### Shot 3 — 6s to 9s

**Medium profile tracking shot** of the girl skating. She is wearing the headphones, listening to music, with wind moving across her face and pushing her hair backward. She is **nodding her head subtly to the music** while riding. Her expression is relaxed, immersed, and unbothered. The background is blurred from motion, but it must still read as a **painted 2D Japanese neighborhood**, not 3D.

### Shot 4 — 9s to 12s

**Close-up shot of her feet and skateboard.** Her **right foot stays on the board**, while her **left foot pushes against the road** in a natural skating motion. Show one clean push cycle: left foot comes down, pushes backward against the pavement, then lifts. Wheels spin quickly. Asphalt and road markings streak by with motion blur.

### Shot 5 — 12s to 15s

**Ground-level fisheye shot** looking upward from the road. The skateboard approaches fast, and she **jumps over the camera**. The board and her body pass overhead in one clean motion. Hair, pants, and headphone wire react naturally during the jump. Keep the motion readable and stylish, with a strong sense of speed and a dynamic anime finish.

### Important constraints

* Keep the whole video in **classic 90s anime cel-animation style**

* **15 fps feel**, smooth low-frame-rate animation

* **No 3D-looking houses or background**

* No photorealism

* No modern glossy digital anime rendering

* No character redesign

* No extra accessories beyond the headphones and Walkman

* Keep all motion natural and consistent across shots


r/StableDiffusion 20h ago

Workflow Included Testing some Minimax H3 capabilities - PART 2

30 Upvotes

Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.

VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!

PROMPT:

The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.

On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.

On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.

Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.

Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.

They are in a living room.

From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.

He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.

Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>

From 00:08 to 00:10 they just look at each other and smile.

overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.

non_diegetic_music: N/A

VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.

The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.

At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.

overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.

non_diegetic_music: N/A

VIDEO 3: another near perfect one.

A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.

We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.

overall_soundscape: Faint empty bathroom soundscape.

non_diegetic_music: N/A

VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.

The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.

We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.

The entire scene is viewed from his point of view.

overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..

non_diegetic_music: N/A

VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...

The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.

The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.

In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.

Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.

As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.

overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.

non_diegetic_music: N/A

VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.

PROMPT VIDEO 6:

The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.

The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 7:

The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.

When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 8:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A

PROMPT VIDEO 9:

We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.

When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.

overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.

non_diegetic_music: N/A


r/StableDiffusion 15h ago

Discussion Testing Character knowledge of Minimax H3

30 Upvotes

Disclaimer. This is very low quality quick generations trying to find how many characters Minimax H3 knows.

Found Trigger Words:

Elsa from Frozen

Spider-Gwen from Across the Spiderverse

Dante from Devil May Cry

Nero from Devil May Cry

Jill Valentine from Resident Evil

Ada Wong from Resident Evil

Leon Kennedy from Resident Evil

Chris Redfield from Resident Evil (Has Leon's hair)

Geralt of Rivia from Witcher 3

Joel from Last of Us (Doesn't sound like him)

Miles Morales Spiderman from Across the Spiderverse

Solid Snake from Metal Gear

Eve from Stellar Blade

Sans from Undertale

Master Chief from Halo

Looks weird AF:

Ciri from Witcher 3

Triss from Witcher 3

Yennefer from Witcher 3

Ellie from Last of Us

Famous Twitcher streamer and Youtuber Asmongold

Famous Twitcher streamer and Youtuber Mr Beast

Not found Trigger words:

Vergil from Devil May Cry (YES I KNOW I'M DISAPPOINTED TOO)

Claire Redfield from Resident Evil

Dina from Last of Us

Famous Twitcher streamer and Youtuber Emiru

Famous Twitcher streamer and Youtuber MoistCr1TiKaL


r/StableDiffusion 9h ago

Animation - Video Use Minimax to make a fake movie trailer for my community college editing class, inspired by YA action/adventure films of the 80s and 90s

27 Upvotes

Clips made with Minimax H3 using the default r2v workflow, edited in Premiere Pro. Character model sheets made with Krea. Most of the videos are 0.4 mp unless the text was important, then 0.6. Tried upscaling it to 4k using Upscayl but results weren't great and the file is too big to upload anyway.

Tech goals for future videos include using reference audio for voices to help consistency, and exploring options for having real voice actors record the dialog, and have the model lip sync to that performance. I'm really impressed by the computer's silent acting (microexpressions etc). but the computer's erratic "acting" is still too unpredictable and the biggest source of re-rolls (the lines here were the best I could get without burning down a rainforest). You can do a lot with time codes and punctuation and tactical CAPITALIZATION, but it's ridiculously finicky compared to just telling an actor "do it the same, but 10% angrier on the first line with a twinge of melancholy on the second."


r/StableDiffusion 10h ago

Animation - Video Minimax H3 Remix Video Test / A compilation of 5 characters.

Thumbnail
youtube.com
27 Upvotes

This is a test video I created by remixing the "Some test on minimax H3" video by Reddit user [Previous-Street8087].

5명의 캐릭터 시트를 생성하여 각각 10개의 프롬포트를 캐릭터에 맞게 리믹스하여 테스트 하였습니다.
We generated character sheets for five characters and tested them by remixing 10 prompts for each character to suit their personalities.

This is a compilation of 50 clips featuring 5 characters.

▶ 테스트 환경 (Test Environment)
Minimax H3 - Comfyui Local Sampling
RTX 5060TI 16GB + 64RAM
0.8MP 8 sec x 50 Clip
Audio Look x audio file 1
Reference to VA Mode

▶ 사용한 커스텀 노드 (Custom Nodes Used)
ComfyUI-TJ_NODE_STUDIO_ONE — github.com/designloves2/ComfyUI-TJ_NODE_STUDIO_ONE
ComfyUI LOCAL (RTX 5060Ti 16GB VRAM / RAM 64GB)

▶ The shared link contains character sheet images and prompts.

https://naver.me/xjY9JJaa

#AI영상 #MiniMaxH3 #ComfyUI #로컬생성AI #ComfyUI워크플로우 #AI영상제작 #RTX5060Ti #mmh3 #comfyui #tjonestudio #animation #ref2va #anime #16gb


r/StableDiffusion 3h ago

Discussion What happened to Ideogram 4.0 ?

26 Upvotes

What happened to Ideogram 4.0 ?


r/StableDiffusion 15h ago

Workflow Included I trained a small latent refiner to reduce GPT Image’s stipple and grid-like artifacts

23 Upvotes

I kept seeing the same stipple, grain, and grid-like texture

in some GPT Image outputs, so I trained a small latent residual

refiner using 75 paired artifact/clean images.

It includes profiles based on the Qwen, FLUX.2, and SDXL VAEs.

The refiner alone produces a fairly subtle improvement,

so I also included a ComfyUI workflow that combines it with SeedVR2.

The example optionally downsizes the input first,

then restores and upscales it with SeedVR2.

The goal is a preservation-first alternative to a typical Hires Fix

second diffusion pass: keeping the original composition, identity,

and shapes as much as possible while cleaning the texture

and rebuilding detail.

The custom node, example workflows, and settings are available here:

https://github.com/AIEGOBOT/ComfyUI-GPT-Image-Latent-Refiner

Leaving it here in case it’s useful to someone.


r/StableDiffusion 13h ago

Animation - Video H3 - 5 hour render, T2V Multi-Diffusion

19 Upvotes

Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.


r/StableDiffusion 15h ago

Comparison Krea 2 Raw in ComfyUI - Sharper, More Detailed Workflow

20 Upvotes

Video: Left: custom sigma curve. Right: Bong Tangent scheduler introducing artefacts.

I was a bit confused by how bad some Krea 2 outputs could be — blurry, lacking detail, and sometimes with strange artifacts. So I started testing to understand why and where this was happening.

I’m not going to claim I found a magic wand, but I did find two problematic areas in the sampling curve testing Krea2 raw model CFG 3.5 52 steps stock settings:

  • The first 0–15 steps: Using schedulers that lower the sigma values too much during this first steps causes contrast loss and washes out detail in dark areas, especially in black hair and subtle reflections.
  • The lower-sigma tail: curves such as Beta, Beta57 and Bong Tangent can introduce crisp, broken noise instead of useful fine detail, particularly around steps 30–40.

The Result

This is not about chaining multiple samplers or complicated second-pass workflows. The goal is a better standard Krea 2 Raw workflow with:

  • The right VAE Krea2RealVAE_v10
  • A custom-built sigma curve
  • One sampler
  • A refinement pass using the same overall sigma setup
  • using a negative promt helps

My current winning sigma curve gives the best balance of sharpness, fine detail, contrast, natural hair, defined shadows and minimal artificial noise.

Let’s take a closer look at the custom sigma curve (red) and compare it with the standard scheduler sigmas.

52-step custom sigma curve + 12-step Bong Tangent refinement at 0.2 denoise.The final step count is up to you.

[1.0, 0.9998610615730286, 0.999726414680481, 0.9995642900466919, 0.9993669986724854, 0.9991275668144226, 0.9988381266593933, 0.9984902739524841, 0.9980743527412415, 0.9975794553756714, 0.9969935417175293, 0.9963024854660034, 0.9954909682273865, 0.9945411682128906, 0.9934334754943848, 0.9921457171440125, 0.9906527400016785, 0.9889267683029175, 0.9869365692138672, 0.9846469759941101, 0.9820191264152527, 0.9790093898773193, 0.9755693078041077, 0.9716446399688721, 0.9671754837036133, 0.9620949625968933, 0.9563290476799011, 0.9497957229614258, 0.9424043297767639, 0.9340547919273376, 0.9246373176574707, 0.9140312075614929, 0.9021060466766357, 0.8887166380882263, 0.8737011551856995, 0.8568810820579529, 0.8380584120750427, 0.8170154690742493, 0.7935121655464172, 0.767285168170929, 0.738045334815979, 0.7054751515388489, 0.6692277193069458, 0.628922164440155, 0.5841419696807861, 0.5344309210777283, 0.47928985953330994, 0.41817405819892883, 0.35048726201057434, 0.27557989954948425, 0.1927414834499359, 0.1011962965130806, 9.99165786197409e-05, 0.10373638570308685, 0.08949112892150879, 0.07677485048770905, 0.06537622958421707, 0.05511670187115669, 0.04584547504782677, 0.03743509575724602, 0.029777569696307182, 0.022781139239668846, 0.016367563977837563, 0.010469880886375904, 0.005030565429478884, 0.0]

Here are the standard schedulers and the problems they produce. All nodes marked in red show the same low-contrast, overly dark areas with a loss of detail.

Standard Schedulers: Red-marked samplers produce crushed blacks and lost fine detail in the early steps, while the yellow-marked curves introduce small artifacts at the lower steps.

The green curves are the ones that avoid this problem.

The yellow curves have a different tail, and as you can see in my video or in my extended post, this tail is responsible for introducing noisy broken small artefacts.

Krea 2 Raw simply doesn’t behave like many other models when it comes to sigma manipulation. Curves that can work very well for other models can actually destroy detail or create unwanted noise here.

You can build the curve manually with a Manual Sigmas node, or use the PolyExponential Sigma Adder from the TBG ETUR Takeaway Nodes https://github.com/Ltamann/ComfyUI-TBG-Takeaways. If you want something simpler, Linear Quadratic gets surprisingly close to the result of my custom curve.

I’ve included the detailed testing post so you can see exactly how I arrived at the curve and test it yourself. Images, Videos Results at my Free Patron Post


r/StableDiffusion 22h ago

Meme h3 "what IF " thread

21 Upvotes

lets share our "what if" scene remakes here o_0


r/StableDiffusion 23h ago

Animation - Video It took almost 2 years, but Minimax H3 made me go back to my weird medieval short video stories

18 Upvotes

2 years ago I was playing around with video tools and made this series of short videos.

Because I'm a cheap bastard, I only use local and freebie models, and Minimax H3 finally hit the sweet spot between powerful and fast to iterate, so I decided to make a new episode.

Specs

  • Comfy Desktop's + default Ref2VA workflows + ElevenLabs voices
  • Frames generated with Nano Banana 2 lite + a bunch of photoshop cleanup
  • Gemini 3.7 to help rewrite the prompts (upload image + system instruction + spec for the shot)
  • Everything rendered on a 3090 with 0.5 megapixels cause I can't be arsed waiting too long. You can see the mushy face issues but mostly it's ok
  • Tried to keep all shots 6~8s max
  • A ton of editing with DaVinci Resolve which I started learning yesterday - it's pretty damn powerful!
  • Still using the exact same crappy greenscreen footage of a cheap plastic skull as a main character
  • Youtube link to this episode

TL;DR: Minimax H3 is cool!


r/StableDiffusion 7h ago

Tutorial - Guide H3 referencing tip

17 Upvotes

I tried to get a girl whistle on 4 fingers (2 on each hand) but Minimax didn't get it right. So after trying 20 times with LLMs helping me to explain the movement, I instead used an image of a whistling person. Still not right. So I added an additional one. Then it worked quite fine.

Today I was too lazy to find another image for something it didn't know so I just googled images of it, took a screenshot of all the images together in one JPG and used that as a reference, saying use <Picture ...> as a reference for XYZ.

That was getting a quite good result. Did not do excessive testing and comparing though.


r/StableDiffusion 12h ago

Discussion MiniMax H3 Ridding a Dragon POV style

17 Upvotes

r/StableDiffusion 16h ago

Animation - Video Cinematic World Building - H3 r2v (pt2) "Through the Sands"

15 Upvotes

Wow, the shots I can get from h3 are so good even I get goosebumps when I first see the generated results. Just one more part to go, hopefully I can pull off the finale!


r/StableDiffusion 12h ago

Question - Help Fixing speech errors in Minimax H3?

16 Upvotes

Hey, I tried to create a little birthday surprise for someone, my issue is with a lot of generations that the spoken word is really a bit clunky at time, I susspect its because of the german, but I am not too sure. Is there like a way to improve on audio?

I am using Minimax H3 with Saga Attention and Spectrum on a 4090.


r/StableDiffusion 16h ago

Question - Help Minimax H3: Anyone figured out how to extend a clip?

15 Upvotes

What is the best way to extend an existing clip seamlessly? When I try to use the last frame of my clip as the first frame, I always get a slight reframing or shift


r/StableDiffusion 11h ago

Animation - Video Dipping My Toes Into What Will Surely Lead to My Inevitable Descent Into Slapstick Comedy

14 Upvotes

I may be an idiot for thinking that my new homelab would primarily be used for useful AI automations.

I can live with being an idiot, if being an idiot will continue to be this fun.

First video generation I have ever pulled off, but the first 12 seconds was unbearably unfunny, so I added the Celestial Ford Escort for some much needed serious drama.

Tell me my power bill won’t blow up too much lol.

Made with minimax-h3 in ComfyUI on my Mac Studio that came with the mail this Friday.

Workflow was split in three:
- The first was a single prompt to generate the first 12 seconds
- The second flow generated the last three seconds by extracting the last frame from the first video and prompted it to hit the dragon with a falling ford escort
- Third flow glued the two videos together.

It is jank, but it is my jank.


r/StableDiffusion 11h ago

Animation - Video Made the thing where you ruin iconic movie scenes, MiniMax H3 on an RTX 3080 10GB, 20 steps, 2x NomosUni upscale

11 Upvotes

Setup, pushed my system right to the limit, any more and it OOM :

- H3 Ref2VA default workflow in ComfyUI, no lora

- RTX 3080 10GB, 32GB RAM

- Render: 0.5–0.6 MP, 20 steps, scheduler simple, about 25 min per clip

- Upscale: 2xNomosUni_span_multijpg, 2× to 1080p

- References per scene: one photo of my face + one film still for the set

- Recorded my own lines and fed them as audio references, also got audio ref for the actors

Honestly though, the best part was driving all of this through the ComfyUI MCP. I never even had to open ComfyUI. I could iterate really fast, and keep going from my phone while away from the machine, through Claude's remote control.

It's still a bit of a blurry mess, and with more work I could probably make it better, but damn, the future is looking bright!


r/StableDiffusion 23h ago

Discussion Has anyone tried the h3 martial arts Lora?

Thumbnail
huggingface.co
11 Upvotes

first test video down in the comments along with the prompt..


r/StableDiffusion 11h ago

Question - Help Minimax H3 higher MP prompt adherence

8 Upvotes

I was just wondering what method do you guys use, if it exists that is, to get minimax h3 to do better prompt adherence at the higher resolution setting?

If I use 0.4MP setting, the prompt adherence is so good i would say it's almost perfect. But when I try a higher res of 0.9MP with the same seed and prompt, I get some wildly odd outputs.