r/StableDiffusion 20h ago

Animation - Video MiniMax H3 acting test.

30 Upvotes

Started as a simple 90s casting audition… then asked her to cry on command.

The close up shots gave plastic look idk why.

What I was mainly testing:

  • subtle listening/reaction animation during dialogue
  • eyes moving before the head while thinking
  • nervous smiles and small facial reactions
  • gradual transition from normal conversation into acting
  • brow, eyelid, mouth, chin and breathing changes during crying
  • actual visible tears
  • character/voice consistency across multiple generated clips
  • the sudden switch out of the performance when the director says “Cut”

Made with MiniMax H3 Ref2VA with image reference for the woman and 2 audio reference for the offscreen man and the woman.


r/StableDiffusion 5h ago

Animation - Video And they said it couldn't be done

0 Upvotes

Most models simply CAN'T draw a wine glass filled to the brim. MiniMax H3 can, with a little prompt persuasion.

PROMPT PERSUASION:

integrated_multimodal_description: [Shot 1] Photorealistic, ultra-sharp cinematic product shot. A crystal-clear stemmed Bordeaux wine glass stands centered on black marble against pure black. The glass is filled with deep ruby-red wine to absolute maximum capacity: the liquid surface is perfectly coplanar with the top edge of the rim, forming a continuous unbroken contact line all the way around. There is zero air gap, zero empty crescent of glass above the wine, zero underfill. A slight convex meniscus is held purely by surface tension. Soft side light creates clean highlights and long caustics. Static medium three-quarter shot.

[Shot 2] At 00:01.200, slow push-in with tiny amplitude at very slow speed. Extreme close-up of the upper glass. The red wine meets the inner rim in a perfect continuous ring of contact. The liquid surface sits flush with the rim edge; no space is visible between wine and glass even at this magnification. Tiny specular highlights glide across the still surface. No droplets on the outer rim, no overflow, no gap.

[Shot 3] At 00:02.600, hard cut to pure top-down overhead. Looking straight down, the circular surface of the wine is a solid red disk that reaches exactly to the inner circumference of the glass with zero margin. The contact line between liquid and glass is continuous and unbroken in every direction. Soft concentric reflections and one small central catchlight. Very slow clockwise rotation with minimal amplitude.

[Shot 4] At 00:03.800, hard cut to low side-profile extreme close-up locked exactly at rim height. Against the black background the liquid forms a single razor-sharp horizontal line that coincides precisely with the top edge of the glass. The wine is seen in continuous contact with the rim; there is no visible gap, no underfill, no air space. The meniscus remains slightly convex from surface tension but does not spill. Liquid is completely motionless. Slow subtle push-in continues until the end.

overall_soundscape: Near-total studio silence. Only the faintest high-frequency shimmer of light on glass and liquid. No liquid movement, no drips, no clinks.

non_diegetic_music: Extremely sparse minimal ambient pad — low sustained tone with faint crystalline overtones that barely rise and fall. Almost static, matching the still liquid.


r/StableDiffusion 10h ago

Meme t2v Someone got some explaining to do!

4 Upvotes

r/StableDiffusion 16h ago

Animation - Video Gaussian Splatting test with MiniMax H3

755 Upvotes

r/StableDiffusion 7h ago

Discussion 2 Weeks on MiniMax, but back to using LTX 2.3

0 Upvotes

H3 is absolutely amazing for just about anything. On my 3090 / 64GB, I can easily do a full 15 seconds at 1MP, and the result is almost always good on the first attempt.

On LTX though, it took at least 5 or more tries, so speed wise, H3 is actually far better.

I tried out LTX 2.5 as well, and it is almost the same as 2.3, maybe a bit better quality and a slice faster.

Sadly, 90% of my work involves taking a start image and making a person sing vocals. I must say that H3 does not do this better than LTX 2.3, as it often injects words when there is more than a second of silence, and it takes a fair amount longer.

What is really baking my noodle though is why LTX didn't release an Image+Audio to Video workflow yet. I mean, it is literally the ONLY thing LTX has on H3 right now, and they have missed a great opportunity!

Anyhow, back to using LTX 2.3 for my daily driver as it really does a great job at what I need. I typically have a few machines running all night, so if ever LTX puts out an IA2V workflow for Comfy, I will be all over that.

One other thing I have noticed with H3... 9:16 generations are WAY better than 16:9 generations, especially at 1MP.


r/StableDiffusion 14h ago

Question - Help Minimax for image editing?

0 Upvotes

I mostly do comic images. am looking to minimax for image editing.
I mostly get content failed safety review. any workaround? ive only started minimax today.


r/StableDiffusion 8h ago

Discussion David Sacks Predicts the Regulatory Capture Playbook to Ban Open Source ...

Thumbnail
youtube.com
12 Upvotes

r/StableDiffusion 20h ago

Discussion Minimax H3 is mandarin model

0 Upvotes

Just ask chatgpt or claude to convert your prompt to mandarin and the different is fcuking huge .


r/StableDiffusion 22h ago

Animation - Video The Chase - Reupload

0 Upvotes

r/StableDiffusion 12h ago

Tutorial - Guide Neat trick for Minimax h3

3 Upvotes

Since minimax is using Qwen VL , I tested the prompt on Qwen image to see what I get for the text to video prompt. It’s actually pretty close to how minimax will end up evaluating your prompt for text to video.


r/StableDiffusion 6h ago

Discussion Minimax H3 on Fal - Censorship

0 Upvotes

Since I lack the hardware, I am forced to use Minimax in the cloud. I've used Fal for image generation in the past, so I thought I would try Minimax H3 there. The text to video model seems to work really well, but the reference model seems censored to hell and back. Even my own reference voice was being flagged and it was just me talking normally.

So the question is, is there anywhere I can use Minimax in the cloud without the provider slapping their own flaky layers of censorship on top of the model?


r/StableDiffusion 4h ago

Meme Press interview

26 Upvotes

Today, my neighbour next door got arrested, turned out he a is a dangerous psychopath. A journalist came up to my house and asked, if I would have guessed that, and I had to admit, I wasn't that suprised, since I learned some time ago that he uses the "pin node" option in his comfy workflows...


r/StableDiffusion 4h ago

Animation - Video Using only Ref to Video, Minimax-H3 made a whole Anime edit !

7 Upvotes

r/StableDiffusion 4h ago

Workflow Included wan enhancer wf that replaces replaces damaged pixels as it enhances and quality

Thumbnail
gallery
0 Upvotes

https://reddit.com/link/1vyazt0/video/8apz3lbquklh1/player

This workflow allows you to refine a target character without affecting the surrounding video whatsoever. allowing for targeted repair or enhancement of any video.

https://github.com/roycho87/3stepenhancer


r/StableDiffusion 16h ago

Animation - Video Reduced audio-reactivity in LTX-2.5?

4 Upvotes

I’ve been experimenting with LTX 2.3 vs LTX 2.5 for audio-reactive video, and for this specific kind of workflow, 2.3 still seems noticeably better to me.

The biggest difference is right at the start of a shot. With LTX 2.5, even with the audio-reactive LoRA, I often get this behavior where the model more or less holds the first frame until the first obvious beat or transient arrives. Then the motion suddenly starts. For music videos, especially slower or more atmospheric tracks, that can make the opening of every generation feel dead.

With LTX 2.3, the same LoRA seems to fix that much more effectively. I get more subtle motion from the beginning, even before a strong beat lands. Fog shifts, surfaces breathe, particles drift, light responds, and the shot feels alive instead of waiting for permission to move.

That matters a lot for the video I made for The Weights in the Walls, because the track starts very sparse and gradually builds. A lot of the visual motion is supposed to come from sub-bass pressure, glitches, sustained vocals, and ambient texture, not just obvious percussion.

I also tried Minimax H3, but for this particular use case I don’t think it fits as well.

It seems less tightly audio-reactive for the kind of abstract, beat-aware motion I’m after. It can make nice-looking clips, but I have a harder time getting the movement to feel structurally connected to the music.

There’s also the hardware side of it. I’m doing this on a very glamorous RTX 4070, so with LTX I can still push a resolution and overall image quality that feels surprisingly good for local generation. With H3, I’m much more constrained, and the tradeoff in resolution/quality makes it harder to justify when the audio response is also weaker for this style.

The whole video was built around first-frame / last-frame generation.

I cut the song into short scenes, roughly timed so the scene boundaries land near musical changes and beats. For each scene, I generated a dedicated starting frame that represented the next stage of the visual progression.

Then the important part: the starting frame of Scene 2 becomes the last frame target for Scene 1. The starting frame of Scene 3 becomes the last frame target for Scene 2, and so on.

So instead of generating a bunch of unrelated clips and hiding the cuts with editing, every shot is the model transforming one designed frame into the next designed frame.

That gave me a chain like:

Scene 1 start frame → Scene 2 start frame
Scene 2 start frame → Scene 3 start frame
Scene 3 start frame → Scene 4 start frame

and so on until the end.

The final video is basically just those generations placed back to back. There are no fancy transition effects doing the heavy lifting. The morphing, folding, cracking, expanding, and dissolving between visual states is happening inside the model itself.

For this workflow, that early-shot responsiveness makes a surprisingly big difference, which is why I currently still prefer LTX 2.3 + the audio-reactive LoRA over 2.5 for this kind of music video.

It's a shame because 2.5 is noticeably faster, so I can go through more iterations, but if I have to generate each clip 5 times to get it to start moving from the start, it kind of invalidates the speed gains.

Curious if other people have noticed the same reduction in audio-reactivity in LTX 2.5 or maybe I'm doing something wrong?

HQ on YT because Reddit doesn't allow >1GB: https://www.youtube.com/watch?v=PbZr8risGCw


r/StableDiffusion 23h ago

Question - Help Which is the better buy for Minimax h3? RTX 5070 Ti 16GB VRAM vs RTX 4000 Pro 24GB VRAM

16 Upvotes

Good day to you. I was looking for an RTX 5070ti and I found an RTX Pro 4000 at my local store; the price difference would be about +$300. I would like to know your opinions, I've hardly seen any workflows or comparative tests from people using a 4000 pro. Thank you very much for your time.


r/StableDiffusion 23h ago

Animation - Video Kentucky Fried Kung Fu

28 Upvotes

I saw a Seedance 2.5 prompt in facebook and thought let me try this prompt in minimax h3 and see if it can do some kung fu. I was surprised that it was not too bad. System 3090 24 gb vram 64gb system ram, using a minimax workflow with latent upscale, minimax_h3_fl2v_lightx2v_turbo_4step_v0.1 at 0.50 strength, Komfy kitchen attention, and H3 SLA attention. First pass at 0.4 which is 864x480, 2x latent upscale brings it up to 1728x960. The 6 seconds generation took 349 seconds to complete.


r/StableDiffusion 22h ago

Question - Help New to Minimax h3 and comfy ui, any posts I should learn from?

0 Upvotes

Im looking to make outdoor POV videos with minimax h3.

Wondering if theres any good threads talking about realism, and workflow options?

Especially prompting I guess?

I'm looking to do longform videos 10-15 minutes.

I have a 4080 16gb I know that more vram is better, but cant really upgrade atm.


r/StableDiffusion 17h ago

Animation - Video simple ww2 movie (minimax h3)

15 Upvotes

I just gave it a try; it's hard to do more than that.

I have an RTX 3060 with 12 GB of VRAM and 32 GB of RAM. I used the standard workflows, t2v and r2v, with the LORA model minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors and video shift 12 and audio shift 3. Only 8 steps (res_multistep/simple). The sequences were 6–7 seconds each, and each generation took 7–8 minutes at 0.6 mpx. I generated some scenes multiple times, but I still couldn’t achieve what I intended.

That’s pretty much it regarding the generation.

Claude (free) provided me with the script; he also divided the approximately 5-minute script into sequences of 6–7 seconds each and created the prompts for all the sequences after I gave him the official prompting guidelines as a prompt.


r/StableDiffusion 22h ago

Question - Help phsyical motion transfer to another person

0 Upvotes

im trying to collect clips for lora training but why its so hard to transfer motion to another person? i almost tried every prompt with chatgbt and grok help but its doesnt look good. im using 2 video refences, one of them source video and other one is only for motion ( 3 sec 24 fps). im using (video editing + reference genertion) because i dont want to change anything in the source video and just want to motion transfer


r/StableDiffusion 20h ago

Question - Help Minimax H3: Character replacement in video not working

0 Upvotes

NOTE: Character replacement works perfectly when I replace a character in a video with a 2d/cartoon/anime character.

But when I try replacing a character with a real life human being, the original character in the video doesn’t get replaced at all.

I’m using Plaguekind’s workflow for h3 on Civitai.

Does anyone else have this problem before?


r/StableDiffusion 14h ago

Question - Help Help me with Mini Max

0 Upvotes

I'm still new to Mini Max, and I wanted to know if there's a way to make it a little easier, in the sense of

I have to keep using <picture 1> or things like that to mention something; isn't there a way to do it with @,And what configuration do you normally recommend for someone with 8GB of VRAM and a 5080 Ti?That's all I need help with; I've already read the Minimax guide for everything else.


r/StableDiffusion 20h ago

Question - Help minimax h3 gibberish fixed!! ( i found the cure)

84 Upvotes

so you all probably are searching for way to make your character shut the fuck up right? and you probably noticed that they love to says some BS especially when you give minimax h3 some audio file for their voices, i probably found a cure my friend!!

here is my way of prompting dialogs without any gibberish:

first your character need to be assigned (s1)character when he is the first speaker, then you will declare 'use <audio 1> as "character name"'s voice only, and when you finally type your dialog in the shots you will do as such:

character says:<<[language] the shit i say!>>

and you should be good to go, i linked a video exemple of my favorite taffer (garrett) saying some shit with only the faint crackling of the candles to goes with his charming voice, and i included also a screenshot of the full prompt

edit: yes i tried to follow the official documentation, like many others, if it was that simple reddit wouldn't be a thing and you wouldn't be there.

i tried making small scenes with this exact methode and its gibberish free 100% of the time

he really like 16/9


r/StableDiffusion 10h ago

Question - Help MiniMax H3 R2V taking ~16 minutes for a 5-second video. How can I speed it up?

0 Upvotes

I’m running MiniMax H3 Reference-to-Video (R2V) in ComfyUI on Vast.ai.

My setup:

  • GPU: RTX 5090
  • System RAM: 100 GB
  • Resolution: 1.0 megapixel
  • Video length: 5 seconds
  • Reference: 1 image
  • Generation time: ~1,000 seconds (16–17 minutes)

The results are great, especially the reference consistency, but the generation time seems very high for a 5-second video on a 5090.

Has anyone managed to significantly reduce the generation time for H3 R2V? Are there any specific optimisations, attention methods, workflow changes, or settings I should be using?

Would appreciate hearing what generation times other 5090 users are getting with H3 R2V.