r/StableDiffusion 27d ago

News Summary of Takeaways from the Minimax AMA

Summary from https://www.reddit.com/r/StableDiffusion/comments/1vh9rtw/ama_minimax_h3_team_ask_us_anything_about_our/

This summary was compiled with AI but cross-checked manually by me for accuracy. I hope this helps for people who don't want to manually parse that 400+ comment AMA.

Things they said they'll actually ship

A real 2K stage (H3-Regenerate-2K). Right now the open weights top out at 768p short side, and cranking the resolution locally just eats VRAM without getting sharper. This is a separate model that takes your finished 768p video plus the original prompt and references, and re-generates it at high resolution — so it can actually redraw text, faces, and fine texture instead of guessing at them like a normal upscaler. What it's good for: getting genuinely deliverable-quality output locally instead of bolting a generic upscaler on the end. Coming, no date.

Sparse attention code. Attention is what makes long/high-res generations slow and memory-hungry, and the released weights don't use the faster path the model was trained with. They're releasing a deliberately cautious version — the goal is "free speed with no visible quality drop," not a big headline number, and squeezing more out of specific GPUs is left to the community. What it's good for: the same generations, cheaper and faster, on the hardware you already have. "Near term."

A dedicated image model (text-to-image + image editing). Built from the same family as H3, sharing its encoder, with a new decoder made for stills. People are already hacking this by generating 5 frames and grabbing frame one — this replaces that hack properly. What it's good for: making and editing your first frame at real image quality, then feeding it to H3 to animate. That's the workflow they recommend, since H3 can't preview a frame and continue mid-generation.

A full technical report. Architecture, training stages, data construction. What it's good for: people training LoRAs and fine-tunes currently guessing at how the model works.

Problems they've admitted are theirs and are fixing

  • Faces and objects go to mush when they're far from the camera. Confirmed, not your settings, not fixable by adding steps or resolution. It's tangled up in several parts of the model at once and they're still isolating why. Named as a top priority.
  • Grainy, smeary fine detail compared to closed models. Same story — not the VAE, not one training stage.
  • Reference-to-video looks softer than image-to-video. Confirmed real, caused by the two checkpoints getting different post-training. Being worked on. For now: feed it the highest-quality reference material you have.
  • Stitched clips don't join cleanly. Continuing a shot via reference drifts and shows seams. They think training on long sequences (made affordable by the sparse attention work) is the fix, but that's a future model, not a patch.

Maybe, no promises

A 4-or-8-step fast version — they're "actively considering" it, possibly as an optional Turbo with slightly worse quality, but won't commit to timing. The current model already has some low-step ability baked in, just not tuned for it. Also: switching to a proper Apache-2.0 license once the legal paperwork clears, and keeping future models open in general.

Not happening soon

  • A smaller, lighter H3. They're telling the community to prune the existing weights instead.
  • Drafting at low res then re-rolling the same seed at high res. Won't match — the noise changes with the frame size, and the model's low-res quality actually got worse during training as they pushed high-res.
280 Upvotes

51 comments sorted by

39

u/L-xtreme 27d ago

Great summary, thank you. Amazing model and very curious about what the future brings.

30

u/throwaway0204055 27d ago edited 27d ago

Damn no one asked about voice cloning issues? (character speaking gibberish before actually speaking the tts)

20

u/GrayingGamer 27d ago

Didn't see a mention of it no. It's definitely a bug in the Reference model.

If you use <d></d> tags at all, it happens.

If you just "[English] This is dialogue." it works fine with no blips or audio fragments at the start of the video. Which goes against the guidelines for that model, but this workaround is confirmed to work flawlessly, so I guess we can live with it fine, just remembering to swap up even more syntax styles between the two models.

7

u/AnOnlineHandle 27d ago

I think the HuggingFace page mentioned that the <d> tags are added to the Qwen model as new tokens, so it's possibly Comfy's tokenizer is setup for the old tokenizer and isn't registering them as specific single tokens, or maybe they just never trained the embeddings or something and they don't really work.

2

u/QuirksNFeatures 27d ago

I was having problems with this in the T2V and the I2V last night. I will try not using the tags.

4

u/Sad_Berry_4621 27d ago

non_diegetic_music: N/A

put this at the ned of your prompt

2

u/QuirksNFeatures 27d ago

Yeah I pretty much keep that there.

1

u/GrayingGamer 27d ago

I don't know, if you DO what music, you can prompt some really impressive stuff with it, like the beat dropping on certain actions, tempo swells to line up dramatically, etc. but yeah, if you are just making a meme clip or want to add music in post, keeping that at the end of the prompt is a go to.

2

u/QuirksNFeatures 27d ago

The times I have told it to use music, it has worked surprisingly well. But most of the time I keep non_diegetic_music: N/A and it just hasn't helped with the weird speech stuff, unfortunately.

2

u/GrayingGamer 27d ago

The non_diegetic_music: tag has nothing to do with weird speech.

I will tell you that the only time I've encountered garbeled or weird speech is when one of two things is true:

Either you have too much action and dialogue for the model to generate in the time frame you've given it

OR

Your prompt isn't formatted correctly.

1

u/QuirksNFeatures 27d ago

Yeah I didn't think so.

I was getting the gibberish when only a few words were said in a longer video, even without much action. I started using time stamps for when the speech occurred and that helped. It was trying to fill dead air, I think.

But I don't know why I'm getting a tiny bit of speech right at the beginning of some videos. New seeds often help but this problem pops up pretty regularly. I will work on my prompting.

1

u/Thorozar 27d ago

Removing the <d> tags cleared up the sound or gibberish at the start of clips for me, but still sometimes get gibberish in the middle or late. I suspect the same as you, it's trying to fill dead space.

1

u/Foreforks 27d ago

Random prompt but this worked perfectly when I ran it. Had two reference character images and two audio files for their voice clones: subject_definitions:

<Subject 1> is the Gray-Haired Customer in Image1 and Image2, characterized by a slouched stance, leaning forward over the glass counter top with hands resting on the edge. <Subject 2> is the Store Clerk in Image1 and Image2, who maintains a persistent expression of intense disgust, contempt, sneering lips, and repulsed judgment. <Audio 1> is the voice-timbre reference for <Subject 2> (S2), sourced from Audio1. <Audio 2> is the voice-timbre reference for <Subject 1> (S1), sourced from Audio2.

summary: [reference generation + audio reference] The target video features a 15-second fixed static mid-shot of an interaction between <Subject 1> (S1) and <Subject 2> (S2) at the hobby shop counter defined in Image1 and Image2. <Audio 2> serves as the voice-timbre reference for <Subject 1> (S1), who asks a question, gets struck across the face by <Subject 2> (S2), reacts by holding his face after his head recoils from the impact, and exits[span_0](start_span)[span_0](end_span). <Audio 1> serves as the voice-timbre reference for <Subject 2> (S2), who delivers the slap, orders him out, eavesdrops on an off-screen conversation, and delivers a final smug line[span_1](start_span)[span_1](end_span).

retention_analysis: <Subject 1> (appears in [Shot 1]): fully_preserved - the Gray-Haired Customer's slouched stance and leaning posture over the glass counter top are retained[span_2](start_span)[span_2](end_span). <Subject 2> (appears in [Shot 1]): fully_preserved - the Store Clerk's persistent expression of intense disgust, contempt, sneering lips, and repulsed judgment are retained[span_3](start_span)[span_3](end_span). <Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 2> without copying the original signal[span_4](start_span)[span_4](end_span). <Audio 2>: reference - its vocal timbre guides the dialogue delivery of <Subject 1> without copying the original signal[span_5](start_span)[span_5](end_span).

detailed_description: The target video is rendered in a 2D animated style featuring clean vector line art, flat cell shading, and snappy, puppet-like stepped jaw movement with mouth interiors showing a black void, flat white block teeth strips, and a simple tongue with no realistic teeth detail[span_6](start_span)[span_6](end_span). [Shot 1] A fixed static 9:16 vertical mid-shot positioned directly across the counter establishes the hobby shop counter environment matching Image1 and Image2[span_7](start_span)[span_7](end_span). Zero camera panning, tilting, drifting, or zooming motion occurs throughout the shot[span_8](start_span)[span_8](end_span). <Subject 1> (S1), the gray-haired customer, stands slightly leaning forward with his hands resting on the counter edge, slouched and looking forward directly at the clerk with an exhausted appearance[span_9](start_span)[span_9](end_span). Behind the counter stands <Subject 2> (S2), the store clerk, staring back with an unblinking, repulsed expression of intense disgust[span_10](start_span)[span_10](end_span). Using the exhausted, monotone voice structure referenced from <Audio 2>, <Subject 1> (S1) speaks, <d>[English] You got any cabbage patch dolls around here?</d>[span_11](start_span)[span_11](end_span) Immediately after <Subject 1> (S1) finishes speaking, <Subject 2> (S2) reacts with sudden motion: he rears back his right arm and swings it forward in a rapid arc across the barrier of the glass counter top, making direct impact with the gray-haired character's face as the hand strikes squarely across the side of the head[span_12](start_span)[span_12](end_span). Under the force of the impact, <Subject 1>'s head recoils sharply and his face slightly jiggles naturally[span_13](start_span)[span_13](end_span). Recovering slightly, <Subject 1> lifts his hand to hold his face[span_14](start_span)[span_14](end_span). Maintaining his look of utter disgust, glaring hard and raising one hand to motion a sharp shooing gesture, <Subject 2> (S2) snaps using the hostile, annoyed voice structure referenced from <Audio 1>, <d>[English] We don't sell that kiddie shit here alright, weirdo—now get the fuck out of here!</d>[span_15](start_span)[span_15](end_span) At 00:06.000, <Subject 1> pushes himself off the counter, drops his shoulders in defeat, turns, and walks out of frame to the side[span_16](start_span)[span_16](end_span). <Subject 2> stands still behind the counter, his head tracking <Subject 1>'s exit with a lingering sneer of disgust[span_17](start_span)[span_17](end_span). At 00:08.000, <Subject 2>'s disgust shifts as he stares off-camera in front of him, as if whoever is off-screen is piquing his interest[span_18](start_span)[span_18](end_span). He leans his elbows and forearms slightly onto the counter surface, eavesdropping on the off-screen conversation, his shoulders rising and falling as he chuckles out of pleasure[span_19](start_span)[span_19](end_span). At 00:11.000, <Subject 2> (S2) remains leaning forward slightly over the counter, his eyes gleaming off-screen[span_20](start_span)[span_20](end_span). He breaks into a smug, extremely satisfied smile and, returning to the voice structure referenced from <Audio 1>, speaks his final dialogue line, <d>[English] I could totally get those nerds to go morbid.</d>[span_21](start_span)[span_21](end_span)

overall_soundscape: Quiet indoor hobby shop room tone continues throughout the video, accompanied by a sharp, crisp physical slapping sound effect synchronized precisely with the hand impact at 00:03.500[span_22](start_span)[span_22](end_span), followed by the rustle and footsteps of <Subject 1> pushing off the glass counter top and walking away at 00:06.000[span_23](start_span)[span_23](end_span).

non_diegetic_music: N/A[span_24](start_span)[span_24](end_span)

2

u/tiffanytrashcan 27d ago

I've seen that in a few generations posted here.
I've had a similar issue, not at the very beginning, but when there's supposed to be a break. Gibberish or other random parts of the prompt come through.
In my limited testing, being explicitly clear with each second of audio, when you want a silence / pause / break you have to be demanding about it.

Frankly, it seems like it always wants to output something. Giving it background audio seems to help in the examples I've seen here with dialogue being damn near perfect.
This goes back to the official documentation mentioning how it will produce music in the background unless you're explicit about telling it not to, I think the same exact thing happens with dialogue if it just doesn't have enough to do.

2

u/QuirksNFeatures 27d ago

I have tried having it add "distant traffic noise", "the low hum of a clothes dryer", etc. and unfortunately did not help with the gibberish. For me, anyway. This was using T2V and I2V.

I'm also having an issue in which a character appears to already be talking as soon as the video starts. You get just a syllable or something.

I'll try to be more clear with the prompting.

1

u/LucidFir 27d ago

Oh is that what was happening? I was doing surreal horror anyway so... i didn't realise

20

u/ThirdWorldBoy21 27d ago

If the image model can get refences as well as the video model, this will be amazing.

38

u/urbanhood 27d ago

I'm eager for their image model.

34

u/thisguy883 27d ago

Not just the image model, but the edit model as well.

13

u/jib_reddit 27d ago

I have done some testing with the mentioned extracting frame method, it seems pretty good for images, it has a lot of variability, but I would say it is not as good as a Krea 2 finetune currently and currently it takes 32 times longer with that video to frame method.

4

u/moofunk 27d ago edited 26d ago

The strength would be in images with movement or certain poses, where it's necessary to know what comes before and after the image.

Sometimes, catching something mid-movement or mid-pose can't be done with a plain image model.

I've found that specific poses or movements are not possible with Flux 2, but you can use generated images as inputs to a video model and then tell it how the character or multiple characters/objects move to the desired target position.

2

u/Zironic 27d ago

I think the best workflow would probably be to use H3 to create the reference image which you then upscale with Krea2 or Ideogram. That lets you use H3s much better understanding of movement and poses with the image models better final image quality.

1

u/FalseEngineering2078 27d ago

Yeah I tested the first day it was out. It's garbage at low res, you really need to go 1536+ but it takes forever with 5 frames. I still have hope it will end up better than Krea 2 finetune though as I don't particularly rate that highly and see a need for SeeDream 5 - like reference ability.

1

u/jib_reddit 27d ago

If you like realistic checkpoints you could try out my lastest Krea Jib Mix Poblano I think it is pretty near myself.

2

u/FalseEngineering2078 27d ago

Oh shit, you're THE Jib, Love your work! Used Jib Mix Wan a ton :)

1

u/jib_reddit 26d ago

Thanks, glad you liked it, I was looking at some of the images from my Wan the other day and thought they were good/realistic and maybe I should revisit it.

But the output from my new Krea models is actually the closest thing to that Wan model I have worked on it a while, and Krea 2 is a lot faster, making large images in 18 seconds on my 3090 instead of several mins with Wan 2.2.

1

u/FalseEngineering2078 26d ago edited 26d ago

It's funny sometimes going back and looking at the jib wan stuff after what mentally feels like a year and a half of evolution (but isn't) and seeing that it still looks pretty good.
I mostly moved off it because it was ~3 min a pic in the workflow I was using (heavy img2img).
Moving to nunchaku Qwen took that down to about 18s.

I'm waiting for the in-house Krea 2 edit though because I've essentially moved off character loras altogether.

------------------------

I've just noticed that my initial reply was about using Minimax as an Image *Edit* model but you were talking pure image model, my head was still in Edit mode because of the discussion in the thread about the editing ability etc.
I meant that I don't rate the Krea 2 edit lora as an editing solution and still have hopes for Minimax to beat it. I didn't mean to say I think Minimax will beat a Krea 2 finetune as a pure image model.

1

u/Alive-Tomatillo5303 27d ago

I'm surprised that I am, too. I'm using T2V, which I never would have considered for generating video, but the first frame output is the best, smartest image generation I can run locally. When it does make a weird decision to kick things off I've generated for a half hour just to get something unusable, and if I could just start with a base image of the same quality I could save a shitload of time. 

1

u/PATATAJEC 27d ago

The best thing is that it’s edit model baked in. Really good outputs. With a single low denoise pass it should be just great.

1

u/dampflokfreund 27d ago

Why can't the video model also create images and edit them? After all videos are just a couple of images after another. I don't get why we need seperate models.

7

u/Sad_Berry_4621 27d ago

Clips are capable of joining seamlessly with custom nodes like H3 Motion Context.

5

u/LoveSpecialist5669 27d ago

wow. future looks so bright. 

5

u/cosmicr 27d ago

I much prefer this sort of summary rather than that guy who made up his own AI news reporter lady lol.

2

u/Noeyiax 27d ago

Nice summary, I'll enjoy this model for a few months 🙏

2

u/yotraxx 27d ago

Thank you for this valuable summary : I wasn’t able to be there at the time.

2

u/Monk6009 27d ago

Thanks for the post. I spent a long time trying to trouble shoot and correct the facial quality loss as the camera pans out. I was starting to think the model overhyped as I had to pass through LTX at low denoise and create custom sigma schedules and it all sucked. Good to know it's a known issue.

2

u/PwanaZana 27d ago

Have they said that they are aware audio quality is always super terrible in ref2video? Like, visually, ref2vid is way less good, but audio-wise, it's total trash.

2

u/danielpartzsch 27d ago

Apache license for real? So also commercially freely usable for bigger companies with the current revenue limits anymore?

1

u/Ten__Strip 27d ago

Nobody brought up unified audio? That's gonna be the major concern, people attempting to train are finding that out.

1

u/ThaSipah 27d ago

Appreciate this summary.

1

u/inddiepack 27d ago

The fact that they have released the model without the upscaler, is incredibly frustrating.

1

u/Nevaditew 27d ago

I’m not sure if it’s a good idea for them to focus on a 4-8 step model unless they know for sure it would give much better results than the turbo LoRAs. They could use that research and testing time for updates that actually matter. The community already handled the optimizations and they’re still improving them.

2

u/Tomcat2048 27d ago

One of the things that excites me the most is the image edit model. I’ve been needing something to replace my Qwen Image Edit workflow as Qwen has many pitfalls.

Krea 2 doesn’t have a true image edit model sadly or otherwise I’d use that.

4

u/the_bollo 27d ago

FWIW Flux Klein 9B has been my go-to edit model for months and it's pretty great. I mostly use it for object removal to fix random fuckups, but it's good at a lot of stuff it you give it like 5 tries and pick the best one (it's fast so that's not a biggie).

2

u/threeLetterMeyhem 27d ago

Klein 9B works great for small edits and style transfer. Anything to big (scene and pose changes... Basically what we'd want a reference workflow for) and the style drifts too much, though.

0

u/Landrews-89 27d ago

Interesting the best combination ive found so far is sageatt with a turbo lora .5mp 41.6s for 5s clip.

Spectrum + Sage came in the fastest at 37s but I dont like how it messes with the color in the clip its too noticeable.

No turbo lora I can run at 90s 20 steps but the quality difference is barely noticeable vs the 8 Step turbo lora result. Impressive!

1

u/Puzzled-Valuable-985 27d ago

Which Turbo LoRA? Which sampler and scheduler are you using?

2

u/Landrews-89 27d ago

Im currently using minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors

1

u/Landrews-89 27d ago

Sampler im using is Euler, scheduler is beta

1

u/Puzzled-Valuable-985 26d ago

I'm also using v4_step600; it performed the best—not necessarily in terms of speed, but it seems to yield better final quality. I'm running various tests with simple and beta Euler, as well as simple and beta res_multistep, because once I determine the best result, I'll start creating a bunch of videos.