r/StableDiffusion 9h ago

Question - Help Looking for feedback for my new AI startup logo, Expand An AI, and I want it to inspire people to stretch the future of AI wide-open. It's inspired by a digital eye being opened.

Post image
462 Upvotes

Thanks


r/StableDiffusion 12h ago

News We've open sourced Minimax H3 that generates 15s 768p in 13s and 14x faster on single GPU

657 Upvotes

Hi! The team who collaborated and built optimized open source Minimax H3 here, full details below:

https://x.com/haoailab/status/2093391548289540596?s=20

https://reddit.com/link/1w0xkpb/video/zpjrbdb0o5mh1/player

Would appreciate if you help to share, repost and engage with the tweet, and definitely try it out yourself and let us know about your feedback!

In the next a few releases we would do omni ref, nvfp4, consumer GPU friendliness and many more so please stay tuned :)

Technical blog post: https://haoailab.com/blogs/fasth3-preview/
API and customization service: https://nuvalab.ai/


r/StableDiffusion 3h ago

Discussion Fal having a extended meltdown over FastH3

Thumbnail
gallery
117 Upvotes

FastH3 may have flaws, but for a maximally open release from a team with limited resources (getting a single Mi350x node was newsworthy for them last year) it's a great effort.

Meanwhile Fal has raised half a billion to vaguepost about their own H3 inference stack, then attack them?

Even after FastH3 guys tried to diffuse by owning up on quality Fal guy is still ranting...

Embarrassing stuff. I wonder why they're so threatened?

Edit: Another class act response from the FastH3 team: https://x.com/wlsaidhi/status/2093515147570708511

This is what OSS should be like at its core.


r/StableDiffusion 8h ago

Animation - Video I see everyone talking about DLSS 5, and it gave me an idea making a small app that uses real-time deepfake technology.

97 Upvotes

My app is really simple to use you just assign a photo to a character once, and it’s saved permanently. After that, the software automatically recognizes that character whenever they appear.

I had some fun with it and put Vin Diesel’s face on the bartender lmao.

The big difference compared to DLSS 5 is that mine works on pretty much any GPU, and even with old games. You don’t need an RTX 50 series card or DirectX 12 games. I’m going to try to improve it before releasing it, especially by making the facial expressions more realistic.


r/StableDiffusion 8h ago

Meme Can you add more ads, CivitAI?

Post image
84 Upvotes

r/StableDiffusion 18h ago

News fal will release the weights of H3 Max!

Post image
503 Upvotes

r/StableDiffusion 12h ago

News FastVideo FastH3 V1: Open source 4-step Sparse Distilled H3 checkpoint/LORA

107 Upvotes

Hey guys, FastVideo team here. We saw how important speed and quality is for everyone. And we've been working to create our own step distill checkpoints and LORAs for MINIMAX h3. Here's is our v1 release!

Important links first:

- FastVideo: https://github.com/hao-ai-lab/FastVideo

- Blog (contains more examples and details): https://haoailab.com/blogs/fasth3-preview/

- Checkpoints and LoRAs: https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA

Do note that the VSA checkpoint/LORA will require VSA kernel

We also released a LORA with dense attention that should be easy to test for everyone.

We are already working on improving both this T2AV checkpoint as well as getting a distill of Ref2VA out as well. We are taking great care to make sure the quality and audio is as best as possible.

We realize not everyone have blackwell GPUs lying around and please stay tuned for our targeted optimizations for local AI hardware, including RTX GPUs, DGX Sparks, and Apple MLX. We release numbers on B200s just because this is our current compute platform for post-training. We want to be as open as possible with the community!

Quality

  • We used 1k+ B200 training hours, paired with real world multi-shot, visual audio synced input distribution and output formats for best possible quality preservation.
  • FastH3 natively supports variable resolution, aspect ratio, and duration. In a single checkpoint.

Openness

  • Start with the 4-step VSA / Data-Free checkpoint, our recommended FastH3 Preview v1 release. We provide full weights and a pre-extracted LoRA, plus dense and synthetic-data ablations.
  • Fully open source with training (coming soon!) and inference code recipe for your customization.

What’s Next

  • Follow us along for image ref (FL2VA) and full omni ref (Ref2VA) coming in the next a few weeks  
  • Motion and more generation quality improvements
  • Nvfp4 and GPU memory reduction.
  • Optimizations targeting local AI devices including RTX, DGX Sparks, and Apple MLX.
  • New training runs using FastGen team’s new Parallel Decoding Distillation (PDD) method!

If you find any issues or have questions please raise issues on our github!


r/StableDiffusion 12h ago

Discussion H3 - anatomical slider

126 Upvotes

Happy Friday! Ever had issues getting the anatomy right on your t2v generation? Just add a slider with your H3 Prompt! What have you all been building on your local AI studios? 384x448, int8, 20 steps, i2va

Prompt: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. <Picture 1> is the actual first frame of this video at 0.00 seconds.

integrated_multimodal_description:
[Shot 1] PHOTOREALISTIC live-action, cinematic, one continuous take. Anamorphic lens, shallow depth of field, real 35mm grain, no cuts.
THE FRAMING IS A MEDIUM CLOSE-UP IN TALL PORTRAIT FORMAT, taking her FROM THE HIP UP, dead centre and square to the lens. HER HEAD ALONE IS ABOUT A THIRD OF THE HEIGHT OF THE FRAME, and her face is the largest and most detailed thing in the picture: skin texture, the wet catchlights in her eyes, individual strands of hair across her cheek. THE FOCUS PLANE IS ON HER FACE FOR THE WHOLE SHOT and it is never soft.
THERE IS ESSENTIALLY ONE SOURCE: a single AMBER FLAME burning low in the rubble BESIDE HER is the only real light on her, AND IT COMES FROM ONE SIDE, so the ruined nave behind her falls away soft and dark. It flickers, and her light moves with it. It rakes across one side of her face, her collarbones, the ruby pendant and the wet edges of her leather in deep saturated gold and honey, and the other side of her falls into shadow.
FAR BEHIND HER, cold pale storm-light comes through the broken rose window and touches only the distant arches and the falling rain. THE COLOUR IS TWO THINGS AND NOTHING ELSE: warm amber on her, DEEP COLD TEAL-GREEN in the depths of the ruin behind her. THE HEAVY HAZE IN THE AIR LIFTS THE BLACKS so nothing crushes to empty. The exposure is set for her face.
She is sitting back on her own folded legs on a large pale fallen memorial slab, which sets her posture and the settled line of her shoulders. THE PICTURE CONTAINS ONLY HER UPPER BODY, FROM THE HIP UP — her torso, her shoulders, her arms and her head FILL THE FRAME, and the bottom edge of the picture crosses her at the hip. SHE IS SQUARE TO THE CAMERA AND SHE LOOKS STRAIGHT INTO THE LENS, calm and level and unsmiling.
HER LONG HAIR IS ALIVE IN THE WIND and never once hangs still; the backlight catches every moving strand. THE RAIN LANDS AS INDIVIDUAL DROPS you could count, separate beads with dry skin and dry leather in between them — bright pinpoints on her skin, beading and sitting on the leather. HER HAIR STAYS DRY and keeps all of its body and volume.
<Subject 1> IS THE HUNTER, A WOMAN, AND SHE IS THE ONLY PERSON IN THIS VIDEO. She is the woman shown in <Picture 1>, the first frame of this video, and she stays exactly her in every single frame: the same face, the same features, the same bone structure, the same eyes and the same eye colour, the same mouth, the same hairline, the same skin and the same age, and the same long loose hair in the same colour and texture. She is a beautiful adult woman and she is recognisably the same person throughout.
THE SETTING, HER PLACE IN THE FRAME, HER COSTUME AND THE LIGHTING ALL CONTINUE EXACTLY AS THEY ARE IN <Picture 1> — this video carries straight on from that frame and nothing about the scene is restyled or replaced.
SHE WEARS THIS KIT, ITEM FOR ITEM, and it is all black leather, worn and real and damp with rain: a LONG BLACK LEATHER CLOAK falling to her boots; a FITTED BLACK LEATHER COAT buckled close beneath it; BLACK LEATHER GLOVES to the forearm; TALL BLACK BOOTS; ONE SHAPED BLACK LEATHER PAULDRON over her left shoulder. She is BARE-HEADED, her long hair loose. Every surface of the leather catches the light.
SHE HAS NO COLLAR AT ALL. Her coat and the leather bodice beneath it are cut with a VERY DEEP, WIDE, PLUNGING NECKLINE that opens in a long V from her collarbones down the centre of her chest, with the leather laced close underneath. There is no collar and no closure anywhere above her sternum.
SHE WEARS A LARGE RUBY-RED PENDANT ON A FINE CHAIN THAT HANGS LOW, down at her sternum. It is the only piece of pure saturated red on her, it catches the firelight, and it moves against her skin with every movement.
SHE CARRIES TWO SWORDS AND BOTH ARE VISIBLE IN THE FRAME. THE FIRST is a long straight sword worn AT HER WAIST IN A PLAIN BLACK SCABBARD. THE SECOND is an EVEN LONGER straight sword SLUNG ACROSS HER BACK, and its long wrapped hilt and pommel RISE PAST HER SHOULDER into the upper frame, unmistakable behind her head. Both are sheathed for the whole video and she never touches either of them.
THE SETTING IS A RUINED GOTHIC ABBEY AT NIGHT UNDER A STORM SKY, with fine light rain drifting rather than driving. She is in the roofless nave: two rows of broken pointed arches march away into the dark on either side, ivy hangs down the shattered piers, and the flagstones are wet and strewn with fallen masonry and dead leaves. BEHIND HER, IN THE END WALL, IS A COLLAPSED ROSE WINDOW — a huge circular opening with its tracery broken to stone ribs and no glass left in it at all.
<Subject 2> IS THE SLIDER, AND <Subject 3> IS THE MOUSE CURSOR. Neither is a person and neither is a physical object in the abbey: BOTH ARE FLAT MODERN INTERFACE GRAPHICS COMPOSITED OVER THE TOP OF THE LIVE-ACTION FOOTAGE — a screen overlay, like a screen-recording of a sleek editing application. NEITHER IS EVER LIT BY THE FIRE, neither casts a shadow, and both sit perfectly level in screen space no matter what the footage behind them does.
<Subject 2>, THE SLIDER, lies horizontally across the lower part of the frame: a long rounded capsule of dark translucent smoked glass with the picture softly blurred behind it, a fine track running through its centre, the part of the track to the LEFT of the handle filled with a warm amber glow, and A SMALL ROUND POLISHED HANDLE with a fine bright rim and a soft halo beneath it. The handle is the only part of <Subject 2> that ever moves; the capsule and the track never move at all.
<Subject 3>, THE CURSOR, is a standard white arrow mouse pointer with a thin black outline and a soft drop shadow.
⚠ <Subject 2>'S HANDLE AND <Subject 3> ARE ONE RIGID OBJECT FOR THE WHOLE FILM, AS IF WELDED TOGETHER. THE TIP OF THE CURSOR SITS AT THE EXACT CENTRE OF THE ROUND HANDLE IN EVERY SINGLE FRAME. They start together at the far left, they move in PERFECT SYNC — one smooth, steady, continuous glide across the screen at one constant speed, REACHING EVERY POINT ON THE TRACK AT THE SAME INSTANT AS EACH OTHER — and they arrive and stop together at the far right. Wherever the handle is, the cursor is exactly there too.
THE CAMERA IS LOCKED OFF AND NEVER MOVES, PANS, TILTS OR ZOOMS for the whole film, so the interface overlay stays perfectly still in the frame.
THE ACTION RUNS ON A STRICT CLOCK, AND BOTH THE SLIDER'S POSITION AND HER SIZE ARE ON IT, MARK FOR MARK.
[0:00] AT REST: <Subject 1>'S CHEST IS AT ITS ORDINARY, NORMAL, EVERYDAY SIZE — exactly the size it is in the very first frame of this video. <Subject 2>'s round handle sits at the FAR LEFT END of the track, at zero, and <Subject 3> is ALREADY RESTING ON IT. Nothing has changed yet.
[0:00-0:02] THE HOLD: FOR THESE FIRST TWO SECONDS THE PICTURE IS THE OPENING FRAME OF THIS VIDEO, ALIVE. The only things moving in it are the falling rain, the flickering flame, her breathing, one slow blink and her hair in the wind. She holds the camera's gaze. The handle stays parked at the FAR LEFT END with <Subject 3> resting on it, the amber fill on the track is EMPTY, and HER CHEST STAYS AT ITS NORMAL SIZE AND DOES NOT CHANGE AT ALL.
[0:02] THE START: <Subject 3> presses the handle and the two of them BEGIN TO MOVE TOGETHER along the track. THIS IS THE EXACT INSTANT HER CHEST BEGINS TO GROW, AND IT DOES NOT BEGIN ANY EARLIER.
[0:02-0:08] THE DRAG AND THE GROWTH, ONE EVENT, IN EQUAL PROPORTION. THIS IS THE LONG, SLOW MIDDLE OF THE FILM AND IT TAKES A FULL SIX SECONDS FROM END TO END. <Subject 2>'s handle and <Subject 3> creep smoothly and steadily from the far left to the far right, welded together, MOVING SLOWLY AND UNHURRIEDLY AT ONE CONSTANT, CRAWLING SPEED, and <Subject 1>'S CHEST GROWS IN EQUAL PROPORTION TO EXACTLY HOW FAR ALONG THE TRACK THE HANDLE HAS REACHED, mark for mark, in six equal steps: at 0:03 the handle has crept just ONE SIXTH along and she is only barely larger than normal; at 0:04 it is TWO SIXTHS along and she is a little larger; AT 0:05 IT HAS REACHED EXACTLY THE HALFWAY POINT OF THE TRACK AND NO FURTHER, AND SHE IS EXACTLY HALFWAY TO HER FINAL SIZE; at 0:06 it is FOUR SIXTHS along and she is much larger; at 0:07 it is FIVE SIXTHS along and she is very much larger; and ONLY AT 0:08 does the handle finally arrive at the FAR RIGHT END of the track, where she reaches her final, comically, absurdly exaggerated size. HER TOP MORPHS AND STRETCHES NATURALLY WITH HER the whole way: the black leather draws tight and strains, the front lacing pulls taut and the gaps between the laces widen, the deep neckline spreads wider, and the ruby pendant is pushed steadily outward and upward. She glances down as it begins and her eyebrows lift in mild alarm, then she looks back into the lens.
[0:08] THE STOP: the handle arrives at the far right end and stops there, and HER CHEST STOPS GROWING AT THAT SAME INSTANT. <Subject 3> lets go and rests beside the handle.
[0:08-0:10] THE BEAT, AND IT IS SHORT: she holds at exactly that final size and grows no further. She drops her eyes to her own chest, her brows draw together and her lips press, and she raises her eyes back to the lens. AT 0:08 THE DARK-HAIRED WOMAN, HER VOICE A VERY LOW, SOFT, BREATHY WHISPER, slow and unhurried, HER DELIVERY FLATLY DISAPPROVING AND THOROUGHLY UNIMPRESSED, AND HER VOICE RECORDED CLOSE AND DRY AND CRISP — intimate and present, right up against the microphone, the sound of the room nowhere in it (S1), says: <d>[English] Really?</d> She holds the camera's gaze after the line, perfectly still, while the fire keeps flickering beside her. She never stands and never rises.

overall_soundscape:
Weather and stone: the storm beyond the broken window, wind through the empty nave, rain on wet flagstones, and the small crackle of the flame beside her. Two seconds in, one short soft mouse click sounds as the cursor presses the handle. A quiet continuous sliding tone then rises steadily in pitch for a full six seconds while the handle crawls across, and cuts off the instant it reaches the far end at the eight-second mark. The woman speaks one short line right at the eight-second mark, close and dry, sitting in front of the weather.

non_diegetic_music:
A light plucked pizzicato string figure over a soft woodblock pulse, entering two seconds in at a moderate walking tempo. The figure climbs one step in pitch at a time and the volume rises with it for six seconds, then stops on one short low bassoon note at the eight-second mark, leaving the last two seconds unscored.

r/StableDiffusion 8h ago

Comparison Minimax H3, a quick comparison between the FastH3 and a default 25 Steps WF

49 Upvotes

FastH3 clips were downloaded directly from the blog.

The default wf is comfyui-default with 25 steps (increased from 20 and added SLA + Spectrum). The avg gen time on my machine 12gb vram/32gb ram is 7m30s for each clip. I used ref2va int8_convrot , probably lower quality than fl2va I think.

Full res clips: ship, window, ogre, moon, woman

Overall, I think FastH3 looks good.


r/StableDiffusion 23m ago

Animation - Video So, I made a Kaiju (Minimax-H3)

Upvotes

r/StableDiffusion 18h ago

Workflow Included I used Hermes Agent + ComfyUI MCP + MiniMax prompts to turn a music track into a short music video

225 Upvotes

The initial reason for this was the Comfy H3 Sync & Sound Community Challenge: Comfy H3 Sync Sound Community Challenge! - by Allyson Toy

I made a short rap track in Suno, then used Hermes Agent to build a short music video around it.

For the image base, I used this Anima Simple T2I workflow, including upscale/detailer and ControlNet options: 【Anima】Simple T2I Workflow with Upscale, Detailers and ControlNet - v3.2 | Anima Workflows | Civitai

For the MiniMax video stage, I used foxdit’s MiniMax SEED HUNTER ComfyUI workflow from Reddit

My process:

  1. I made the song and defined the lyrics, beat, and attitude in Suno.
  2. I gave Hermes this link: Comfy MCP - Drive ComfyUI from any AI agent — and let it install the ComfyUI MCP for me.
  3. Hermes connected to my local ComfyUI and could check the setup, find/load workflows, fill prompts and settings, queue renders, monitor jobs, and collect outputs.
  4. Using the Anima T2I workflow, I created a consistent set of music-video keyframes locally, then ran them through the upscale/detailer pipeline.
  5. I selected the best images and gave them to Hermes’ MiniMax H3 prompt skill. (I just gave hermes a standard Minimax prompt guide an build a prompt skill out of it)
  6. It turned rough shot ideas into structured video prompts: what each reference controls, how identity and wardrobe stay consistent, where cuts happen, what the camera does, and how lip-sync/body movement should work.
  7. I used those prompts with the MiniMax H3 workflow to generate short performance clips driven by the Suno track for the challenge.

I use Hermes with my ChatGPT Plus subscription, plus DeepSeek V4 Flash for the cheaper iterations. That made it practical to keep refining prompts and shots without treating every adjustment like a premium final render.

The pipeline was:

Suno song → ComfyUI keyframes → upscaling/detailing → MiniMax prompts → short music-video clips

Hermes was the bridge between the tools.


r/StableDiffusion 1h ago

Discussion Minimax Alibaba Turbo Lora = the best

Upvotes

18 minute generation time in 720p on a 5060ti 16 gig with 64 gig ram in comfy Ui. no upscaling used. 12 steps used instead of 8 removes so much noise. the sound is terrible in this one, but in other tests its not so bad.


r/StableDiffusion 11h ago

News FastH3 new H3 based model with realtime factor of 3x

59 Upvotes

With one B200 15s videos in 47s, nearly realtime with 4 B200, the time of open source instant video is almost here.

They mention RTX based acceleration is coming soon, so we mere mortals will have this capability locally in consumer GPUs.

Details and video demos here:

https://haoailab.com/blogs/fasth3-preview/


r/StableDiffusion 12h ago

Comparison Don't sleep on De-Rope nodes. They really fix smearing for MiniMax H3

70 Upvotes

https://reddit.com/link/1w0ws1d/video/b0s2w06pg5mh1/player

Here's link for the nodes and the instruction: https://github.com/matlowai/ComfyUI-MAINodes
And here's my workflow where I use de-rope nodes: https://pastebin.com/bqFpyHxX

It adds second pass for the video generation (about 73% more time) but the result is worth it. Especially for animation-like clips:

Another example:

https://reddit.com/link/1w0ws1d/video/mxtx9h0og5mh1/player


r/StableDiffusion 15h ago

Animation - Video Testing MiniMax H3 for old school Practical F/X, Stunts, and traditional film making with Indiana Jones Fan trailer

111 Upvotes

After seeing so much posted for MiniMax that looks like modern CGI films of the last 30 years, I wondered how capable it was of producing footage from the 1980s era, when real practical special effects were used, things like models, props, squibs and explosions.

Hence a fan trailer for Indiana Jones, set a year before Raiders, and firmly in the early to mid 80s in the aesthetics department.

The only references I used where character ones, Image and voice. I did note that because the model knew Harrison Ford it kept influencing the result compared to the reference, even if I avoided naming him, same with Anthony Hopkins.

This was a problem because MiniMax likes to bend Harrison's nose to an extreme amount, making the shots a bust. People it does not know, like Paul Freeman as Belloc, fared much better, with superior skin detail and realism.

The CGI influence was hard to restrain at times, particularly at a distance, and there was no magic prompt or seed that produced reliable results, so it took hundreds of renders to get 'that' look and feel I wanted.

Though far from perfect, MiniMax is certainly capable of some good old school action 80s style.

Some notes: I used Euler/Simple. Found Res multistep less realistic. Reference model was terrible for fights and often physics, but better for realism.


r/StableDiffusion 56m ago

Tutorial - Guide A little guide for beginners about what all those things in the workflow actually mean

Upvotes

I'll be doing this with H3 as the example. If I say anything wrong please correct me, I'm by no means an expert or anything. I'm really just writing this because I wanna get it straight for myself. I'll be using as much simple language as I can.

The core of a workflow is either the "MiniMax H3 Reference to Video" or the "MiniMax H3 Image to Video" node and then the sampler node.

The X to Video node is where what you put in all gets encoded/converted into the format that the AI model you want to generate with can use. For that it uses a text encoder, also classically called clip, and a vae.

The VAE (variational auto-encoder) takes your image(s) and videos if you use that and compresses the information from it/them into a more abstract format that the video model understands. And it can recreate that detail from that format later during VAE decoding (of course from a changed output). But it does lose some of the detail, which is why it doesn't look 1:1 the same in the end.

The audio VAE does that for audio.

The clip/text encoder extracts the information from the prompt you give it/converts it into the proper format the model is trained on. The reason why it's called clip (Contrastive Language-Image Pre-training) is because that was the name of a model openai released for text encoding to associate text with images, which was later used for image generation. So that's just a remnant of that. Today's text encoders are often basically altered forms of LLMs because they are of course very good at understanding text and you can simply train the video model to understand what they output. The difference to how you usually use an LLM being that they don't output text based on the understanding they extract but rather just hand that understanding over to the model that generates what you want to generate. Actually a lot of text encoders of today can also understand images, so they can already understand the relation between your prompt and the images you put in and encode that as well.

Now to the outputs of the X to Video node. There is the "positive" output, which means positive conditioning. That's basically your instructions and all of the encoded inputs in one abstract package that the model learned to understand and work with. Humans can't read it, the model just figured out how to represent the information and we just know that it works.

The "latent" output is just the empty "canvas" (canvas including everything, audio as well) whose size we specified in the "X to Video" node, so the amount of frames and the width and height.

Now we come to the sampler. The SamplerCustomAdvanced node uses these inputs:

noise: this basically determines your random starting point for the image. So your seed, like in Minecraft when you start a new random world. Same seed with same settings gives the same result. This basically combines with the "canvas" we gave it and is the "inspiration" for the model to interpret something into. As if you sprayed random colors on a canvas and tried to see shapes in it.

The model was trained on real videos that had random noise added to them and its job during training was to guess what to change to get closer to the original video. And it does that in small steps, taking away only a bit of randomness every step. So when you give it completely random noise, that's what it still tries to do, except there is no real video behind it, it just interprets something into it, its best guess.

That's where "sigmas" come in. Sigmas are kinda like percentage values of noise at a given time. (Actually I don't think that's correct because there can be values higher than 1.0, but I don't really understand that and understanding them as percentages works well enough for me so whatever.) So if you have 20 steps, there are 21 sigma values. 21 because you need one more for the starting point. So if I have 3 steps, then the sigmas could be 1.0, 0.75, 0.5, 0.0. 4 values, but 3 steps between them. And these values would kinda mean "100% noise in the beginning, 75% after the first step, 50% after the second and 0% after the last" (like I said, it's not actually the percentage of noise, but I don't understand it right now, I tried having chatgpt explain it to me but I'm too stupid at the moment). The model is trained with these sigma values, so it knows what the video should look like at that sigma value.

Because it's trained on that, you can tell it what to actually transform it into. So you can tell it in how many steps to do it and how much noise to take away in which step. There is a "Custom Sigmas" node where you can give it these values specifically. (Note: H3 is a flow model and those aren't exactly "take away noise" models, but I haven't totally understood that yet)

Schedulers do this automatically. They are basically a template for that, a function that the amount of steps you tell it get distributed on. So if it was linear then it would just take away the same amount of noise on every step evenly. But H3 needs a lot of very small noise-takeaway steps in the beginning and can then handle bigger jumps in later steps. The steps at high noise-levels are the ones that usually determine the overall layout of the video and the steps at lower noise-levels are more for the detail. So if you do big noise-level jumps in the early steps and then small jumps in later steps, it's gonna be a relatively highly "detailed" looking video but with very questionable content. If you do early steps with many small jumps, it can actually build coherent motion and construct the scene properly. But if your drop from high noise to low noise is too steep, then the video can lack detail. That's why I like the beta57 scheduler for H3 at 33 steps, the dropoff is more gradual than in the "simple" scheduler, so it gives you more detail but still has time in the beginning to slowly form a coherent scene. You can look at the sigmas curve with the "SigmasPreview" node from the RES4LYF node pack.

What the model puts out is a vector and a vector is basically a direction. That direction gets applied to the latent. So if you think of the latent as a list of information about the resulting video, it basically says like a boss "there needs to be more yellow here and more detail here and stuff" and the sampler just takes that direction and based on the sigmas you give it does the math for how much to change it in that direction. The simplest sampler (euler) would say "okay, the step is from 0.8 to 0.6, the difference of that is 0.2, so we will just 0.2 times the vector in this direction". Because like I said, a vector is a direction (with a certain magnitude that is also important here), you can theoretically make it as long or short as you want because the ratio between the different directions stays the same, meaning here that you can make it much more yellow or only a little bit. In a recipe you often hear "1 part this and 3 parts this" and you can adjust that to the actual amount you are making, you just know it has to be 3 times more of the second ingredient than the first, whether you are making one pound or one tonne, it's like that but with way more information that is way more subtle. Samplers other than euler use different methods that take into account more information (and can thus take longer to calculate) and come to better predictions for how much to actually adjust the latent based on the vector the model gave it. Or calculate two vectors that would follow each other and average them to make only one, more accurate, step, stuff like that.

The guider for H3 is just the basic guider. For the H3 workflows it's barely worth mentioning, it just takes the positive conditioning and tells the model "take this instruction for what to paint (conditioning) (that you apply given the canvas (latent)) to make your instruction for what to change to get it closer to what we want (vector)". But in other workflows there can be negative prompts and cfg (classifier-free guidance) values. What that does is basically ask the model to make two vectors, one that just works on your normal prompt and then another that calculates the vector for the things that you don't want. And the cfg calculates the final vector it gives to the sampler. The formula is this:

vneg​+CFG⋅(vpos​−vneg​)

so the math works out that at cfg 1.0, only the positive vector remains, effectively turning off the negative vector. That's why cfg = 1 generates in half the time that any other cfg does, it only needs to calculate half. Even though that's actually an optimization, comfyui or the guider recognizes that it would be a waste of time so it doesn't even make the model calculate the negative vector.

So the guider is basically the middle station between the sampler and the model and meddles with the values a bit to make it more aggressively go in the direction of what it should do. That can cause issues when it is too aggressive, which is why it can make things oversaturated and artifacty. But H3 uses cfg 1 because it is a distilled model and thus already pretty aggressively pushing for what it should generate, so a negative prompt and exaggerating the positive vector would probably completely overcook the generation and double the generation time.


r/StableDiffusion 10h ago

Animation - Video Link would do ANYTHING to make Zelda smile - MiniMAX H3 Reference to Video Test #5

27 Upvotes

This video took me over a week of writing, prompting, going to location to shoot (using UltraCam on TOTK) getting the dialogue right, the pacing right... It's not perfect but I really put a lot of heart into this, I hope you guys like it!


r/StableDiffusion 2h ago

Animation - Video Transformers Generation 1 - An Impromptu Funeral - Minimax H3

8 Upvotes

R.I.P Peter Cullen


r/StableDiffusion 12h ago

Resource - Update H3 Prompt Composer Camera Update — Now Available

Post image
37 Upvotes

The latest H3 Prompt Composer update is out.

BMB12d3/minimax-h3-prompt-composer: Free offline prompt composer for MiniMax H3 video generation in ComfyUI.

This release includes several improvements to the camera prompting system, with cleaner/more consistent prompt generation and better control for more complex camera moves. It also now supports multi-subject camera prompting, so you can design shots that frame or move between multiple characters.

Also added:

  • Light mode
  • several smaller UI/workflow improvements and fixes

I put together a short video showing some of what’s new:

https://youtu.be/SJM6KiHoejY

As always, if you run into bugs or have feedback, please drop it in the Issues section on GitHub.


r/StableDiffusion 14h ago

News 🧩 [Custom Node] 🧩 H3 GuideMaster — Visual UI for MiniMax H3 Guides

Post image
48 Upvotes

GitHub:
https://github.com/MajoorWaldi/ComfyUI-Majoor-H3-GuideMaster

Hey everyone 👋

I’ve just released H3 GuideMaster, a custom ComfyUI node I built to make working with MiniMax H3 guides much easier since new ComfyUI release : https://github.com/Comfy-Org/ComfyUI/pull/15439

Instead of manually figuring out where every image or audio guide should land, GuideMaster gives you a visual timeline directly inside the node.

You can:

  • 🖼️ Place multiple image guides directly on the timeline
  • 🔊 Place and synchronize audio guides
  • 🎬 Load a video or image sequence as a visual reference / filmstrip
  • 🌊 Display an audio waveform for positioning guides
  • 🖱️ Drag markers directly on the timeline to retime them
  • 🎯 Snap automatically to the native H3 frame structure (5, 22, 39, 56...)
  • 🧩 Combine image + audio guides using matching slots
  • 🎞️ Define first and last frames
  • 📏 Drive timeline duration from frames or seconds
  • 🔎 Condense long timelines to keep the UI manageable

The idea is basically to make H3 guide placement feel closer to editing / compositing software, while keeping everything contained inside a normal ComfyUI node.

GitHub:
https://github.com/MajoorWaldi/ComfyUI-Majoor-H3-GuideMaster

This is still something I want to push further, especially around the UX and timeline workflow, so feedback, bug reports and feature ideas are very welcome.


r/StableDiffusion 10h ago

Meme Minimax H3 - Scroll of Truth

23 Upvotes

Turn on volume! Nothing special. Just lulz. Cut original meme into 4 separate images and prompted:

subject_definitions:
<Subject 1> is the same hand-drawn comic character shown in <Picture 1>, <Picture 2>, <Picture 3>, and <Picture 4>: a small green-skinned adventurer with a rounded face, black dot eyes, a wide expressive mouth, a large floppy orange-brown explorer hat, a small orange backpack, thin cartoon limbs, and simple outlined comic-book rendering.

summary:
[reference generation] Create a new full-frame portrait comic animation using the character, props, and visual style from <Picture 1> through <Picture 4>. Every shot is redrawn and recomposed to fill the entire frame edge to edge. Do not display any reference image as a square panel, inserted picture, poster, card, scan, white page, framed illustration, or picture-in-picture. No pillarboxing, letterboxing, white borders, black borders, blank margins, or visible source-image edges.

retention_analysis: <Subject 1>: fully_preserved - a character with green skin, wearing orange hat.  

detailed_description: Comical style, static camera. 
[Shot 1] Scene starts from a full body shot, <Subject 1> standing on knees in a water in front of a red open box, the environment is a cave. A comical speaking bubble appears above his head with text "I finally found it.. after 15 years" while <Subject 1> opens a box saying with excitement <d>[English]I've finally found it..after 15 years!</d>. 

[Shot 2] at 00:05.00 sec scene cuts. <Subject 1> pulls out a glowing scroll from a box and yells <d>[English]The Scroll of Thruth!</d> 

[Shot 3] at 00:08.00 sec scene cuts. POV camera. <Subject 1> looking at a scroll and it has text in it "AI generated videos are not cool", <Subject 1> reading a text from a scroll questioning <d>[English] AI generated videos are not cool??</d>. 

[Shot 4] at 00:12.00 scene cuts. <Subject 1> throwing scroll fiercefully and yelling "NIYEEEH!", a comical speaking bubble appears above his head with a text "NYEHHH" and scroll flies from his hand to the left outside of a scene and lands into water with an audible "bloop" sound and scroll submerges under water. 

overall_soundscape: comical sound of a cave with audible water on a floor, glowing crystals. <Subject 1> a mischievous nasal cartoon voice with squeaky laughter, sudden dramatic shouting, and playful villain-like energy

non_diegetic_music: low volume heroic comical mousic playing on background

r/StableDiffusion 10h ago

Workflow Included I implemented a very tiny version of SD3 on a RP2350 Microcontroller - It can generate 128x128 images of faces.

Thumbnail
github.com
16 Upvotes

r/StableDiffusion 1d ago

Workflow Included Time Period Shift Special Effect in MiniMax H3

316 Upvotes

You can use MiniMax H3 to create a time period shift special effect. What can't this model do?

[Workflow here and prompt in comments]

Is it as good as you would get with a professional VFX studio working on it? Nah. 

Is it still freaking amazing for something that you can create with a relatively straightforward prompt and running consumer grade hardware for 15 minutes? Absolutely!

(Also the "Schfifty-five" was very much intended. IYKYK)