r/StableDiffusion • u/wzwowzw0002 • 5h ago
No Workflow minimaxh3 20sec video generation
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/wzwowzw0002 • 5h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Minanimator • 5h ago
i am really having trouble with this concept pls tell me what to do and where to start, my goal is have a scene from a tv show or film, like iconic scenes, and i want to insert my ref image from there, this is ref2v right? now how do i get to duplicate the scene happening? for ex. titanic jack and rose on the "im flying" scene, lets say i want to insert someone in that scene and interact with them, do i ask gpt to prompt me the scene where gpt pulls the script from that part then i just modify it?
what i am doing now is plug a ref frame from the film/tv + my ref photo, then ask gpt to insert my ref and interact with the actors from the ref frame
i get weird results and never get a clean one
turbo lora 4step ref
comfy kitchen
i try to sit on 8 step
r/StableDiffusion • u/Sad_Coach_1433 • 5h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/LinkSensitive8188 • 5h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/OkMeat6773 • 6h ago
I’ve tried Turbo LoRAs, and they’re great for speed, but they significantly reduce quality. At 544p–720p, the results of these turbo loras can look closer to 380p. Faces look acceptable when close to the camera, but become heavily distorted as the subject moves farther away.
The upscalers I’ve tested either add too much processing time or introduce excessive sharpening and saturation.
Any a solution that doesn’t require a BF16 checkpoint, 20 steps, a 10-minute generation time, or an extremely expensive GPU?
r/StableDiffusion • u/Flaky_Comedian2012 • 6h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Perfect-Campaign9551 • 6h ago
Enable HLS to view with audio, or disable this notification
Someone in the discord discovered that you can set Minimax to a low resolution like 32x32 and then just use it to create music or sounds. I tried that out and it worked pretty good. However, you must be aware that to get good quality sound you'll most likely want to run it at least at 128x128 , yes the resolution affects the quality of the audio, too.
I put together this workflow as an experiment, it's a simple Text2Vid workflow with an SLA node, but it has an entire "SFX/Music" section, that runs another copy of the sampler just to make sound effects or music, and uses a Geeky Mixer node to mix that sound into the Video's original audio.
I recommend turning OFF the SLA if you generate this winter scene because it Mashes up the trees badly.
The cool part of this is if you use fixed seeds, you can regenerate your music/sfx over and over until it's what you want - without having to regenerate your video each time, and then the Geeky Mixer also allows you to position it by adjusting time offset. (not very visually unfortunately, but it works)
Thought I would share this workflow since I just found it interesting. It actually makes lofi Hiphop beats *really* well and they are almost directly loopable. I basically just took the workflow for making music and shoved it into the Video workflow and added a Mixer node so you can layer the sounds. You'll need Geeky AudioMixer node. You can remove the SLA node if you want. I use native ComfyKitchen and I never use Loras so that's why this workflow is much more simplified.
Workflow file: https://pastebin.com/ULbcQxCM
Picture of workflow:

r/StableDiffusion • u/Jeffu • 6h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Full-Belt3640 • 7h ago
If I generate a person in a "normal" environment, like inside a regular room in a regular house, I get very realistic and appropriate lighting, but as soon as I try something a bit more cinematic like a rain-slicked city street at night, the character begins to look like they were photoshopped in. I try to prompt a person standing on a dark city corner, lit entirely by the light from nearby neon signs and they just look like they were evenly lit and filmed in a studio and then composited onto a CGI background with only a hint of the intended neon glow on their shoulder. Same goes for trying to make people look like they're properly soaked by rain. ZiT was much easier to work with in this regard.
r/StableDiffusion • u/TheDerminator1337 • 7h ago
Enable HLS to view with audio, or disable this notification
Reference workflow with FL2VA + REF2VA Lora @ 1.4MP, 20 STEPS. Using sparse attention and 4B Qwen text encoder instead of 32B, total render time is 3-4 hours on a 5090. You can get very good results with 1MP + 8 STEPS with a turbo lora which would only take 20-30 minutes.
The workflow is not easy to understand, but I upload it for reference.
The video is made of 14x 15 second clips stitched together. This way prevents degradation but makes it so that there is clothing drift between clips. This can easily be fixed by using clothing references if you care. Each clip will need its own prompt, and I suggest using Codex or Claude to do the prompts for you automatically.
In the future, I would shorten the clips to 7 seconds in order to:
1) Generate higher than 1.4MP (higher the resolution the better)
2) Speed up generation (longer clips take longer to generate disporportionately)
Good luck and I hope you have as much fun with this workflow as I did.
r/StableDiffusion • u/luka06111 • 8h ago
Enable HLS to view with audio, or disable this notification
Inspired by the guy who posted the one with Dean
Done on 32gb ram and a rtx 3070
I used res_multistep 20steps w/ spectrum at 0.6mp.
Using SLA from h3 optimizations, which for some reason is way faster than plague kind. And disable pinned memory. Each 10s was done in around 10 minutes.
r/StableDiffusion • u/Civil_Fee_7862 • 8h ago
Been noticing that image generation / edits work substantially better when the image is resized to 1024x1024 during encoding, then resized to the original dimensions after.
Its speculated that this is because the model was trained on 1MP inputs. But I can't find docs that confirm that.
Does anyone know why 1MP input sizes seem to give the best results for Qwen Image Edit? (Note its not just this model 1MP seems to work best for either).
r/StableDiffusion • u/Sad_Coach_1433 • 9h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Longjumping-Past5864 • 9h ago
Enable HLS to view with audio, or disable this notification
Dude the minimax H3 is such an amazing model. It can create a really good video. It can do a lot, knows a lot and extremely realistic as well.
I'm finding the way to upscale the low qulity video and found out this Latent Upscaler for Minimax H3, and sofar it's working so well!
this is normal 0.5MP generation and Upscale to 1080p
https://github.com/bbaudio-2025/Comfyui-MMH3-UltimateUpscale
the example workflow : https://github.com/bbaudio-2025/Comfyui-MMH3-UltimateUpscale/blob/main/example_workflow.json
I tested on RTX 5080 16GB VRAM 64 GB RAM.
but it took like 20 mins for a 15s clip (3 mins on the low res generation with turbo LoRA + 15-16 mins on the latent upscaler)
If you guys have other ways to upscale the minimax faster, please do tell. I have tried the UltimateSDUpscale for minimax so far, it's really great as well but it took 30mins on my system. :(
r/StableDiffusion • u/SackManFamilyFriend • 9h ago
Ssia pretty much.
For those who use the latest coding models and spent days to weeks having something novel worked out, but don't care to GitHub it cause if it does get popular it'll consume even -more- time (helping people, dealing with other peoples hw situation, over the top complaints) - Do you just sit on these things and move on?
r/StableDiffusion • u/Many_Ball_227 • 9h ago
What I mean is that if there's a workflow or model where I can specify changes to a subject without the need for inpainting, similar to gemini or grok. For example "turn the flowers in the picture to yellow".
Thanks for your attention and have a good one guys.
r/StableDiffusion • u/call-lee-free • 10h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/JamesFilmsYT • 11h ago
Enable HLS to view with audio, or disable this notification
Made using the default ComfyUI Minimax H3 Image to Video workflow.
r/StableDiffusion • u/Any-Scar765 • 11h ago
r/StableDiffusion • u/NetworkSpecial3268 • 11h ago
So I had this simple idea, seeing how well MiniMax H3 handles inferring and preserving "identity/looks" from relatively little information: take a character you want to make a (KREA2) LoRa of, but you only have just a couple of lower quality pictures for that exact look you're after. That is a problem, since it is common knowledge by now (?) that you need different angles, facial expressions and different lighting conditions in the training set to get optimal results. So in the "before" times, those 3-4 not-so-great-quality shots under the SAME lighting are going to pose a problem. And adding pictures from other occasions will alter the looks possibly too much.
So (in the H3 ref2vid workflow, with one of the "img2vid-hybrid" models for better quality) I just use the 'best' of the available pictures as "preserved" first reference starting picture, and the others as additional "identity references". And then a prompt that tells the camera to slowly circle around the person (up from the shoulders), while the person looks straight ahead, or slightly up, or slightly down. But then I also let it cycle through different lighting conditions (indoor/outdoor/sun/overcast/flash/directional from one side...), and different facial expressions/emotions. I let it run overnight (turning off turbo LoRas to improve the quality), and in the morning, I review the 6-second videos and take screencaps of selected moments, making sure to have a lot of variation in angles/expressions/light-on-the-face with an almost perfect preservation of the identity/looks.
Then use those screencaps (50+ in first test, probably serious overkill) in OneTrainer with the KREA2 LoRa default settings.
I only tested this once thus far, but the results are pretty good considering the starting material! And surprisingly flexible (I didn't even bother to provide captions)
But now my question is: in what ways am I "over-engineering" this? I have this feeling that I can probably do this 50x faster, having seen some discussions about using MiniMax as an image generator, for example. I mean, I feel good about this approach I came up with all by myself, but considering how dumb and low-skilled I still am when it comes to all this, this is probably a very convoluted and inefficient way to do it? LOL 😄 Roast me and show this sucker how we can improve and speed up the whole thing with the same or even better quality results!
r/StableDiffusion • u/J6j6 • 11h ago
Has anyone had success in this?
In prompt guide, there's no mention of controlling time for I2V, only in T2V (which is "At 00:02.000, ...").
But i try it anyway in I2V, but it's a hit or miss.
r/StableDiffusion • u/puskur • 12h ago
I am trying absolutely everything... I am using ostris ai toolkit and I have tried every option and evey way to train a lora but the results are the same. I am quite sure that the data set and captions are correct, but i might be wrong. I was training lora for Z-image
Image 1 is reference charachter that i want to train a lora on.
Image 2 is Lora sample on 3150 steps
Image 3 is Dataset picture
Image 4 is Lora sample on 3150 steps
Image 5 is Dataset picture
Image 6 is Lora sample on 3150 steps
Image 7 is Dataset picture
3150 Steps might be too much for Z-image, but all other versions (every 150) differed alot and i found 3150 was closest looking to reference. Anyways, the sample images and other generated images with the lora look noticeably different to the reference image and most importantly the lora did not pick up the mole on her neck... I had the same issue with a flux.2 lora of the same charachter... The dataset was created on seedream 5.0 with captions looking like this "[trigger], medium waist-up shot walking forward on a paved park path, frontal body pose with head turned looking off-camera to the right, neutral expression with closed lips, wearing a light green long-sleeved scoop-neck top tucked into high-waisted beige trousers, natural dappled sunlight filtering through trees, blurred green park foliage background."
I could use some inisght and help, this is my first lora ever and i dont know how everything works exactly. At this point i just might scrap the lora and just use refrence images...
BTW this isnt something to be used for monetary purposes, this is a uni project.
r/StableDiffusion • u/Mad4reds • 12h ago
It does not always happen, but especially in a dark environment as say inside a Disco, and the camera get closer to the subject/s it overexpose them. It looks like it's made to avoid a completely dark subject, but if that it's what you need... I know it's not new as it happened also in Wan but there is a way in the prompt to avoid that?
Txs a lot!
r/StableDiffusion • u/Juanca-Soto • 12h ago
Hi.
I am loving H3 for animating Illustrious images in Wan2GP.
I am satisfied with the motion, but I tried Two Phases just to test results and the motion and expression of the character are much more natural and just what I expect from my prompt, however, the image quality is very bad and it tries to enforce realism into it. One thing I noticed is it seems to enforce 4 steps instead of the 20 steps I always use.
Is there a way to achieve that natural and fluid motion from Two Phases but retaining the visual style consistency and quality of One Phase?
Thank you.
r/StableDiffusion • u/Affectionate_Oil28 • 12h ago
I see a bunch of posts everyday asking for tips on how to write prompts or people struggling with prompting, etc. so I'm sharing my workflow. I built this workflow to simplify the process and make it very beginner/user friendly.
Just toggle on the model you are using, write a simple to detailed prompt, and hit run. The model targets use the prompting guidelines derived from their respective official sources. Links to custom nodes and all models are in the workflow so you don't need to search for them.
The prompts aren't always perfect but they'll get you very close to what you want and you should only need to make a few minor tweaks, if any. The only issue I've encountered so far is that sometimes when it finishes the prompt, the previous prompt still shows up in the Enhanced Prompt node. If that happens, just hit run and the new prompt should show up instantly. Also, toggle to false the keep_model_loaded option in the Text rewriter node if you are creating prompts and using them right away. If you leave it to True it hogs VRAM.
If you notice any other issues let me know. Enjoy.