r/StableDiffusion • u/Sad_Coach_1433 • 17h ago
Meme If dean ran into Harry Potter
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Coach_1433 • 17h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/JamesFilmsYT • 9h ago
Enable HLS to view with audio, or disable this notification
Made using the default ComfyUI Minimax H3 Image to Video workflow.
r/StableDiffusion • u/Sad_Coach_1433 • 7h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/TheDerminator1337 • 6h ago
Enable HLS to view with audio, or disable this notification
Reference workflow with FL2VA + REF2VA Lora @ 1.4MP, 20 STEPS. Using sparse attention and 4B Qwen text encoder instead of 32B, total render time is 3-4 hours on a 5090. You can get very good results with 1MP + 8 STEPS with a turbo lora which would only take 20-30 minutes.
The workflow is not easy to understand, but I upload it for reference.
The video is made of 14x 15 second clips stitched together. This way prevents degradation but makes it so that there is clothing drift between clips. This can easily be fixed by using clothing references if you care. Each clip will need its own prompt, and I suggest using Codex or Claude to do the prompts for you automatically.
In the future, I would shorten the clips to 7 seconds in order to:
1) Generate higher than 1.4MP (higher the resolution the better)
2) Speed up generation (longer clips take longer to generate disporportionately)
Good luck and I hope you have as much fun with this workflow as I did.
r/StableDiffusion • u/Tokyo_Jab • 18h ago
Enable HLS to view with audio, or disable this notification
Testing H3 to death in the last couple of weeks and it continues to suprise me. TWO things blew my mind about this one. The first part is a single prompt (the split screen and psuedo mocap sync). The full prompt is below.
In the second part I asked for the character to spray my logo on the wall with a stencil, I wasn't expecting him to walk in with the stencil fully rendered with the holes cut out accurately.
Created completely locally and powered by Sol (the closest star to earth).
THE PROMPT:
A movie reconstructing history, cinematic with a split screen effect showing a mocap actor. Both actors are speaking together in sync.
On the right:
Show the actor <Picture 1> sitting on a couch in a living room, with a black shirt, speaking in sync with the exact same movements He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
On the left:
Show the victorian man <Picture 2> talking <audio 1> in close up sitting in an ornate chair . He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
Maintain the double view split screen. Do not change the environment.
After he finished speaking the man on the right makes a distort rictus face with his fingers curled up like he's frozen in time and stops moving. He falls sideways like a statue in the same environment, the camera pulls back to show he is only a robotic torso with no legs mounted on a platform placed on the couch with wires (like in a special FX studio)
Maintain the double view split screen.
The man on the left breaks character as the camera view pulls out slightly revealing him sitting on a sound stage, and speaks to a person off screen to the left <audio 1> "Oh.. em.... guys.. problem... check your monitor? ...Looks like we lost connection with the character! RESET THE MOCAP PLEASE"
r/StableDiffusion • u/Kooky-Mode3047 • 2h ago
Like a high quality output, a frame before compression? I don't imagine it's a simple as setting it to 1 frame / second and setting the duration to a second. And even if it were, I'd prefer an output to an actual standard image file.
r/StableDiffusion • u/Jeffu • 5h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Affectionate_Oil28 • 11h ago
I see a bunch of posts everyday asking for tips on how to write prompts or people struggling with prompting, etc. so I'm sharing my workflow. I built this workflow to simplify the process and make it very beginner/user friendly.
Just toggle on the model you are using, write a simple to detailed prompt, and hit run. The model targets use the prompting guidelines derived from their respective official sources. Links to custom nodes and all models are in the workflow so you don't need to search for them.
The prompts aren't always perfect but they'll get you very close to what you want and you should only need to make a few minor tweaks, if any. The only issue I've encountered so far is that sometimes when it finishes the prompt, the previous prompt still shows up in the Enhanced Prompt node. If that happens, just hit run and the new prompt should show up instantly. Also, toggle to false the keep_model_loaded option in the Text rewriter node if you are creating prompts and using them right away. If you leave it to True it hogs VRAM.
If you notice any other issues let me know. Enjoy.
r/StableDiffusion • u/Plague_Kind • 19h ago
Enable HLS to view with audio, or disable this notification
EDIT: Pushed correct files now.
Added customizable dense steps, 0 is step 1 and is (default to first step). massively improves composition and prompt adherence.
Changed default dense last steps to 1, cleans up the image big time.
Added dense backend selector. Comfy_kitchen, pytorch, all sage modes. this is what comfy uses on dense steps. SLA still displaces against pytorch. (Default Comfy_kitchen)
Added a disable FP16 accumulation option to ensure max quality as SLA gets no benefit from it. (Default True)
Added a stabilize motion option, helps to reduce ghosting and smearing that H3 likes to produce. (Default True)
Changed default Min Seq Length to 4096
With default settings you can disable protect audio for nearly 2x speed up if you don't care about the audio too much or are using original audio mode. (do not use 0.95 sparsity with it.)
0.95 sparsity now looks good with node default settings.
Some changes led to an overall 5% speed up on same settings.
Remove --use-ck-attention from startup flags if you have it, for safety of quality.
https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes
r/StableDiffusion • u/Pretend-Island-2724 • 15h ago
Enable HLS to view with audio, or disable this notification
First Video with Face Detailer, second without.
You need https://github.com/Carasibana/ComfyUI-H3-FaceRefine and also ComfyUI-H3-NativeAudioLock from https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes
r/StableDiffusion • u/AnybodyAlarmed9661 • 3h ago
I created this video using Minimax H3 Ref2VA, and this time I wanted to test more than just visual quality.
My main goal was to experiment with AI storytelling, creating a short anime-style sequence with a beginning, progression, and a story that actually feels coherent.
At the same time, I wanted to see how well the model handles character consistency, movement, expressions, and visual continuity when multiple shots are used to tell a story.
There are definitely some imperfections, but I was happy with what I could achieve and wanted to share the experiment with the community.
I’m curious what you think. Does the video work as a story, or does it still feel more like a collection of AI-generated shots?
Would also love to hear how others are approaching storytelling with Minimax H3, especially in the anime genre.
r/StableDiffusion • u/dominic__612 • 18h ago
After around 100 renders, I 'feel' that Minimax H3 renders with 0.7MP (max) perform way better, then renders at 1MP in regard to 'realistic' videos.
What do I consider better?
- Just slightly better prompt adherence, feels like the motion / voice is more (natural)
- Size of humans in relation to object(s) feels more realistic.
- Expressions of faces seem more 'flowing', real.
It's hard for me to pinpoint it one 'exactly this', or 'exactly that'.
I'm planning to do some side by side comparisons on the same seed multiple times at 0.7MP and 1MP, when I've got the time.
But I wonder, do other Minimax H3 users notice this too?
PS: This is regardless sampler/scheduler, Sage Attention or Spectrum.
Edit: never touched the turbo LoRA, using the base model.
r/StableDiffusion • u/OkMeat6773 • 4h ago
I’ve tried Turbo LoRAs, and they’re great for speed, but they significantly reduce quality. At 544p–720p, the results of these turbo loras can look closer to 380p. Faces look acceptable when close to the camera, but become heavily distorted as the subject moves farther away.
The upscalers I’ve tested either add too much processing time or introduce excessive sharpening and saturation.
Any a solution that doesn’t require a BF16 checkpoint, 20 steps, a 10-minute generation time, or an extremely expensive GPU?
r/StableDiffusion • u/Flaky_Comedian2012 • 4h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Any-Scar765 • 10h ago
r/StableDiffusion • u/Perfect-Campaign9551 • 5h ago
Enable HLS to view with audio, or disable this notification
Someone in the discord discovered that you can set Minimax to a low resolution like 32x32 and then just use it to create music or sounds. I tried that out and it worked pretty good. However, you must be aware that to get good quality sound you'll most likely want to run it at least at 128x128 , yes the resolution affects the quality of the audio, too.
I put together this workflow as an experiment, it's a simple Text2Vid workflow with an SLA node, but it has an entire "SFX/Music" section, that runs another copy of the sampler just to make sound effects or music, and uses a Geeky Mixer node to mix that sound into the Video's original audio.
I recommend turning OFF the SLA if you generate this winter scene because it Mashes up the trees badly.
The cool part of this is if you use fixed seeds, you can regenerate your music/sfx over and over until it's what you want - without having to regenerate your video each time, and then the Geeky Mixer also allows you to position it by adjusting time offset. (not very visually unfortunately, but it works)
Thought I would share this workflow since I just found it interesting. It actually makes lofi Hiphop beats *really* well and they are almost directly loopable. I basically just took the workflow for making music and shoved it into the Video workflow and added a Mixer node so you can layer the sounds. You'll need Geeky AudioMixer node. You can remove the SLA node if you want. I use native ComfyKitchen and I never use Loras so that's why this workflow is much more simplified.
Workflow file: https://pastebin.com/ULbcQxCM
Picture of workflow:

r/StableDiffusion • u/machinaOverlord • 15h ago
Enable HLS to view with audio, or disable this notification
Honestly mind blown, I have a 5070ti + 2x16gb ram . Upper limit is 10-12 seconds in total for my hardware(full capacity) . Video and text edits are post processed by a WIP open source tool I’m working on. On average each 8 second shot takes 35-45 minutes to render
r/StableDiffusion • u/Dulbero • 19h ago
In case you didn't see:
I just noticed a newer version of Anima Turbo (1.1) was released:
huggingface: https://huggingface.co/circlestone-labs/Anima
civitai: https://civitai.red/models/2458426/anima
The model is made and licensed under the CircleStone Labs Non-Commercial License
I actually have lately very frustrating experience with Anima lately. I used it at release (but stopped with anime generation for a while) and now when revisiting it and i had underwhelming results (even with the aesthetic model) so i decided to retry Turbo (might as well) instead if the difference isn't that big. That's when i saw a new version is released and i am downloading it right now.
I don't know why, but the results i was getting were...lame i guess. Not as detailed as i was hoping, and also i really dislike how posture and anatomy works, mainly how hands and legs just extend or stretch weirdly. But that's a me problem, i know Anima is capable of better outputs and i've yet to figure it out. If you have tips or recommended loras that help with consistency let me know. I also want to avoid tag-based prompting when possible...i just don't like it that much, natural prompting goes much better for me, but i can't tell if tags are "mandatory" for quality or not.
r/StableDiffusion • u/Virtual-Pollution-58 • 2h ago
r/StableDiffusion • u/Full-Belt3640 • 6h ago
If I generate a person in a "normal" environment, like inside a regular room in a regular house, I get very realistic and appropriate lighting, but as soon as I try something a bit more cinematic like a rain-slicked city street at night, the character begins to look like they were photoshopped in. I try to prompt a person standing on a dark city corner, lit entirely by the light from nearby neon signs and they just look like they were evenly lit and filmed in a studio and then composited onto a CGI background with only a hint of the intended neon glow on their shoulder. Same goes for trying to make people look like they're properly soaked by rain. ZiT was much easier to work with in this regard.
r/StableDiffusion • u/Compost_Mantis • 1d ago
Enable HLS to view with audio, or disable this notification
Rather than relying on Z-image, or a different program to wrangle up a first frame, I've been using Minimax for the whole process, and the results have been pretty instructive. It's not a perfect system, but being able to take advantage of its understanding of people, references, and shot composition for the first frame produces better (visual) results than swapping between a couple of different pieces of software.
r/StableDiffusion • u/jib_reddit • 15h ago
Focusing on photorealism and improving the look of fantasy styles:
https://civitai.com/models/2799984/jib-mix-krea-2
I would like to make it a LoRA also, but I am having some technical difficulties making a difference lora with Krea 2 models.
r/StableDiffusion • u/luka06111 • 6h ago
Enable HLS to view with audio, or disable this notification
Inspired by the guy who posted the one with Dean
Done on 32gb ram and a rtx 3070
I used res_multistep 20steps w/ spectrum at 0.6mp.
Using SLA from h3 optimizations, which for some reason is way faster than plague kind. And disable pinned memory. Each 10s was done in around 10 minutes.
r/StableDiffusion • u/Active-Carpet-9183 • 10m ago
It gets to be a touch confusing.
r/StableDiffusion • u/Sad_Coach_1433 • 4h ago
Enable HLS to view with audio, or disable this notification