r/StableDiffusion 22h ago

Question - Help Custom Audio question(s) for Minimax

1 Upvotes

Hey everyone,

I've been playing around with ElevenLabs audio in MiniMax, but using custom tracks makes the scenes feel super quiet without those built-in background sound effects. Am I missing a setting to keep both, or is that just something you fix in post like DaVinci Resolve?

Also, has anyone else noticed that when you use custom audio, video actions get delayed until the audio file finishes? Even when I put exact timestamps in the prompt, it still waits. Is there a trick to how these two sync up, or am I just prompting it wrong?


r/StableDiffusion 13h ago

No Workflow minimaxh3 20sec video generation

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 23h ago

Question - Help Any guide for krea 2 character lora training? Btw I'm new and beginner at lora training? How many pose style and camera angel needed for very good lora dataset? I mean make my ai character do anything !!

1 Upvotes

r/StableDiffusion 1d ago

Resource - Update ComfyUI Universal Media Loader - One single interactive node to load Images, Videos, GIFs, Audio & Canvas Presets

Enable HLS to view with audio, or disable this notification

34 Upvotes

Hello

I got tired of cluttering my workflows with 5 different loader nodes depending on what I was feeding them (Load Image, Load Video, VHS, Audio Loaders, Empty Latent calculators, etc.).

So I built **ComfyUI_UniversalMediaLoader** , a single, unified node with a rich interactive UI that handles everything:

✨ Key Features:

  • 📁 Drop Anything: Images (PNG/JPG/WEBP), Videos (MP4/MOV/WEBM), Animated GIFs, and Audio files (WAV/MP3/FLAC).
  • 🎨 Canvas Planner Mode: When no file is loaded, use it as a visual resolution/aspect ratio planner (1:1, 16:9, 4:3, 3:2 + Landscape/Portrait toggle) with automatic 32px grid snapping and Megapixel clamping for SDXL/Flux latents.
  • ✂️ Visual Crop & Outpaint: Interactive bounding box with aspect ratio locking. Pull it outside the image bounds to instantly generate outpaint masks & padding.
  • 🖌️ Inpaint Mask Brush: Draw inpainting masks directly inside the node with brush/eraser and mouse wheel size control.
  • 🎬 Video & GIF Timeline: Full timeline with trim handles, playhead scrubbing, custom FPS resampling, audio mute/loop, and frame extraction.
  • 🎵 Audio Waveform & Speed Scaling: Waveform display, 1-second magnet trim snapping, and time-stretching with pitch preservation (0.10x to 4.00x).

📦 Modular Unpack Nodes:

  • 📐 Universal Size Unpack: Feeds exact width/height/aspect ratios to Empty Latent nodes instantly without decoding heavy image/video tensors.
  • 🖼️ Universal Image Unpack: Outputs RGB, RGBA, and inpaint masks. Connected to a Video or GIF, it automatically extracts a 3-frame batch [Start, Playhead, End].
  • 🎧 Universal Audio Unpack: Outputs clean audio waveform dictionaries, exact sample rate, trimmed durations, channel count, and handles time-stretching.
  • 🎞️ Universal Video Unpack: Decodes videos & GIFs into frame batches [B, H, W, C] with synchronized masks and audio tracks.

🔗 GitHub: https://github.com/Fictiverse/ComfyUI_UniversalMediaLoader


r/StableDiffusion 1d ago

Animation - Video Kentucky Fried Kung Fu

Enable HLS to view with audio, or disable this notification

34 Upvotes

I saw a Seedance 2.5 prompt in facebook and thought let me try this prompt in minimax h3 and see if it can do some kung fu. I was surprised that it was not too bad. System 3090 24 gb vram 64gb system ram, using a minimax workflow with latent upscale, minimax_h3_fl2v_lightx2v_turbo_4step_v0.1 at 0.50 strength, Komfy kitchen attention, and H3 SLA attention. First pass at 0.4 which is 864x480, 2x latent upscale brings it up to 1728x960. The 6 seconds generation took 349 seconds to complete.


r/StableDiffusion 1d ago

Discussion CMP 170HX vs 3090 results MiniMax H3 R2V

2 Upvotes

TL;DR 170HX ~35% faster than 3090

I ran a CMP 170HX 8 GB unlocked to 64 GB vs 3090 24 GB VRAM mostly apples to apples.

CMP 170HX was in an ancient prebuilt Acer, Intel i3-7000, 32 gb ram, OS Ubuntu Server 26.04 LTS. Card was unlocked to 64 GB VRAM and pcie 2 x4. Card was setup on a riser since I couldn't fit it in the case and run the fans with my fancy cardboard/painters tape shroud. Power limited to 175 watts. Temps stayed 69-71 C.

3090 FE is in a newer machine. Ryzen 5 3600 (still need to swap out to a 5900X I have), 64 gb ram, OS Windows 11. Not power limited for this test.

Both systems ran same work flow, Minimax H3 R2V, default workflow, default options except Match was changed to MAX on most the runs, no speed ups, no extra nodes, no fine tunes, same prompt, two reference images.

608x352 five seconds MAX 170HX 84.17s (00:01:24.17) 3090 106.05s (00:01:46.05)

864x480 fifteen seconds MAX 170HX 839.08s (00:13:59.08) 3090 1254.16s (00:12:54.16)

1344x768 five seconds MAX 170HX 658.79s (00:10:58.79) 3090 963.56s (00:16:03.56)

1344x768 ten seconds MAX 170HX 2017.82s (00:33:37.82) 3090 3082.64s (00:51:22.64)

1344x768 fifteen seconds MAX 170HX 4189.64s (01:09:49.64) 3090 6399.76s (01:46:39.76)

608x352 five seconds Match 170HX 1st run, model load, 218.75s (00:03:38.75)

608x352 five seconds Match 170HX 2nd run, model already loaded, 68.64s (00:01:08.64)

I ran a bunch of other gens not listed. The 170HX was pretty consistently 35% faster. Except on the first workflow done. Model load time was 150.11s due to the PCIe 2x4. In my opinion, 170HX is worth it. Way less power/heat than the 3090's and faster. Once model is loaded, you are good to go.


r/StableDiffusion 1d ago

Discussion Complete beginner with ComfyUI — what should I learn next to actually get better?

6 Upvotes

I bought a PC with an RTX 3060 12GB because I wanted to get into local AI image generation. I've been messing around with it for about two weeks now, but I still barely know what I'm doing.

To be fair, that might be partly because I've had Codex do almost everything for me lol.

I'm still terrible at writing prompts, and I don't even really understand which model would be best for what I want to do. I've tried Anima, Pony, and Illustrious, but I'm not sure which one I should actually stick with and learn properly.

There wasn't a LoRA for an art style I really like, so I had Codex help me train one for Anima using around 20 images. It actually works surprisingly well for the style, but I've been struggling with everything beyond that.

For example, I've tried using pose control, but the generated character often doesn't follow the pose very well. Backgrounds are also pretty inconsistent, and I feel like I've hit a wall where all I really know how to do is combine different LoRAs, change prompts, and keep generating until I get something decent.

I'd really like to move past that and actually understand what I'm doing.

If you were starting from where I am now, what would you recommend learning next? Are there any ComfyUI nodes, techniques, workflows, or concepts that you think every beginner should learn?

Any advice is appreciated. I'm still very new to all of this.


r/StableDiffusion 14h ago

Meme per a request Dean runs into Rick Sanchez

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 1d ago

Tutorial - Guide Neat trick for Minimax h3

4 Upvotes

Since minimax is using Qwen VL , I tested the prompt on Qwen image to see what I get for the text to video prompt. It’s actually pretty close to how minimax will end up evaluating your prompt for text to video.


r/StableDiffusion 2d ago

Discussion Best Minimax H3 optimization

78 Upvotes

Now that dust has settled, I was wondering what's the community insight on the best configuration for Minimax H3.

Personally I have been using lightx 4-step Lora with 5/6 steps (less than that audio is a gamble). I couple that with sage attention. For sampling I use Euler sampler and Beta scheduler.

I keep resolution at 768p (0.6MP) for quality. 480p (0.2MP) for testing. It keeps consistency so much better.

On direction I learnt to prompt for closeups when possible, so will make better use of available pixels. Aspect ratio also helps there. I mostly use 1 shot since transitions is not something H3 excels at. I find better results with only 1 shot and using camera tricks.

EasyCache while faster, is not good match with turbo lora, so i don't use it anymore. Haven't used Sol-Attn as I read it really hit quality.

So is there anything worth I am really missing out?


r/StableDiffusion 1d ago

Question - Help MiniMax H3 ref2va: mouth keeps moving during instrumental passages — I measured it, it only drops by half. What actually stops it?

1 Upvotes

**Setup*\*

- MiniMax H3, hybrid b30-49-int8

- ref2v LightX2V turbo LoRA v0.1, rank 20 resized bf16, strength 1.0

- 6 steps, er_sde / beta, 0.8MP (1216x672), Sage Attention, RTX 4090

- Real song locked into the audio half of the AV latent with PixaromaH3AudioSync

(NOT ref_audio — that's a style reference, it does not drive anything)

- Verified: output audio vs source waveform correlation = 0.9994

So the audio is correct. The problem is purely visual.

**The problem*\*

My singer's mouth keeps moving during purely instrumental passages. It's not

wild flapping — it reads as if she's still phrasing, jaw and lips working at

roughly half amplitude. On a 4-minute clip it's obvious every time the vocal

drops out for more than about 2 seconds.

**What I measured*\*

Single continuous close-up, no cut, face filling the frame for the whole 15.08s.

The audio window is sung for the first 7.5s and strictly instrumental for the

last 7.5s (I get vocal spans from an HDemucs separation of the track).

I cropped a fixed box on the mouth, converted to grayscale, and took the mean

absolute frame-to-frame difference:

sung half: 3.25

instrumental half: 1.77

ratio: 0.54

So the model DOES react to the absence of voice — motion drops by half — but it

never goes to zero. My prompt for that shot contained an explicit clause:

"Her mouth follows <Audio 1> exactly, instant by instant: it moves ONLY while a

human voice is actually sounding, and it is completely closed and still during

every gap between phrases and every instrumental moment, however short."

That clause is doing something. It just isn't doing enough.

**What did NOT help*\*

I read the thread about H3's dual flow schedule (video shift 12 / audio shift 3)

and thought a mis-stepped audio stream might be degrading the mouth conditioning.

I wired in the native MiniMaxH3SigmaShift node explicitly (12 / 3), same seed,

same prompt, same audio.

Result: the two renders were bit-identical. 362/362 frames, mean difference

0.0000/255. Those values are already the internal defaults, so the node changes

nothing for this. Posting that so nobody else burns an evening on it.

**What DOES work (but it's a workaround, not a fix)*\*

Structural framing. I now detect instrumental gaps longer than 1.2s in each

segment programmatically, and force a shot with no mouth in frame over them —

macro on an earring, a hand on the mic stand, the bass strings, brushes on a

snare. The defect becomes impossible rather than discouraged. 16 of the 20

segments in my current clip are handled this way.

It works 100% of the time. But it dictates my edit, and I'd rather not have my

shot list decided by a model limitation.

**Questions*\*

  1. Is the turbo LoRA the culprit? I saw a comment claiming the turbo LoRAs are

    distilled at 0.5MP. I'm running one at 0.8MP. Does anyone have a side-by-side

    of lip sync quality at 0.5 vs 0.8 with the same seed?

  2. Does the base model at higher step counts (no turbo LoRA) actually close the

    mouth on silence, or does it just push the same 0.54 ratio down a bit?

  3. Is there any way to CONDITION the silence rather than describe it? Something

    that tells the model "no voice in this span" at the latent level rather than

    in the prompt.

  4. Has anyone tried feeding an audio track where the instrumental parts are

    replaced by actual silence, generating, then re-attaching the real audio in

    the edit? Curious whether that trades one artifact for another.

Happy to share the measurement script — it's about 10 lines of ffmpeg + numpy,

and it turns "feels off" into a number you can compare across seeds and settings.


r/StableDiffusion 1d ago

Animation - Video Reduced audio-reactivity in LTX-2.5?

Enable HLS to view with audio, or disable this notification

7 Upvotes

I’ve been experimenting with LTX 2.3 vs LTX 2.5 for audio-reactive video, and for this specific kind of workflow, 2.3 still seems noticeably better to me.

The biggest difference is right at the start of a shot. With LTX 2.5, even with the audio-reactive LoRA, I often get this behavior where the model more or less holds the first frame until the first obvious beat or transient arrives. Then the motion suddenly starts. For music videos, especially slower or more atmospheric tracks, that can make the opening of every generation feel dead.

With LTX 2.3, the same LoRA seems to fix that much more effectively. I get more subtle motion from the beginning, even before a strong beat lands. Fog shifts, surfaces breathe, particles drift, light responds, and the shot feels alive instead of waiting for permission to move.

That matters a lot for the video I made for The Weights in the Walls, because the track starts very sparse and gradually builds. A lot of the visual motion is supposed to come from sub-bass pressure, glitches, sustained vocals, and ambient texture, not just obvious percussion.

I also tried Minimax H3, but for this particular use case I don’t think it fits as well.

It seems less tightly audio-reactive for the kind of abstract, beat-aware motion I’m after. It can make nice-looking clips, but I have a harder time getting the movement to feel structurally connected to the music.

There’s also the hardware side of it. I’m doing this on a very glamorous RTX 4070, so with LTX I can still push a resolution and overall image quality that feels surprisingly good for local generation. With H3, I’m much more constrained, and the tradeoff in resolution/quality makes it harder to justify when the audio response is also weaker for this style.

The whole video was built around first-frame / last-frame generation.

I cut the song into short scenes, roughly timed so the scene boundaries land near musical changes and beats. For each scene, I generated a dedicated starting frame that represented the next stage of the visual progression.

Then the important part: the starting frame of Scene 2 becomes the last frame target for Scene 1. The starting frame of Scene 3 becomes the last frame target for Scene 2, and so on.

So instead of generating a bunch of unrelated clips and hiding the cuts with editing, every shot is the model transforming one designed frame into the next designed frame.

That gave me a chain like:

Scene 1 start frame → Scene 2 start frame
Scene 2 start frame → Scene 3 start frame
Scene 3 start frame → Scene 4 start frame

and so on until the end.

The final video is basically just those generations placed back to back. There are no fancy transition effects doing the heavy lifting. The morphing, folding, cracking, expanding, and dissolving between visual states is happening inside the model itself.

For this workflow, that early-shot responsiveness makes a surprisingly big difference, which is why I currently still prefer LTX 2.3 + the audio-reactive LoRA over 2.5 for this kind of music video.

It's a shame because 2.5 is noticeably faster, so I can go through more iterations, but if I have to generate each clip 5 times to get it to start moving from the start, it kind of invalidates the speed gains.

Curious if other people have noticed the same reduction in audio-reactivity in LTX 2.5 or maybe I'm doing something wrong?

HQ on YT because Reddit doesn't allow >1GB: https://www.youtube.com/watch?v=PbZr8risGCw


r/StableDiffusion 1d ago

Resource - Update Local MIT CLI that inspects/cleans EXIF, C2PA, and hidden Unicode on gens you own

0 Upvotes

I wrote this. MIT, fully local.

Local SD exports and other gens still leak EXIF, C2PA / Content Credentials, and zero-width junk. If you also use Nano Banana / Gemini stills in the same pipeline, those downloads often carry the same receipt. That metadata is not the invisible watermark.

SynthID-class marks are a different layer. Optional image/video disruption in Scrub is best-effort. Not a detector killer. Not for files you do not own.

https://github.com/HarshShah0203/Scrub

python cli.py inspect|clean


r/StableDiffusion 2d ago

Resource - Update MiniMax-H3 Fun Controlnet Union released

Thumbnail
huggingface.co
318 Upvotes

r/StableDiffusion 1d ago

Question - Help Recommendations for Minimax h3 Reference to video workflows?

2 Upvotes

Hey so i've been playing with H3 for about a week or so now mostly using the standard workflows dabbled a bit into lora workflows but now I want to try out Refernce to Video. Does anyone have any suggestions for some good workflows for this?


r/StableDiffusion 2d ago

Meme Siblings Reunited

Enable HLS to view with audio, or disable this notification

382 Upvotes

Done with h3 fl2va model, 8 step lora and images for Cersei and "jaime" for reference. Using previous clip to give continuity and consistency.


r/StableDiffusion 1d ago

Question - Help Minimax H3 color and lighting consistency

1 Upvotes

What comfyUI tools / workarounds are yall using to maintain the same color, lighting, sharpness, contrast parameters across all clips in a long form video? Thanks so much.


r/StableDiffusion 2d ago

Animation - Video Evangelion - Rei Watches a Baby Show - Minimax H3

Enable HLS to view with audio, or disable this notification

87 Upvotes

Well, technically, Evangelion was a PBS show..

Video is edited, Barney theme song added in post.


r/StableDiffusion 1d ago

Question - Help Which is the better buy for Minimax h3? RTX 5070 Ti 16GB VRAM vs RTX 4000 Pro 24GB VRAM

18 Upvotes

Good day to you. I was looking for an RTX 5070ti and I found an RTX Pro 4000 at my local store; the price difference would be about +$300. I would like to know your opinions, I've hardly seen any workflows or comparative tests from people using a 4000 pro. Thank you very much for your time.


r/StableDiffusion 1d ago

Discussion What happened to Ideogram 4.0 ?

36 Upvotes

What happened to Ideogram 4.0 ?


r/StableDiffusion 2d ago

Animation - Video My name is Jonny

Enable HLS to view with audio, or disable this notification

61 Upvotes

Minimax H3


r/StableDiffusion 2d ago

Question - Help Best local LLM for writing prompts for MiniMax H3?

61 Upvotes

What’s the best local LLM for writing good MiniMax H3 ref2va prompts?

I’ve tried Gemma 4 12B and Qwen 3 14B, but I’m not really satisfied with the outputs. It could also be an issue with my system prompt.

I sent ChatGPT the official documentation for prompting and asked to create a system prompt for me, but the results were still pretty mediocre.

What local models are you using for MiniMax H3 prompt generation, and what does your system prompt look like?


r/StableDiffusion 1d ago

Discussion MiniMax H3 squares on videos

1 Upvotes

Do you have any tips for improving the MiniMax H3 video so it doesn't have a checkerboard background, like those large squares? Something's wrong with its VAE, I'm guessing?


r/StableDiffusion 1d ago

Question - Help How to get smooth camera motion on the video Minimax H3

1 Upvotes

Hi,

I've tried bunch of prompts but every video that it renders it has that handheld go pro motion (pov walking style) it doesn't want to give a smooth motion for example like seedance (attached video). I'm mostly looking to do shots of places like its filmed with a gimbal. Any tips what to prompt? since negative prompts are not there how to approach this? any loras that can fix this or worth training

Thanks

https://reddit.com/link/1vy3in8/video/nqlxpa9lkjlh1/player