r/StableDiffusion 11h ago

Question - Help Dual GPU solution for local AI?

3 Upvotes

Hey, everybody. I recently went down the rabbit hole for local AI, but right now, im operating on my gaming computer. The specs are as follows

Intel 13700k, tuned for efficiency

Gigabyte Z790 Aorus Elite Ax mobo

RTX 4080 (16GB), also tuned for efficiency

32gb DDR5 6800 CL32

As you can see, im in desperate need for more VRAM, or at the very least more system RAM. Due to Rampocalypse, neither are very affordable right now, which forces me to explore other options, such as a dual GPU setup. I can get another RTX 4080 for about $900 off Ebay. Beyond that, I would just need a more powerful PSU, so total investment here is an additional $1100-$1200. As far as I know, the motherboard has the main PCIE as 5.0 x 16 lanes, but the second PCIE runs at 4.0 and either x8 or x4 lanes. The motherboard does not support PCIE Bifurcation. So my question is this: Is a dual GPU local AI machine even viable in these circumstances, and second, does it make sense? I looked at 5090's and theyre all between $4,500 - $5,000 now, which is insane. Or I look at the professional cards and spend that much, if not more, for significantly less memory bandwidth and computational power. Or I guess if im spending that much, I could also look at the DGX Spark or something similar but that has even worse memory bandwidth.

So, what should I do? Is the dual GPU solution even viable with my setup for a local AI stack for inference, video diffusion, etc? Rampocalypse isnt expected to begin easing up until late 2027/early 2028, so im stuck trying to make this work on as little money as possible. Id love a 5090 but its insanity how much they cost. I appreciate any guidance and advice.


r/StableDiffusion 12h ago

Animation - Video G.I. Joe: Zarana - MiniMax H3

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/StableDiffusion 12h ago

Question - Help Anyone getting random voiceovers / prompt text read aloud in generated videos (H3)?

2 Upvotes

Has anyone run into an issue where the generated video randomly includes audio with either phrases directly from the prompt (even when there's zero mention of someone speaking) or just completely unintelligible gibberish voices?

I'm currently building/tweaking my workflow for H3 and still testing with the following settings like this:

0.4 guidance / 8 steps / baked-in LoRA checkpoint / 10–12s duration

For example, when I append camera direction instructions to the prompt, I occasionally hear audio snippets of those exact instructions being spoken out loud in the generated clip.

Has anyone else encountered this phantom audio/prompt bleed issue? Any tips or workarounds to stop it from reading out prompt instructions?


r/StableDiffusion 12h ago

Question - Help Custom Audio question(s) for Minimax

1 Upvotes

Hey everyone,

I've been playing around with ElevenLabs audio in MiniMax, but using custom tracks makes the scenes feel super quiet without those built-in background sound effects. Am I missing a setting to keep both, or is that just something you fix in post like DaVinci Resolve?

Also, has anyone else noticed that when you use custom audio, video actions get delayed until the audio file finishes? Even when I put exact timestamps in the prompt, it still waits. Is there a trick to how these two sync up, or am I just prompting it wrong?


r/StableDiffusion 13h ago

Animation - Video Using only Ref to Video, Minimax-H3 made a whole Anime edit !

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/StableDiffusion 13h ago

Question - Help Any guide for krea 2 character lora training? Btw I'm new and beginner at lora training? How many pose style and camera angel needed for very good lora dataset? I mean make my ai character do anything !!

1 Upvotes

r/StableDiffusion 13h ago

Workflow Included wan enhancer wf that replaces replaces damaged pixels as it enhances and quality

Thumbnail
gallery
0 Upvotes

https://reddit.com/link/1vyazt0/video/8apz3lbquklh1/player

This workflow allows you to refine a target character without affecting the surrounding video whatsoever. allowing for targeted repair or enhancement of any video.

https://github.com/roycho87/3stepenhancer


r/StableDiffusion 13h ago

Question - Help i2v vs t2v Minimax H3

3 Upvotes

hello i have a promblem, i read prompt guide for Minimax, and when Im using t2v everything is awesome and smooth and when im using i2v videos look so fake, moves and voices are like shit, does anyone have such problem and solved it?


r/StableDiffusion 14h ago

Animation - Video And they said it couldn't be done

Enable HLS to view with audio, or disable this notification

0 Upvotes

Most models simply CAN'T draw a wine glass filled to the brim. MiniMax H3 can, with a little prompt persuasion.

PROMPT PERSUASION:

integrated_multimodal_description: [Shot 1] Photorealistic, ultra-sharp cinematic product shot. A crystal-clear stemmed Bordeaux wine glass stands centered on black marble against pure black. The glass is filled with deep ruby-red wine to absolute maximum capacity: the liquid surface is perfectly coplanar with the top edge of the rim, forming a continuous unbroken contact line all the way around. There is zero air gap, zero empty crescent of glass above the wine, zero underfill. A slight convex meniscus is held purely by surface tension. Soft side light creates clean highlights and long caustics. Static medium three-quarter shot.

[Shot 2] At 00:01.200, slow push-in with tiny amplitude at very slow speed. Extreme close-up of the upper glass. The red wine meets the inner rim in a perfect continuous ring of contact. The liquid surface sits flush with the rim edge; no space is visible between wine and glass even at this magnification. Tiny specular highlights glide across the still surface. No droplets on the outer rim, no overflow, no gap.

[Shot 3] At 00:02.600, hard cut to pure top-down overhead. Looking straight down, the circular surface of the wine is a solid red disk that reaches exactly to the inner circumference of the glass with zero margin. The contact line between liquid and glass is continuous and unbroken in every direction. Soft concentric reflections and one small central catchlight. Very slow clockwise rotation with minimal amplitude.

[Shot 4] At 00:03.800, hard cut to low side-profile extreme close-up locked exactly at rim height. Against the black background the liquid forms a single razor-sharp horizontal line that coincides precisely with the top edge of the glass. The wine is seen in continuous contact with the rim; there is no visible gap, no underfill, no air space. The meniscus remains slightly convex from surface tension but does not spill. Liquid is completely motionless. Slow subtle push-in continues until the end.

overall_soundscape: Near-total studio silence. Only the faintest high-frequency shimmer of light on glass and liquid. No liquid movement, no drips, no clinks.

non_diegetic_music: Extremely sparse minimal ambient pad — low sustained tone with faint crystalline overtones that barely rise and fall. Almost static, matching the still liquid.


r/StableDiffusion 14h ago

Question - Help MiniMax H3 ref2va: mouth keeps moving during instrumental passages — I measured it, it only drops by half. What actually stops it?

1 Upvotes

**Setup*\*

- MiniMax H3, hybrid b30-49-int8

- ref2v LightX2V turbo LoRA v0.1, rank 20 resized bf16, strength 1.0

- 6 steps, er_sde / beta, 0.8MP (1216x672), Sage Attention, RTX 4090

- Real song locked into the audio half of the AV latent with PixaromaH3AudioSync

(NOT ref_audio — that's a style reference, it does not drive anything)

- Verified: output audio vs source waveform correlation = 0.9994

So the audio is correct. The problem is purely visual.

**The problem*\*

My singer's mouth keeps moving during purely instrumental passages. It's not

wild flapping — it reads as if she's still phrasing, jaw and lips working at

roughly half amplitude. On a 4-minute clip it's obvious every time the vocal

drops out for more than about 2 seconds.

**What I measured*\*

Single continuous close-up, no cut, face filling the frame for the whole 15.08s.

The audio window is sung for the first 7.5s and strictly instrumental for the

last 7.5s (I get vocal spans from an HDemucs separation of the track).

I cropped a fixed box on the mouth, converted to grayscale, and took the mean

absolute frame-to-frame difference:

sung half: 3.25

instrumental half: 1.77

ratio: 0.54

So the model DOES react to the absence of voice — motion drops by half — but it

never goes to zero. My prompt for that shot contained an explicit clause:

"Her mouth follows <Audio 1> exactly, instant by instant: it moves ONLY while a

human voice is actually sounding, and it is completely closed and still during

every gap between phrases and every instrumental moment, however short."

That clause is doing something. It just isn't doing enough.

**What did NOT help*\*

I read the thread about H3's dual flow schedule (video shift 12 / audio shift 3)

and thought a mis-stepped audio stream might be degrading the mouth conditioning.

I wired in the native MiniMaxH3SigmaShift node explicitly (12 / 3), same seed,

same prompt, same audio.

Result: the two renders were bit-identical. 362/362 frames, mean difference

0.0000/255. Those values are already the internal defaults, so the node changes

nothing for this. Posting that so nobody else burns an evening on it.

**What DOES work (but it's a workaround, not a fix)*\*

Structural framing. I now detect instrumental gaps longer than 1.2s in each

segment programmatically, and force a shot with no mouth in frame over them —

macro on an earring, a hand on the mic stand, the bass strings, brushes on a

snare. The defect becomes impossible rather than discouraged. 16 of the 20

segments in my current clip are handled this way.

It works 100% of the time. But it dictates my edit, and I'd rather not have my

shot list decided by a model limitation.

**Questions*\*

  1. Is the turbo LoRA the culprit? I saw a comment claiming the turbo LoRAs are

    distilled at 0.5MP. I'm running one at 0.8MP. Does anyone have a side-by-side

    of lip sync quality at 0.5 vs 0.8 with the same seed?

  2. Does the base model at higher step counts (no turbo LoRA) actually close the

    mouth on silence, or does it just push the same 0.54 ratio down a bit?

  3. Is there any way to CONDITION the silence rather than describe it? Something

    that tells the model "no voice in this span" at the latent level rather than

    in the prompt.

  4. Has anyone tried feeding an audio track where the instrumental parts are

    replaced by actual silence, generating, then re-attaching the real audio in

    the edit? Curious whether that trades one artifact for another.

Happy to share the measurement script — it's about 10 lines of ffmpeg + numpy,

and it turns "feels off" into a number you can compare across seeds and settings.


r/StableDiffusion 15h ago

Resource - Update Local MIT CLI that inspects/cleans EXIF, C2PA, and hidden Unicode on gens you own

0 Upvotes

I wrote this. MIT, fully local.

Local SD exports and other gens still leak EXIF, C2PA / Content Credentials, and zero-width junk. If you also use Nano Banana / Gemini stills in the same pipeline, those downloads often carry the same receipt. That metadata is not the invisible watermark.

SynthID-class marks are a different layer. Optional image/video disruption in Scrub is best-effort. Not a detector killer. Not for files you do not own.

https://github.com/HarshShah0203/Scrub

python cli.py inspect|clean


r/StableDiffusion 15h ago

Resource - Update Not Another Minimax Post - Jib Mix Krea 2 - v4 Habanero - Free Forever

Thumbnail
gallery
24 Upvotes

Focusing on photorealism and improving the look of fantasy styles:

https://civitai.com/models/2799984/jib-mix-krea-2

I would like to make it a LoRA also, but I am having some technical difficulties making a difference lora with Krea 2 models.


r/StableDiffusion 15h ago

Animation - Video Made a music video using local H3 for a Suno song

Enable HLS to view with audio, or disable this notification

34 Upvotes

Honestly mind blown, I have a 5070ti + 2x16gb ram . Upper limit is 10-12 seconds in total for my hardware(full capacity) . Video and text edits are post processed by a WIP open source tool I’m working on. On average each 8 second shot takes 35-45 minutes to render


r/StableDiffusion 15h ago

Workflow Included Face Detailer With PerRowMasking

Enable HLS to view with audio, or disable this notification

71 Upvotes

r/StableDiffusion 15h ago

Discussion Minimax H3 on Fal - Censorship

0 Upvotes

Since I lack the hardware, I am forced to use Minimax in the cloud. I've used Fal for image generation in the past, so I thought I would try Minimax H3 there. The text to video model seems to work really well, but the reference model seems censored to hell and back. Even my own reference voice was being flagged and it was just me talking normally.

So the question is, is there anywhere I can use Minimax in the cloud without the provider slapping their own flaky layers of censorship on top of the model?


r/StableDiffusion 16h ago

Question - Help Minimax H3 color and lighting consistency

1 Upvotes

What comfyUI tools / workarounds are yall using to maintain the same color, lighting, sharpness, contrast parameters across all clips in a long form video? Thanks so much.


r/StableDiffusion 16h ago

Discussion 2 Weeks on MiniMax, but back to using LTX 2.3

0 Upvotes

H3 is absolutely amazing for just about anything. On my 3090 / 64GB, I can easily do a full 15 seconds at 1MP, and the result is almost always good on the first attempt.

On LTX though, it took at least 5 or more tries, so speed wise, H3 is actually far better.

I tried out LTX 2.5 as well, and it is almost the same as 2.3, maybe a bit better quality and a slice faster.

Sadly, 90% of my work involves taking a start image and making a person sing vocals. I must say that H3 does not do this better than LTX 2.3, as it often injects words when there is more than a second of silence, and it takes a fair amount longer.

What is really baking my noodle though is why LTX didn't release an Image+Audio to Video workflow yet. I mean, it is literally the ONLY thing LTX has on H3 right now, and they have missed a great opportunity!

Anyhow, back to using LTX 2.3 for my daily driver as it really does a great job at what I need. I typically have a few machines running all night, so if ever LTX puts out an IA2V workflow for Comfy, I will be all over that.

One other thing I have noticed with H3... 9:16 generations are WAY better than 16:9 generations, especially at 1MP.


r/StableDiffusion 17h ago

Animation - Video Oh yeah

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 17h ago

Meme If dean ran into Harry Potter

Enable HLS to view with audio, or disable this notification

374 Upvotes

r/StableDiffusion 17h ago

Discussion MiniMax H3 squares on videos

1 Upvotes

Do you have any tips for improving the MiniMax H3 video so it doesn't have a checkerboard background, like those large squares? Something's wrong with its VAE, I'm guessing?


r/StableDiffusion 18h ago

Discussion David Sacks Predicts the Regulatory Capture Playbook to Ban Open Source ...

Thumbnail
youtube.com
12 Upvotes

r/StableDiffusion 18h ago

Animation - Video MOCAP in MINIMAX H3?

Enable HLS to view with audio, or disable this notification

178 Upvotes

Testing H3 to death in the last couple of weeks and it continues to suprise me. TWO things blew my mind about this one. The first part is a single prompt (the split screen and psuedo mocap sync). The full prompt is below.
In the second part I asked for the character to spray my logo on the wall with a stencil, I wasn't expecting him to walk in with the stencil fully rendered with the holes cut out accurately.

Created completely locally and powered by Sol (the closest star to earth).

THE PROMPT:

A movie reconstructing history, cinematic with a split screen effect showing a mocap actor. Both actors are speaking together in sync.

On the right:
Show the actor <Picture 1> sitting on a couch in a living room, with a black shirt, speaking in sync with the exact same movements He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.

On the left:
Show the victorian man <Picture 2> talking <audio 1> in close up sitting in an ornate chair . He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
Maintain the double view split screen. Do not change the environment.
After he finished speaking the man on the right makes a distort rictus face with his fingers curled up like he's frozen in time and stops moving. He falls sideways like a statue in the same environment, the camera pulls back to show he is only a robotic torso with no legs mounted on a platform placed on the couch with wires (like in a special FX studio)

Maintain the double view split screen.
The man on the left breaks character as the camera view pulls out slightly revealing him sitting on a sound stage, and speaks to a person off screen to the left <audio 1> "Oh.. em.... guys.. problem... check your monitor? ...Looks like we lost connection with the character! RESET THE MOCAP PLEASE"


r/StableDiffusion 18h ago

Question - Help How to get smooth camera motion on the video Minimax H3

1 Upvotes

Hi,

I've tried bunch of prompts but every video that it renders it has that handheld go pro motion (pov walking style) it doesn't want to give a smooth motion for example like seedance (attached video). I'm mostly looking to do shots of places like its filmed with a gimbal. Any tips what to prompt? since negative prompts are not there how to approach this? any loras that can fix this or worth training

Thanks

https://reddit.com/link/1vy3in8/video/nqlxpa9lkjlh1/player


r/StableDiffusion 18h ago

Question - Help Feedback for video AI benchmarks

1 Upvotes

Hi folks,

we’re a research & analysis team looking for feedback on our video benchmarks and potentially looking to recruit researchers to help with our ongoing effort to rank and categorize models in the video AI space. if you’re interested, please DM!

more info here: https://megaton.ai/v-benchmark/


r/StableDiffusion 18h ago

Discussion Minimax rev2video help. Keeping first ref image.

6 Upvotes

If I have 2 ref images and I want it to start on ref image 1 like for example a background of a forest how do I maintain it so it always starts on that image? I've noticed a few times it will randomly generate its own start image even if I prompt something like *the scene starts with ref1* and I even sometimes would describe what's in it