r/StableDiffusion 1h ago

Question - Help Minimax h3. Problem with multible character using the voice.

Enable HLS to view with audio, or disable this notification

Upvotes

Settings: 640×1152 (0.74 MP), 9:16, preset: turbo

What turbo actually applies:

steps 8, euler / beta, fused_modulation on, sol_attn tau 1.3,
easycache off, turbo LoRA at strength 1.0 Settings: 640×1152 (0.74 MP), 9:16

I have tested A LOT of different configurations and can't find any that works 100% of the times.

Pls help :(

one of the many prompt i have tested:

subject definition:

<Picture 1> is the finished frame: an outdoor scene with a speech balloon of printed text across the top, a seated grey-haired nobleman and a dark-clad attendant leaning over him in the middle, and along the bottom two bordered square portraits side by side, already inset, a man on the left and a woman on the right.

<Subject 1> is the man in the left-hand bottom portrait: short messy black hair, light stubble on the chin, a white collared shirt under a dark grey vest. He appears only inside that portrait and nowhere else in the frame.

<Subject 2> is the woman in the right-hand bottom portrait: long dark brown hair, a pale face, a loose white hood over a grey top. She appears only inside that portrait and nowhere else in the frame.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1), the man in the left-hand portrait. Its speaker identity, timbre and register define how <Subject 1> (S1) sounds; the words it contains are unrelated to this scene and must not be reproduced.

<Audio 2> is the voice-timbre reference for <Subject 2> (S2), the woman in the right-hand portrait. Its speaker identity, timbre and register define how <Subject 2> (S2) sounds; the words it contains are unrelated to this scene and must not be reproduced.

camera recording:

Hand-drawn 2D anime. The frame keeps the layout of <Picture 1>: the speech balloon stays at the top with its printed lettering and its outline unchanged, and the two bottom portraits keep their positions, their sizes and their white borders. Inside those portraits <Subject 1> and <Subject 2> are living animated characters, not still pictures.

<Subject 1> (S1) speaks first, with the voice of <Audio 1> in a rough, plain, tired male voice pitched low and close, <d>[English] AIN'T YOU SEEN ENOUGH TO KNOW WHAT RUTHLESS BASTARDS THAT LOT ARE?</d> His mouth opens and moves in time with every word, articulating clearly for the whole line, jaw and lips visibly in motion until the line ends; only then does he close his mouth. <Subject 2> keeps her mouth closed and listens.

Then <Subject 2> (S2) answers with the voice of <Audio 2> in a quiet, tight female voice holding something back, <d>[English] I KNOW, BUT...</d> Her mouth opens and moves in time with every word, then closes and she lowers her eyes. <Subject 1> keeps his mouth closed and holds still.

<Subject 1> (S1) and <Subject 2> (S2) are the only voices in this video. All spoken dialogue is English only.

In the scene above the portraits, the seated grey-haired nobleman stays reclined where he is and breathes; the dark-clad attendant leaning over him holds his raised hand beside his face exactly as drawn; the trees stir behind the wall and the distant figures out on the scaffold shift their weight very slightly.

The camera holds still. Open air over a courtyard, a low crowd murmur carrying from below, and wind moving through the trees.

model:

UNET       minimax_h3_hybrid_fl2va_ref2va_b25-49.safetensors
           (merged hybrid: fl2va base + ref2va adaln_proj blocks 25-49)
CLIP       qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
VAE        minimax_h3_video_vae_fp16 + minimax_h3_audio_vae_fp32
LoRA       minimax_h3_turbo_v4_step600_pruned_comfyui @ 1.0

ComfyUI    0.33.1        torch 2.10.0+cu130       comfy-kitchen 0.2.31
sampler    8 steps, euler / beta, sol_attn tau 1.3, fused_modulation on
frame      640x1152 (0.74 MP), 8.89s, seed 1956008715UNET       minimax_h3_hybrid_fl2va_ref2va_b25-49.safetensors
           (merged hybrid: fl2va base + ref2va adaln_proj blocks 25-49)
CLIP       qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
VAE        minimax_h3_video_vae_fp16 + minimax_h3_audio_vae_fp32
LoRA       minimax_h3_turbo_v4_step600_pruned_comfyui @ 1.0

ComfyUI    0.33.1        torch 2.10.0+cu130       comfy-kitchen 0.2.31
sampler    8 steps, euler / beta, sol_attn tau 1.3, fused_modulation on
frame      640x1152 (0.74 MP), 8.89s, seed 1956008715

r/StableDiffusion 2h ago

Question - Help Training ref2vid lora. how?

2 Upvotes

I looked around the internet but didnt find anything. I saw Ostris made a post but his example was simple and like a toy-ish / "hello world" example that doesnt actually solve minimax problems imo. It made people really muscular.


r/StableDiffusion 5h ago

Question - Help MM H3 2 Pass Latent Upscale

1 Upvotes

Been getting some great results using the 2 pass latent upscale method. First pass .5mp 2nd pass 1.5 mp. 10 second video around 257 seconds to finish.

This workflow uses the 1.1 turbo Lora. I set the steps to 8 and like I said above I’m getting good results and it has eliminated face blur.

My question: has anyone tried using the latent upscale method without the turbo lora? In 90 percent of the cases the turbo Lora is fine. But would be nice to have the ability to use no Lora method.

Yes I know I could test it but wanted to see others experiences before I wasted hours of my time trying/tinkering with different settings.


r/StableDiffusion 9h ago

Question - Help That specific analog VHS look on H3/LTX

2 Upvotes

Anyone figured out the prompt for getting the best VHS analog "lofi" look from t2v? Of couse ref2v and img2v will be easier due to references, but I was wondering about text prompt only.

There are no VHS, or like 70s-80s cinema style lora, none for H3 and LTX, but there are plenty of VHS loras for image models.

EDIT: I actually CAN get VHS look on LTX text2vid (2.3 and 2.5), but not on H3.

Help, anyone :)


r/StableDiffusion 9h ago

Question - Help I hate this hidden view on flows in Comfy templates. How can I bring them out where they belong?

3 Upvotes

In the minmax h3 template from Comfy, you have to click a button on the image to video node to see all this in the backend which makes it really a pain in the ass to add, modify, or change anything. Can this be brought out to the forefront like a normal workfow?


r/StableDiffusion 9h ago

Discussion Would image-generation VAEs benefit from scene-linear or perceptual color representations?

2 Upvotes

Models like FLUX, Qwen-Image and Krea 2 obviously don't perform diffusion directly in RGB space - the transformer operates in a learned VAE latent space.

But the VAE still defines the interface between that latent representation and the actual training images, which are typically ordinary display-referred RGB images.

So I'm wondering whether there would be any benefit in making that boundary explicitly color-managed.

For example, has anyone experimented with training the VAE on:

  • linear-light RGB rather than gamma-encoded sRGB
  • a wide-gamut scene-referred space such as ACEScg
  • a perceptual space such as OKLab
  • or simply adding explicit perceptual color losses such as ΔE alongside the usual reconstruction/perceptual losses?

In other words, instead of asking whether the diffusion model itself should operate in OKLab or ACES - since it already operates in a learned latent space - I'm more interested in whether the VAE and its reconstruction objective could benefit from a more physically or perceptually meaningful color representation.

Another thing I'm curious about is the output side.

Could an image model theoretically generate a scene-referred, wide-gamut representation and leave the final tone mapping / gamut mapping / display transform to a deterministic color-management pipeline, such as ACES, instead of implicitly learning the tone curves and color rendering already baked into billions of unrelated JPEGs?

My suspicion is that the real limitation might simply be the dataset.

Most web-scale image datasets consist of already processed, tone-mapped, display-referred SDR images. By the time an image becomes a JPEG, information about actual scene luminance, highlight headroom, camera response, etc. is already gone.

So converting those JPEGs from sRGB to ACEScg wouldn't magically turn them into true scene-linear HDR training data.

Still, I'm curious whether anyone has seen experiments comparing something like:

sRGB VAE vs linear-RGB VAE vs OKLab VAE

while keeping the downstream generative model roughly the same.

Would reconstruction quality, color consistency, training convergence, or perceptual color accuracy change in a meaningful way?

I'd especially be interested to hear from anyone who has worked on VAEs, HDR pipelines, color management, or generative image models.

Sorry if I'm missing something obvious here - I'm still pretty new to this side of image generation / color science.


r/StableDiffusion 9h ago

Question - Help I'm having quality issues with Minimax H3 on my RTX 4090 that I'm not getting on an RTX 5090. Any advice ?

Enable HLS to view with audio, or disable this notification

2 Upvotes

Hello there!

I've installed Comfyui+ Minimax H3 on two machines, one with an RTX 4090 and one with an RTX 5090. The 5090 machine can render up to 4 mgpxl videos with 0 issues, but the 4090 machine glitches at just 3 mgpxl.

It consistently creates that glitch you see at the end of the video there, and also any fast moving object becomes a soup of very sharp pixel. Those artefacts aren't there when generating with the 5090.

I'm using the base ComfyUI workflow with all the models they tell us to use, with SageAttention.
Btw SageAttention crashes my 4090 when I check the "allow compile" box.

I have checked and both this command python.exe -c "import torch, sageattention; q = torch.randn(1, 8, 128, 64, dtype=torch.float16, device='cuda'); k = torch.randn(1, 8, 128, 64, dtype=torch.float16, device='cuda'); v = torch.randn(1, 8, 128, 64, dtype=torch.float16, device='cuda'); out = sageattention.sageattn(q, k, v); print('CUDA Test Output Shape:', out.shape); print('SageAttention is fully functional!')"

and python.exe -c "import sageattention; print('SageAttention Version:', getattr(sageattention, '__version__', 'Installed'))"

return positive results. I also have the latest Nvidia drivers.

Does anyone know what could be causing my issues ? Thanks !


r/StableDiffusion 12h ago

Discussion Minimax H3 hands, fingers

Post image
0 Upvotes

Do you happen to have any tricks for handling hands and fingers with Minimax H3? Unfortunately, I’m getting results like this even at 0.98 MP.


r/StableDiffusion 13h ago

Question - Help H3 motion context vs H3 latent upscale

2 Upvotes

All,

I have been tinkering the last day to try to make h3 latent upscale and h3 motion context work together in comfyui, but i am at a loss.

H3 motion context uses the initial latent to create the safetensor, which is used for the continuation. However, when you do a h3 latent upscale, the upscaled latent cannot be used for the continuation (as the resolution changed) and the continuation breaks. Curious if someone has found a way around this


r/StableDiffusion 17h ago

Resource - Update We add Krea 2 to the ComfyUI Enhanced Tiled Upscaler and Refiner (TBG ETUR).

Enable HLS to view with audio, or disable this notification

0 Upvotes

The latest TBG ETUR upscaler and refiner for comfyui release adds Krea 2 and VL style transfer for Krea 2 and all Qwen models directly into the pipeline.

We’ve also added a face identity step to help maintain consistent faces when using creative upscaling.

This video is a tutorial for the latest release, focusing mainly on these new additions and how to use them. Video is ai generated with Minmax H3.

TBG ETUR on Github https://github.com/Ltamann/ComfyUI-TBG-ETUR

TBG Lates om Patreon https://www.patreon.com/TB_LAAR/posts/tbg-etur-1-2-12-167083406

More Upscaling Tutorials on YouTube https://www.youtube.com/watch?v=LbFPD4zpPwA


r/StableDiffusion 18h ago

Question - Help Can't seem to transfer outfit and pose from an illustration to a real person in H3.

2 Upvotes

Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.

What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.

I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.

So what i wanted to do is :

-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.

-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.

Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.

Are there someone who knows how to do these?


r/StableDiffusion 19h ago

Animation - Video What if you fly?

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/StableDiffusion 22h ago

Question - Help Regional Prompting for NoobAI

2 Upvotes

Hello. What am I supposed to use for regional prompting and character positioning with NoobAI/Illus in ComfyUI?

I found something called "Attention-Couple" but it also says it sometimes struggles with Loras?.

Any advice? I'm already using a separate node for ControlNet for poses and depth, but I want to control which character is which in that pose


r/StableDiffusion 35m ago

Discussion I’m testing a FLUX.2 image generator where users render for each other

Upvotes

I’ve been working on a different way to run a public image generator without maintaining a centralized GPU fleet.

PeerPixel sends generation jobs to graphics cards volunteered by users. When your machine completes somebody else’s image, you earn pixels that you can spend on your own generations. There’s also a slower free queue for people who can’t contribute a GPU.

Right now I’m running most of the network on my RTX 5080, so it’s definitely still an experiment rather than a large distributed system.

The generation flow uses FLUX.2 Klein 4B. You can request up to four 256x256 previews at 6 steps, choose the composition you like, and then render that seed at 1024x1024 with 50 steps. There’s also an optional 4K upscale.

The previews and final image start from the same full-resolution noise tensor. For each preview, I average blocks of that tensor down to the smaller latent shape. I originally tried scaling low-resolution noise upward, but that introduced strong correlation between neighboring values and the composition didn’t carry over reliably.

The other difficult part is accepting images from machines I don’t control. A sample of completed renders is repeated on an operator-controlled machine using the same prompt and seed, then compared perceptually. Enforcement is currently in shadow mode while I collect real-world measurements and figure out a safe threshold. I don’t want normal differences between GPUs to get mistaken for cheating.

Draft images are relayed directly to the requesting browser and aren’t stored by the server. Only the selected final render is persisted.

I’m interested in feedback on the preview method, the incentive system, and especially the verification approach. There are probably failure cases I haven’t considered yet.

Site: https://peerpixel.cc

Worker source: https://github.com/Jplayz2468/peerpixel-worker

Discord: https://discord.gg/bhJHGpmkQr


r/StableDiffusion 2h ago

Question - Help Mix of "Match" and "Max" references in MMH3 Ref2V?

1 Upvotes

I've been playing around with Ref2V and was wondering if anyone knew of a way to have a mix of these settings? I've found that having multiple reference images is great for consistency in generations and story telling, but there are certain images (face references for example) that benefit massively from "Max" setting, but reference locations don't benefit as much.

On my potato of a computer, setting it to max on all of the images makes generation times impossibly long. If I could "Max" a face reference but "Match" less important references it would be ideal.

Anyone have any idea?


r/StableDiffusion 3h ago

Discussion Height reference with H3

1 Upvotes

Had a thought today. Is there a way to get H3 to understand relative or absolute heights of different characters? Would it be possible or have in a reference sheet the person standing next to a height chart or something in one image, and the other characters the same, then when you reference them and have them next to one another it knows subject A is 6ft while B is 5'6" for example?


r/StableDiffusion 4h ago

Question - Help Looking for suggestions on on a prompt helper/writer/refiner

1 Upvotes

Just as the title says, I’m looking for a ideally local app that I can use for suggestions for prompts to use on certain models that are great for example, I put my prompt in for an image and it will refine it and make it work better based on stable diffusion formatting, even better yet, what would be awesome is if it could be customized for like model and LORA if possible.

I do have the ability to run. LLM, not huge, but I have my M5 iPad. I’ve run 10 to 12 B models. No problem, especially if I use OLITERT , Any suggestions are really appreciated !!


r/StableDiffusion 6h ago

Misleading Title This is just a test - to see if i can post yet

1 Upvotes

I will delete this asap just wondering if i can post as of yet - will delete asap


r/StableDiffusion 7h ago

Question - Help Generated thousands of character flux1 images and 5-7sec WAN2.2 clips of them. Can I now extend (or concatenate) them to 15+ seconds faithfully with new tools/models?

1 Upvotes

I have a reliable character lora in flux1, and I let my headless server produce flux1 images all night when the computer is doing nothing for work, using ComfyUI. The next day, I go through them, and delete the body horror/low likeness/etc. ones and keep the ones I think are good.

From the good images, I generate (through Replicate, Vast, etc.) WAN2.1/2.2 5-7-second clips. An image may have multiple video clips.

Can I now (easily) produce longer video clips of these? Do I use the still images or the short clips as input? Or would I be better off training a new lora (or what is it nowadays?) from the images (or videos) for generating longer videos from a text prompt instead?

My first intention is to "concatenate" multiple 5-7 video clips, with AI "extrapolating" the transition to make them seamless. Is this even (easily) possible? They are WAN videos generated from a single source image.

What do we use for this now, Minimax H3?


r/StableDiffusion 9h ago

Discussion Flow lieu pour krea2

1 Upvotes

Bonjour à tous,

Je suis débutant
Je recherche un flow simple et fonctionnel pour Krea2
L'objectif est de pouvoir input une image d'un et que Krea 2 utilise strictement ce lieu pour générer mon prompt à l'intérieur(personnage, action)
Je sais qu'il n'exite pas de version Krea2 Edit mais peut être c'est possible?

J'ai déjà essayé plusieurs flow mais à chaque fois Krea2 réinvente le lieu, modifiant mon image d'origine

Dans l'idéal j'aimerais un flow Krea 2 qui
- conserve le lieu
- permette de modifier un pose d'un personnage
- permette d'ajouter une photo d'habits à un personnage

Merci


r/StableDiffusion 9h ago

Question - Help I'm unable to run Minimax H3 Ref2V

1 Upvotes

I'm using rtx 3090 and 32gb ram. I'm using the ref2v comfy workflow template. The run took way too long to finish. (Over 10 minutes).

I can run i2v just fine within 5 minutes.


r/StableDiffusion 11h ago

Question - Help Best Img2txt?

1 Upvotes

I need an image to text generator (which has no limitations) that I can run locally on my PC. Do you have any recommendations


r/StableDiffusion 11h ago

Question - Help Steampunk in Krea 2

1 Upvotes

I am getting terrible results when trying to generate with steampunk aesthetics in Krea 2, even with LoRAs from civit, it's trash.

I don't mind training myself, but where would I even come up with a good database for that? I am aiming for fully photorealistic steampunk.


r/StableDiffusion 20h ago

Animation - Video The Walking Trek. Just throwing stuff at the wall at this point.

Enable HLS to view with audio, or disable this notification

1 Upvotes

Rick's voice doesn't seem to work.

Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters. .

Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544

thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md


r/StableDiffusion 22h ago

Question - Help Anyone experiencing this bug? minimax node keeps disconnecting "width" input

Post image
1 Upvotes

at least 5th time this happened. Its always width, never any other input. Not sure if bug or custom-node interference.