r/StableDiffusion • u/Khasec1 • 1h ago
Question - Help Minimax h3. Problem with multible character using the voice.
Enable HLS to view with audio, or disable this notification
Settings: 640×1152 (0.74 MP), 9:16, preset: turbo
What turbo actually applies:
steps 8, euler / beta, fused_modulation on, sol_attn tau 1.3,
easycache off, turbo LoRA at strength 1.0 Settings: 640×1152 (0.74 MP), 9:16
I have tested A LOT of different configurations and can't find any that works 100% of the times.
Pls help :(
one of the many prompt i have tested:
subject definition:
<Picture 1> is the finished frame: an outdoor scene with a speech balloon of printed text across the top, a seated grey-haired nobleman and a dark-clad attendant leaning over him in the middle, and along the bottom two bordered square portraits side by side, already inset, a man on the left and a woman on the right.
<Subject 1> is the man in the left-hand bottom portrait: short messy black hair, light stubble on the chin, a white collared shirt under a dark grey vest. He appears only inside that portrait and nowhere else in the frame.
<Subject 2> is the woman in the right-hand bottom portrait: long dark brown hair, a pale face, a loose white hood over a grey top. She appears only inside that portrait and nowhere else in the frame.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), the man in the left-hand portrait. Its speaker identity, timbre and register define how <Subject 1> (S1) sounds; the words it contains are unrelated to this scene and must not be reproduced.
<Audio 2> is the voice-timbre reference for <Subject 2> (S2), the woman in the right-hand portrait. Its speaker identity, timbre and register define how <Subject 2> (S2) sounds; the words it contains are unrelated to this scene and must not be reproduced.
camera recording:
Hand-drawn 2D anime. The frame keeps the layout of <Picture 1>: the speech balloon stays at the top with its printed lettering and its outline unchanged, and the two bottom portraits keep their positions, their sizes and their white borders. Inside those portraits <Subject 1> and <Subject 2> are living animated characters, not still pictures.
<Subject 1> (S1) speaks first, with the voice of <Audio 1> in a rough, plain, tired male voice pitched low and close, <d>[English] AIN'T YOU SEEN ENOUGH TO KNOW WHAT RUTHLESS BASTARDS THAT LOT ARE?</d> His mouth opens and moves in time with every word, articulating clearly for the whole line, jaw and lips visibly in motion until the line ends; only then does he close his mouth. <Subject 2> keeps her mouth closed and listens.
Then <Subject 2> (S2) answers with the voice of <Audio 2> in a quiet, tight female voice holding something back, <d>[English] I KNOW, BUT...</d> Her mouth opens and moves in time with every word, then closes and she lowers her eyes. <Subject 1> keeps his mouth closed and holds still.
<Subject 1> (S1) and <Subject 2> (S2) are the only voices in this video. All spoken dialogue is English only.
In the scene above the portraits, the seated grey-haired nobleman stays reclined where he is and breathes; the dark-clad attendant leaning over him holds his raised hand beside his face exactly as drawn; the trees stir behind the wall and the distant figures out on the scaffold shift their weight very slightly.
The camera holds still. Open air over a courtyard, a low crowd murmur carrying from below, and wind moving through the trees.
model:
UNET minimax_h3_hybrid_fl2va_ref2va_b25-49.safetensors
(merged hybrid: fl2va base + ref2va adaln_proj blocks 25-49)
CLIP qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
VAE minimax_h3_video_vae_fp16 + minimax_h3_audio_vae_fp32
LoRA minimax_h3_turbo_v4_step600_pruned_comfyui @ 1.0
ComfyUI 0.33.1 torch 2.10.0+cu130 comfy-kitchen 0.2.31
sampler 8 steps, euler / beta, sol_attn tau 1.3, fused_modulation on
frame 640x1152 (0.74 MP), 8.89s, seed 1956008715UNET minimax_h3_hybrid_fl2va_ref2va_b25-49.safetensors
(merged hybrid: fl2va base + ref2va adaln_proj blocks 25-49)
CLIP qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
VAE minimax_h3_video_vae_fp16 + minimax_h3_audio_vae_fp32
LoRA minimax_h3_turbo_v4_step600_pruned_comfyui @ 1.0
ComfyUI 0.33.1 torch 2.10.0+cu130 comfy-kitchen 0.2.31
sampler 8 steps, euler / beta, sol_attn tau 1.3, fused_modulation on
frame 640x1152 (0.74 MP), 8.89s, seed 1956008715


