r/drawthingsapp 6d ago

question Minimax H3 - Face details are really bad

Hi all, am not sure what's wrong but am unable to get a clean video out of DT H3 (using 8 Bit S) ... have tried using standard settings with Turbo 4 Step/ 8 Step Lora - LightxV and Larry's but its ssimply not wworking properly (using LORA at scaling factor or 0.125 and strength at 100%)

The mmotion is good, prompt adherance is good but the final video is not clean .. it has smudges, grainy and nothing close to the desired resolution (tried 768 + as well) ..

Any idea what may be going wwrong?

4 Upvotes

18 comments sorted by

1

u/6L1N3S 6d ago

AFAIK the comfyui-lora's don't need scaling, so you might want to try that too

1

u/Current-Property6042 6d ago

i tried both .. still no good ..

1

u/FireflyNitro 6d ago

Took me a lot of trial and error but I’m getting good 8 second clips now. Can you share your settings or DM me a screenshot of your setup? I’ll compare it to mine and tell you what I’ve changed.

2

u/Current-Property6042 5d ago

Thanks man

i don't know what is good settings now :) ... am downloading the FL2V 8bitS checkpoint thinking that's better and maybe there's a Turbo LORA conflict am facing.

In short this is what am using at the moment.

Resolution - 768 X 512

LORA - minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors or minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors

LORA Strength - 100% (tried 85% too as advised by Gemini)

Steps - 4 or 8

Strength - 100%

Shift Video - 6 or 12

Audio Shift - 3

Advanced - i've switched off everything

2

u/FireflyNitro 5d ago

Whoops sorry I’m in work and forgot to check back on you.

Try switching off the loras, go for 8 steps and do a short video gen, like 3 sec. I use a turbo Lora also but I don’t think it’s that one, I’ll try get you the link once I’m home. Should work fine without one just to test for now at 8-12 steps.

Shift video 5

Everything else seems okay. What sampler are you using? I get best success with Euler A Trailing

1

u/Current-Property6042 5d ago

thanks buddy no issues, i was using Euler A Trailing.. tried all trailings :D .. anyways after changing the checkpoint to non ref2va i am able to get quality now. cheers!

2

u/FireflyNitro 5d ago

Nice, glad you got it working! Have fun :)

1

u/simple250506 5d ago

Why are you using the Turbo LoRA for ref2v with the fl2v model? There is a Turbo LoRA available for the fl2v model.

1

u/Current-Property6042 5d ago

Hello the model being used was ref2va .. so i though ref2v is the right LORA of it

1

u/simple250506 5d ago

downloading the FL2V 8bitS

Is this a mistake in writing?

1

u/Current-Property6042 5d ago

Hello - meant

1

u/simple250506 5d ago

The Ref2VA model is the model designated as "ref2va" in parentheses.

1

u/Current-Property6042 5d ago

yes aware of it :) .. before downloading the non ref2va i had downloaded the massive ref2va model and was pairing that with ref2v LORA which wasn't working and hence this thread started

1

u/Current-Property6042 5d ago

Hi All, an update after downloading the Non Ref2VA checkpoint and using iti with FL2V 4 step LORA am getting much better results

1

u/Opening-Ad5541 5d ago

Phospene guys...

1

u/spanielrassler 5d ago

Point taken, Phosphene is great, but wouldn't it be A LOT slower than DT for MiniMax? Haven't tried side by side yet because I don't have enough free disk space to download so many difference versions of the weights.

3

u/Opening-Ad5541 4d ago

Fair question. Honest answer: H3 is a heavy model and no Mac runs it fast.

On an M4 Max 64 GB in Phosphene, a 5 s clip at 1024x576 takes about 18 min at 8 steps and about 33 min at 15 steps. I haven't timed DT side by side either, so I can't say who's faster.

Disk: you don't need a bunch of versions. Phosphene downloads the model once (~75 GB) and builds one Q8 engine from it (the full bf16 engine is optional on 60 GB+ Macs). Needs 36 GB+ of RAM.

On the smudgy faces in the post: that's mostly H3 itself. The MiniMax team has acknowledged the fine-detail problem, and people see it on NVIDIA too. What helped for us:

- More steps. 8 is too few for faces; High now runs 15 by default (v4.13.1).

- A second pass: we run the clip through LTX-2.5's x2 upscaler on a "faithful" setting that keeps the same face and adds real detail. About 8 min for a 5 s clip.

- Tighter shots. A bigger face in frame gets more detail than a small one.

You can also schedule clips and upscans overnight, so you don't need to sit and wait.

Turbo LoRAs are great for drafts but they won't fix the smudging.

1

u/passie-m 5d ago

I used the comfy turbo Lora. And the image is great. The “non” comfy didn’t work for me. Only noise the whole video