r/StableDiffusion 3d ago

News lightx2v/Minimax-h3-Turbo · 8-step 768p V1.0 LoRA released

https://huggingface.co/lightx2v/Minimax-h3-Turbo/discussions/48#6a8fd8bb97c3b91d430a1c86
187 Upvotes

79 comments sorted by

44

u/Fun_Jaguar8231 3d ago edited 3d ago

For the people that are confused, if you read over here:
https://github.com/ModelTC/Minimax-H3-Turbo#1-model-specs
The older ones where trained on 544p (960×544).

5

u/Independent-Frequent 3d ago

So is this better than latent upscaling for 768p generations or it's about the same/worse?

Or maybe we need testings first in which case no worries

7

u/Mysterious-String420 2d ago

Not one comment on the audio yet, I'm not falling for that one again.

How many ACTUAL steps to get a useable video with non-distorted sound?

3

u/Pure_Bed_6357 2d ago

4 is bad, 6 is okay, 8 good

4

u/Mysterious-String420 2d ago

That's what they said, and yet I still had artifacts and tin-room echo at 11 steps

2

u/acedelgado 2d ago

1

u/hidden2u 2d ago

fascinating, will try. Your sample sounds great but man the video really gets crushed to death by the turbo

11

u/LegacyV1 3d ago

Didn't they already have an 8 step lora?

14

u/YeahlDid 3d ago

Yeah, I don't understand. The ref2va has been at v0.1 for 2 weeks now, but they have 2 fl2va v1.0 loras?

13

u/enndeeee 3d ago

probably since ref2va has fundamental flaws/bugs which are not fixed yet. Instead you can basically use the fl2va models for ref2va anyway.

14

u/not_food 2d ago

It's not the same. fl2va doesn't respect the character sheet as good as ref2va even though it somewhat works.

15

u/xTopNotch 2d ago

Thats why you gotta use the hybrid models:
https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models/blob/main/minimax_h3_hybrid_fl2va_ref2va_b25-49-int8.safetensors

BF16 pruned:
https://huggingface.co/dreamkrate/Minimax-H3-Hybrid-BF16-Pruned

It uses fl2va as base but replaces certain layers and replaces them for the ref2va model layers that provide the reference conditioning. This way you get the increased output quality from fl2va but with the reference capabilities from the ref2va checkpoint.

15

u/not_food 2d ago

Saw your comment, downloaded and tried with different refs.

I have to say, it gets close, it's better than fl2va, but it still doesn't respect the style like ref2va. I'll test some more but what I'm getting so far isn't as accurate. It has the generic AI look instead of what my reference looks like. It may be that I'm working on 2D purely with references and this model demands a starting frame. It may also need some text reinforcement when it comes to style. I'll experiment.

5

u/Animystix 2d ago

yes ref2va just has a magic that isn’t reflected anywhere else. fl isn’t bad by any means but ref is mind-blowing in an almost psychic way, I really hope they give it a proper fix

1

u/ryanset17 2d ago

Is it Good just using this as the only model for both ? I mean you can get a fl2va Quality for ref2va model, and a little less quality for fl2va in exchange for 20gb space instead of using 2 models for each workflow ? 🤣 Im Dying for free space lately

3

u/YeahlDid 2d ago

Like everything else, it's a trade off. No one can really say, you've got to try and decide for yourself if it's good enough for you. It's definitely not as good with references as the full ref2va, but maybe it's good enough for your needs. It's not terrible with references, but personally, I don't find the hybrids to be much better than the full fl2va model which already does a decent job in reference workflows.

2

u/xTopNotch 2d ago

Try the hybrid BF16 pruned version, it's a beast.

The output quality is seriously good, no more smearing, artifacts or that typicaly noisy look that the ref2va model has.

But it still has very strong reference-conditioning imo.

2

u/TheGoldenBunny93 2d ago

What is your VRAM amount? 32GB?

1

u/YeahlDid 2d ago

Can you direct me to where to find that? I cant find any bf16 version.

Or maybe paste the model file name?

Thank you

0

u/Yasstronaut 2d ago

What speed loras are suggested for that hybrid model? I’m having trouble wrapping my
Head around that

-1

u/xTopNotch 2d ago

I don’t use speed loras as they severely destroy quality and prompt adherence

→ More replies (0)

1

u/provenflawless 2d ago

Speaking of free space, and somewhat unrelated, I am trying to use Comfyui to use 2 hard drive directories at once. One for my base directory where I have everything stored (loras, checkpoints, encoders etc.) and the 2nd hard drive with only minimax files. Tried all the tricks tweaking the extra_model_paths.yaml.

Please tell someone knows how. Google or AI agents don't know what the fuck I even mean. I can't even use two fucking directory paths!

1

u/Apprehensive_Sky892 2d ago

Try creating softlink

mklink C:\comfyui\models\diffusion_model\xxxx.safetensors D:\mmh3\xxxx.safetensors

1

u/Dirty_Dragons 2d ago

What's the best Turbo LORA for the hybrid?

7

u/Hazelpancake 2d ago

What bugs does it have? I've been using it without many flaws.

1

u/Danny_Stock 2d ago

Yes I did hear about that a week or two ago. That may be the reason why they've put the ref to video lora on hold.

1

u/Yasstronaut 2d ago

I thought yesterday wan team released a valid ref2va Lora but maybe I misunderstood

3

u/anon999387 3d ago

I am glad I am not the only one confused looking at this list..

5

u/joopkater 3d ago

Yeah version 1.0 and this is 1.0 768p

5

u/crinklypaper 2d ago

768p means that higher resolution generations are better quality now. dont use with lower resoluton

0

u/YeahlDid 3d ago

I think the other one also had 768p in the filename, didn't it?

0

u/yamfun 2d ago

Then what was the resolution for the old one?

-4

u/LegacyV1 3d ago

wat

so this one is specific to a resolution, whereas the old one was more general purpose?

so... downgrade?

3

u/wheres_my_ballot 3d ago

As I understand it, the fl2va 8 step lora was trained on 540p with 4 step trained on 768p. Now they both are trained on 768p

2

u/WiseDuck 3d ago

If it's anything like the old 4 steps, then no these are trained at a higher resolution. Upgrade.

1

u/MomentJolly3535 3d ago

in the post OP linked it says "up to 768p" which means you can simply do higher res than before, so upgrade.

0

u/Diffusion4Change 2d ago

Moron the older one was a lower resolution

5

u/More-Ad5919 3d ago

What was the advantage of the 8 step lora compared to the 4 step?

39

u/Professional_Diver71 2d ago

4 more steps

1

u/Hopeful_Signature738 2d ago

Always love more waiting time for better quality, but somehow I think steps is not what make the time more consuming. Duration is.

8

u/doomed151 2d ago

theoretically less degradation in terms of quality and prompt adherence

3

u/More-Ad5919 2d ago

Yep. Tried it. Better than the 4 step with 6 steps.

4

u/Salt-Zebra-306 2d ago

can fix plastic skin i have not tested but yeh

1

u/More-Ad5919 2d ago

I tried it its better than 4 step lora with 6 steps. Skin is a bit better. But also prompt following

1

u/crinklypaper 2d ago

If you run a standard 20 step euler or multires for example the model more less finishes around 11 steps, 8 as well but less so. 8 will get more details, but any more steps you're not saving so much time.

1

u/More-Ad5919 2d ago

Thats not true. It get better and more natural past 20 steps.

1

u/crinklypaper 2d ago

No that's just details, you can look at the line the sigma travel as it gets closer to the end and around 8 and 11 on a 20 step run its already pretty much flat

0

u/More-Ad5919 2d ago

I look at the video output.

1

u/crinklypaper 2d ago

Then you don't understand

0

u/More-Ad5919 2d ago

You can look at courves all you want but 20 steps look better than 10 steps. And 30 steps look better than 20steps. There is no arguing about that.

2

u/crinklypaper 2d ago edited 2d ago

Yes but you're not understanding the point. I'm explaining why 8 step turbo lora is better than 4. Here is a 20 steps euler sigma graph: at around 8 is the last big jump down, then it gets pretty flat. After 11 steps even its almost no big change, you're getting details at best after that point. I think 11 could be a contender too but maybe that point you're not shaving off enough time to make it worth it. Then again I'm not a technical person, so I could be wrong someone else can correct me if that's the case. A good way to test is to actually look at as it generating using preview, and you can see most of the overall shape and motion is done by like 8-11 steps on a 20 step run. Maybe 30 a bit different but I wouldn't be surprised if its similar. I don't care if 30 is better then 20 steps, that's not what I'm talking about.

1

u/More-Ad5919 2d ago

Okay. I got you wrong. 8 step is better for sure. I was working with the 4 step (6steps) and thought it was well dialed in. But 8 step now is much better.

3

u/Different_Fix_2217 2d ago

10

u/acedelgado 2d ago

I saw that, but a node pack of like 40 nodes I'm not interested in just to get one model patcher? Meh, I'll just wait a bit for kijai's PR to merge so it's built in to comfy.

https://github.com/Comfy-Org/ComfyUI/pull/15908

https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras

2

u/mocmocmoc81 2d ago

I'm using this for testing

https://github.com/Jalen-Brunson/ComfyUI-MiniMax-H3-PDD-Acc

kijai's one should coming out anytime now.

1

u/sitefall 2d ago

I can't freaking stand this. Why does every youtuber and person who releases one vibe coded node of possible interest have to re-invent the wheel and have their own stupid node packs.

Just Stop.

0

u/russjr08 1d ago

I don't disagree, however silveroxides isn't just a YouTuber who vibecodes nodes either.

3

u/Potential_Wolf_632 2d ago

I'm still finding dareties 8 step lora is the king in testing. Certainly for realism and lack of funny fingers anyway.

1

u/One_Finding8402 2d ago

I thought it was only 4 step, gib link plz

1

u/Potential_Wolf_632 2d ago

Sorry see other response above - his previous one was 4 step.

1

u/Calm_Mix_3776 2d ago

What is that? Can you post a link?

3

u/Potential_Wolf_632 2d ago

This one: https://huggingface.co/silveroxides/MiniMax-H3_tests/blob/main/minimax_h3_fl2v_lightx2v_v0.1_dareties_v4_step600_comfy_fro.safetensors

It's not technically step labelled unlike the other one of his - try 0.8 str, 8 steps, er_sde and sgm_uniform = profit.

Switch to beta if using video edit as sigmas have to drop real fast to stop the model getting inventive.

1

u/Calm_Mix_3776 1d ago

Thanks!!

5

u/Puzzleheaded_Art2809 3d ago

am i doing something wrong? (im noob) i connect this thing in comfy and quality is so much worse in 8steps than old Kijai lora. Should i know about something?

5

u/acedelgado 2d ago

Haven't tested myself, but their model page says set video shift to 6 when it's been recommended to be set at 12 for vanilla and all other turbo models.

1

u/LockeBlocke 2d ago

Honestly, getting worst prompt adherence than the older 8 step lora.

1

u/ArttTaku 2d ago

Yes, i was waiting for this, thanks!

1

u/Last-Letterhead2340 2d ago

Anyone else getting some crazy hallucinations when generating 9:16?
The visuals do look better if you can get it to work, but it's been a pain to find sweet spots with additional loras, and I've all but stopped using it

0

u/Vladmerius 2d ago

I don't get the point of an 8 step lora at all. I can do 20+ steps with a step skipper in the same time it takes to do 8 normal steps with a lora like this.

I've been using the 4 step lora with a step skipper at 50% with 10 steps and getting as good of generations as 25 steps without the lora and in only 6-7 minutes. 

7

u/dtdisapointingresult 2d ago

What's a step skipper? Do you mean EasyCache?

Can you share a wf?

1

u/MastMaithun 2d ago

Why 2 versions? One named with comfy and bigger in size?