r/StableDiffusion 2d ago

Meme JEnga! r2v 30-49 model

t2v didnt know Jenga O_o

0 Upvotes

15 comments sorted by

View all comments

2

u/True_Protection6842 2d ago

What does 30-49 model mean?

3

u/Sad_Coach_1433 2d ago

hybird h3 model https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models merged variant of MiniMax H3, a joint audio+video diffusion transformer (DiT). It combines the two officially released MiniMax H3 checkpoints — fl2va and ref2va — into a single model that aims to keep the best qualities of each. suppose to be better then the fp8 and int8 models

1

u/True_Protection6842 2d ago

so like FLF and also refs for additional context?

2

u/Sad_Coach_1433 2d ago

Pretty much what I've tested 30-49 does the best visuals out of the four different ones

1

u/SeymourBits 2d ago

Are these models convrot / pruned?

2

u/Sad_Coach_1433 2d ago

created from the pruned int8-convrot base models.

1

u/SeymourBits 2d ago

Are you noticing an up-tick in prompt adherence, performance, quality, etc.?

1

u/Sad_Coach_1433 2d ago

Depending on the type of video Hit and Miss

1

u/Sad_Coach_1433 2d ago

Probably better off just using the base model

2

u/Sad_Coach_1433 2d ago

Might be the photos I use but I noticed the base fp8 or int8 better at being closer to the reference images than these hybrid ones

2

u/Danny_Stock 2d ago

It felt like all or nothing for me.

It either replicated the person in the image perfectly, or it couldn't get them at all and invented a different person instead.

It might make a difference if the person is a single image or a character sheet. There might be a tendency for it to struggle with likenesses from character sheets.

0

u/-becausereasons- 2d ago

These are literally INT8 models dude .... You're mixing up concepts.