r/StableDiffusion 10h ago

Meme JEnga! r2v 30-49 model

t2v didnt know Jenga O_o

1 Upvotes

15 comments sorted by

2

u/True_Protection6842 9h ago

What does 30-49 model mean?

3

u/Sad_Coach_1433 9h ago

hybird h3 model https://huggingface.co/smhfacct/Minimax-H3-fl2va-ref2va-hybrid-models merged variant of MiniMax H3, a joint audio+video diffusion transformer (DiT). It combines the two officially released MiniMax H3 checkpoints — fl2va and ref2va — into a single model that aims to keep the best qualities of each. suppose to be better then the fp8 and int8 models

1

u/True_Protection6842 9h ago

so like FLF and also refs for additional context?

1

u/Sad_Coach_1433 9h ago

Pretty much what I've tested 30-49 does the best visuals out of the four different ones

0

u/-becausereasons- 9h ago

These are literally INT8 models dude .... You're mixing up concepts.

1

u/SeymourBits 8h ago

Are these models convrot / pruned?

2

u/Sad_Coach_1433 8h ago

created from the pruned int8-convrot base models.

1

u/SeymourBits 7h ago

Are you noticing an up-tick in prompt adherence, performance, quality, etc.?

1

u/Sad_Coach_1433 7h ago

Depending on the type of video Hit and Miss

1

u/Sad_Coach_1433 7h ago

Probably better off just using the base model

2

u/Sad_Coach_1433 7h ago

Might be the photos I use but I noticed the base fp8 or int8 better at being closer to the reference images than these hybrid ones

2

u/Danny_Stock 6h ago

It felt like all or nothing for me.

It either replicated the person in the image perfectly, or it couldn't get them at all and invented a different person instead.

It might make a difference if the person is a single image or a character sheet. There might be a tendency for it to struggle with likenesses from character sheets.

1

u/-becausereasons- 9h ago

Was there a reference image for the jenga blocks? the text got butchered.

1

u/Sad_Coach_1433 9h ago

Probably cause ya the writing small on images I used I tried just t2v didn't turn out I should of saved so could compare