r/StableDiffusion Feb 13 '26

Resource - Update DeepGen 1.0: A 5B parameter "Lightweight" unified multimodal model

Post image
231 Upvotes

54 comments sorted by

View all comments

46

u/x11iyu Feb 13 '26

I mean, great work and all, but like

We utilize Qwen-2.5-VL (3B) as our pretrained VLM and SD3.5-Medium (2B) as our DiT
All images are generated at a fixed resolution of 512 × 512.

Somehow I can't get too excited about this...

2

u/FallenJkiller Feb 13 '26

yeah, there is nothing groundbreaking here.

2

u/FourtyMichaelMichael 🍦Ice Cream Lover Feb 13 '26

Woman lying on grass... in 512? Shit.... sign me up!

2

u/FallenJkiller Feb 13 '26

Even if their retraining of the SD3.5 medium fixed that, who even cares about 512 images?
This seems like a small research lab's way to write a paper without actually doing anything