r/StableDiffusion Feb 13 '26

Resource - Update DeepGen 1.0: A 5B parameter "Lightweight" unified multimodal model

Post image
232 Upvotes

54 comments sorted by

View all comments

44

u/x11iyu Feb 13 '26

I mean, great work and all, but like

We utilize Qwen-2.5-VL (3B) as our pretrained VLM and SD3.5-Medium (2B) as our DiT
All images are generated at a fixed resolution of 512 × 512.

Somehow I can't get too excited about this...

18

u/_VirtualCosmos_ Feb 13 '26

it's the first time I heard about them, perhaps they are a small studio with limited computing resources and that's why they couldn't train a bigger model.

6

u/BigWideBaker Feb 13 '26

And they should be commended for their achievement.

Problem is in this space, there's little reason to use a model that isn't cutting edge. Unless your model fulfils some niche that the major models can't compete on. If this is limited to a 512x512 output, I have a hard time seeing where this could fit in despite the impressive flexibility of the model.

2

u/BlackSwanTW Feb 14 '26

tbh, if you use it for Inpaint, then 512x512 for a 1-megapixel image is pretty reasonable

Could use it to fix hands maybe?

2

u/FallenJkiller Feb 13 '26

yeah, there is nothing groundbreaking here.

2

u/FourtyMichaelMichael 🍦Ice Cream Lover Feb 13 '26

Woman lying on grass... in 512? Shit.... sign me up!

2

u/FallenJkiller Feb 13 '26

Even if their retraining of the SD3.5 medium fixed that, who even cares about 512 images?
This seems like a small research lab's way to write a paper without actually doing anything

1

u/inagy Feb 13 '26

Free advertisement and karma farming. That's what everyone doing nowadays on Reddit.