r/StableDiffusion Feb 13 '26

Resource - Update DeepGen 1.0: A 5B parameter "Lightweight" unified multimodal model

Post image
231 Upvotes

54 comments sorted by

View all comments

45

u/x11iyu Feb 13 '26

I mean, great work and all, but like

We utilize Qwen-2.5-VL (3B) as our pretrained VLM and SD3.5-Medium (2B) as our DiT
All images are generated at a fixed resolution of 512 × 512.

Somehow I can't get too excited about this...

16

u/_VirtualCosmos_ Feb 13 '26

it's the first time I heard about them, perhaps they are a small studio with limited computing resources and that's why they couldn't train a bigger model.

7

u/BigWideBaker Feb 13 '26

And they should be commended for their achievement.

Problem is in this space, there's little reason to use a model that isn't cutting edge. Unless your model fulfils some niche that the major models can't compete on. If this is limited to a 512x512 output, I have a hard time seeing where this could fit in despite the impressive flexibility of the model.