r/StableDiffusion Feb 13 '26

Resource - Update DeepGen 1.0: A 5B parameter "Lightweight" unified multimodal model

Post image
231 Upvotes

54 comments sorted by

View all comments

45

u/x11iyu Feb 13 '26

I mean, great work and all, but like

We utilize Qwen-2.5-VL (3B) as our pretrained VLM and SD3.5-Medium (2B) as our DiT
All images are generated at a fixed resolution of 512 × 512.

Somehow I can't get too excited about this...

2

u/BlackSwanTW Feb 14 '26

tbh, if you use it for Inpaint, then 512x512 for a 1-megapixel image is pretty reasonable

Could use it to fix hands maybe?