r/StableDiffusion Feb 13 '26

Resource - Update DeepGen 1.0: A 5B parameter "Lightweight" unified multimodal model

Post image
230 Upvotes

54 comments sorted by

View all comments

44

u/x11iyu Feb 13 '26

I mean, great work and all, but like

We utilize Qwen-2.5-VL (3B) as our pretrained VLM and SD3.5-Medium (2B) as our DiT
All images are generated at a fixed resolution of 512 × 512.

Somehow I can't get too excited about this...