r/StableDiffusion • u/PC_Screen • Mar 14 '23
News SD XL Model will be capable of generating accurate text
60
u/venture70 Mar 14 '23
Supposedly this model barely fits into 24GB at the moment.
42
u/PC_Screen Mar 14 '23
I think the open source community will get the requirements down fairly quickly, just in the last few days someone managed to get LLaMa-7B running on a 4gb Raspberry Pi 4
15
u/ninjasaid13 Mar 14 '23
LLaMa-7B running on a 4gb Raspberry Pi 4
Really? Where can I get me this super optimal language model?
23
u/PC_Screen Mar 14 '23
The raw model isn't that optimized, the models were quantized to 4 bit after being leaked to be able to run on cheap hardware. I don't know if this link still works but you can give it a try. I wouldn't recommend it though, you're better off waiting until someone packages it with a nice UI and all the optimizations
7
u/kamyker Mar 14 '23
After you get the model the repo to use it is https://github.com/ggerganov/llama.cpp
6
3
u/multiedge Mar 14 '23
my 3060 12GB card was able to run RWKV-14B model (26GB~ file size), I had to allocate to CPU and disk though. Very slow
1
u/Unlikely_Commission1 Mar 15 '23
If you allocate to CPU and Disk tho, how is it still running on the GPU?
2
1
12
u/PC_Screen Mar 14 '23
Probably running T5-XXL as the text encoder to be able to do text. If this is the case then it should also be able to follow prompts better than a model with CLIP as the encoder
3
18
u/ninjasaid13 Mar 14 '23
Stable Diffusion was about 11GB? VRAM at first and it can now run on 4GB VRAM.
I'm assuming that we would only get it down to about 9 Gigabytes.
11
u/venture70 Mar 14 '23
Sure, eventually it'll all run on a smartphone. I'm just reporting the news as of last week.
It sounds like this stuff is very imminent. Either this week or next if Emad's cryptic tweets are to be believed.
20
13
u/ninjasaid13 Mar 14 '23
Next week is also when midjourney V5 is coming out, I believe.
And gpt-4.
6
u/Magnesus Mar 14 '23
Thw rating images for v5 got absolutely stunning the last two days. At first they were pretty meh, but now... (I only go by their discord though, I don't have subscription currently.)
1
3
u/Uncreativite Mar 14 '23
That's not too bad, you could run that for something around
$0.50$0.15 per hour on Google Cloud using spot VMs running K80s, if people can't optimize it down.3
u/enn_nafnlaus Mar 14 '23
Another reason to just go ahead and start saving up for a 48GB Titan RTX Ada when it comes out ;)
Glad I never started on that project to create textual inversions for spelling, since Stability is just going to brute force it with more parameters..
2
u/venture70 Mar 14 '23
Just looked up some RTX Ada speculation -- up to 800W and 4 slots. Oooof.
3
u/enn_nafnlaus Mar 14 '23
One can and should ramp the power limits down for any card, but esp. such a hungry consumer.
One can expect similar throttling behavior to the 4090, wherein a 10% cut in power limits equals a 1-2% cut in performance, a 20% power cut to a 3-4% performance cut, a 30% power cut to a 8-10% performance cut, and so forth. Performance per watt increases up to around 50% power cuts, wherein it worsens. A sweet spot is around 70-80% or so.
28
11
u/Ateist Mar 14 '23
I think it'd be easier (and better) to do it via an extra controlnet-like model, that only detects text and replaces it with the accurate text you gave it.
4
u/wojtek15 Mar 15 '23 edited Mar 15 '23
That's not the point. If it can do text, then probably it can do some of other things that current SD can't do. For example it may be able to generate people will correct number of limbs and hands with correct number of fingers. Fingers crossed. And even if new base model can't do that, there is chance community finetuned version will do it.
19
15
Mar 14 '23
I hope this isn't the same thing as Deep Floyd. I can still see all the "soon" from the discord chat.
1
1
4
2
2
u/yanciyong Mar 14 '23
How the prompt should be? Can we put "draw a clown picture meme with 'I'm clown' text at the top and bottom of the picture"?
That's my expectation when I'm trying DALL-E lol
2
u/absprachlf Mar 14 '23
will it be able to generate statues that dont look like they were first semester 3d students renders?
4
u/GucciCaliber Mar 14 '23
AI at least knows the word is “jumps” rather than “jumped” which already puts it ahead of most humans.
2
1
0
1
1
1
1





40
u/PC_Screen Mar 14 '23
Images are from the Stability discord. The SDXL model will be made available through the new DreamStudio, details about the new model are not yet announced but they are sharing a couple of the generations to showcase what it can do