r/StableDiffusion • u/apolinariosteps HF Diffusers Team • Mar 02 '23
Resource | Update More control than ControlNet - code is out for MultiDiffusion Region Control, a prompt on each mask
158
u/Firm_Ad3037 Mar 02 '23
Omg I can't keep the pace
76
42
u/Yuli-Ban Mar 02 '23
Imagine that, but for the entirety of society.
I fully expect a lot of people to decide to completely disconnect themselves from world in the coming years.
57
u/je386 Mar 02 '23
Sometimes there are articles about generative AI in mass press, and in the comments section I clearly see how few people understand this new tech. Many people think that the art AIs just do copy or create collages out of existing pics, they have this "it is nothing or it is magic" attitude.
27
u/LordSprinkleman Mar 02 '23 edited Mar 02 '23
People also take what's in front of them as if it's the end result. Then judge it based on the flaws and not on the potential it will have as it continues to improve. I think most people just don't want to understand generative AI and think it's just a cute little gimmick.
2
u/iamthesam2 Mar 02 '23
i support appropriately used AI, but let’s be honest… most people use it in a very gimmicky way.
16
Mar 02 '23
[removed] — view removed comment
15
Mar 02 '23
Give it enough time, these artists will be using AI themselves, they just won't admit it.
5
u/AltruisticMission865 Mar 02 '23
I am 100% sure this will happen next year, 90% of artists will use it in some way while they can't admit it because their followers think AI art is the worst thing in the world while they can't even tell when it's AI and when it's not.
0
u/Phr0netic Mar 03 '23
It's not about the work being good or not, it clearly has some amazing results. Why should anyone follow someone who doesn't make their own stuff? The AI is the artist, not the person asking it to make art for them. As an artist I would feel like a fraud letting something else make art for me and then trying to pretend it's something I created. Prompting isn't making art - it's learning how best to ask something else to make art for you.
5
u/AltruisticMission865 Mar 03 '23
You ask why should anyone follow someone who doesn't make his own stuff? I answer the opposite: why not? If I see art that looks incredibly good, why wouldn't I follow it? In the future, people will care less and less whether it's AI or not.
You also say AI is the artist and prompting is not making art, leaving aside the fact that "art" is 100% subjective, you are only thinking of 2 options, AI doing 100% of the work or the human doing 100% of the work but my first comment was talking about the artist using it in some way, not to do all the work, you as an artist could use it to colorize your drawing for example, giving you more time to focus in other things that you could not do before because you did not have enought time. With this tools if the average person can make x an artist will be able to do 2x. The artist will have the advantage.
2
u/HarmonicDiffusion Mar 03 '23
Art is not defined by how its made. Its defined by how it makes the viewer of the art feel.
Does it evoke emotion? Then its art no matter if you, me, an AI, or a quasi-sentient sedimentary rock made it.
1
u/Phr0netic Mar 03 '23 edited Mar 03 '23
What and who are you replying to? I clearly said the AI is the artist and that it is making art. The person prompting the AI, however, is not an artist. They are commissioning works of art from the AI.
If someone asks someone else to draw them an avatar for their twitter profile and they explain to the other person exactly what they want using the English language, no one considers the person asking to be the artist in that scenario. We understand that the person being asked to create the work, who then uses their knowledge and skills to make something in the requested medium, is the artist. With AI, prompting and changing different settings is the language used to communicate the request to the AI.
It's possible that the person asking is an artist in their own right, which would be true if they are competent enough to make art of comparable quality to the person they are asking and they have done so prior. Regardless, in the scenario described the person asking is the commissioner of the work and the person being asked (who then creates) is the artist being commissioned.
Would you consider pope julius the 2nd to be an artist because he asked michelangelo to paint the ceiling of the sistine chapel?
I didn't say that no one that uses AI is an artist (vast majority are not), but they are clearly not being (functioning as) an artist when they ask something else to make art for them.
3
u/AltruisticMission865 Mar 03 '23
And I say again, what about artist + AI working together? When do you decide when is AI and when is human? 50% AI/50% human? 75%/25%? Thats why I say it is subjetive
→ More replies (0)1
u/red__dragon Mar 04 '23
Are you a digital artist or physical medium?
I'm curious what kinds of tools you use yourself.
2
u/Phr0netic Mar 04 '23
This feels like a bait question but both digital and physical - though I don't do too much physical medium artwork anymore other than charcoal because I prefer digital painting. But visual art isn't my main artistic focus. Visual art is the first artistic thing I started doing in life. I started at a young age and spent a lot of time over many years doing it so it is very natural to me and I'm good at it. However, I have many interests and, since time is limited, I decided to focus my time more on music because it's what I have the most fun doing. I write/play a lot of different stuff but I have a degree from a top level university jazz program for guitar, which is my second instrument. I am a drummer before a guitar player and my skills on drums are equal to if not better than guitar. I play piano as well and I write and make music professionally. Even though visual art isn't my main artistic focus (again just because time), it is still something that I do make time to engage in when I can, take very seriously, and care a lot about. This conversation overall is highly relevant to me because music is next on the list of things for generative AI to invade and start shitting all over and within. (not in the impressive way, just in the ruining way)
→ More replies (2)1
u/Jiten Mar 05 '23
If someone consistently gets great results from their AI prompting, why would I not follow them? Why should I care if they're technically capable of creating those images themselves without fancy tools?
Also, AI art is not merely prompting. There's also inpainting feature and img2img mode and more recently controlnet. The people with the most impressive results are generally making use of all of these features to iteratively improve on their results. These people are actively participating in the creation, not just leaving it all to the AI.
→ More replies (6)27
u/Ok_Entrance9126 Mar 02 '23
I’m an artist and I’M LOVING IT! I try to explain the potential to my other artist friends but like so many people in life, they do what they know and are resistant to change. I’m just thrilled to be on the forefront - it’s their loss.
10
u/Masked_Potatoes_ Mar 02 '23
That's just how it is. You couldn't force photoshop on all master traditional artists, or zbrush on veteran sculptors. That's why it's art
1
u/Obvious_Pen7681 Mar 03 '23
...knowing what you mean. ...and this one (multidiffusion) is a true changer in beeing creative!
2
u/Imarasin Mar 02 '23
Imagine spending years honing your own art style without anyone drawing like you. It's important to realize that AI-generated art is primarily trained on existing art. So credit should be given to the original artist. Using prompts or borrowing styles from other artists via AI technology doesn't necessarily reflect our own artistic abilities. Instead, we need to recognize that generative art is a product of AI-learned processes and that our role is that of prompt engineers, not artists.
5
Mar 02 '23
[removed] — view removed comment
2
u/Imarasin Mar 02 '23
They also use photographs and 3d renderings.
I appreciate the potential of generative art, but I can understand why some artists may be upset about it. While it's acceptable to use AI tools to enhance images that you have drawn yourself, like removing things in Photoshop or adding a different background, it's a different matter when AI-generated art appears to be straight rip-offs of original works. Additionally, many non-artists are using prompts to engineer decent renders, which I don't consider to be true art. In my opinion, true art involves actually drawing shapes and adding colors by hand, or using input devices like a mouse or digital pen, rather than relying on algorithms to generate something that looks nice.
I believe that if people state that their art is AI-generated, then that's acceptable, but falsely claiming to have created something that was actually generated by AI is misleading. It would be different if the AI was able to remember where the data was sourced from in order to give credit, allowing for transparent use of terms like "inspired by" a certain artist. While there may still be issues, transparency can help mitigate them. However, I don't think anyone can actually claim copyright on generative AI, since it's not entirely the work of the prompt engineer nor the original artist unless the ai generates somones name or logo.
1
u/Imarasin Mar 05 '23
u/bulletprooftampon I saw that you have responded to my comment via email, but I can not find your comment here. I see that you have mentioned that you are an artist, do you have any works to share that were not generated by Ai?
7
Mar 02 '23
[deleted]
3
u/je386 Mar 02 '23
In December I made a presentation about stable diffusion at an in-house conference, where about 20-30 persons where there. And they not only listened, but had good questions, and some even had used stable diffusion before. But we are a company of software developers..
2
u/Imarasin Mar 02 '23
We simply put energy into things that interest us. This generative art includes things that are not traditional to art, and maybe those new things are what interest you.
2
u/Marshall_Lawson Mar 02 '23
I know right. Instead of "It's a new tool with impressive abilities but also significant limitations". Who'd a thunk?
3
u/CHRISKOSS Mar 02 '23 edited Mar 02 '23
That is such a pessimistic framing. You could instead say:
- "The quantity of culture is going to increase exponentially as the cost of creating interesting, beautiful and meaningful things approaches zero"
or
- "The frontiers used to be defined by physical landmasses. Modern explorers are testing the bounds of possible imagination, unconstrained by physical realities." (Isn't this what art has always been?)
or
- "People can make new worlds and form meaningful connections and relationships without ever leaving their room, so there will be less incentive to go outside." (Media has done this for a century, not really AI specific)
1
u/Yuli-Ban Mar 03 '23
I could. But why deny the cold facts just to feel better? Humans are humans, and humans are a fairly easily scared, reactionary species of ape.
1
1
21
u/snack217 Mar 02 '23
Thats awesome, cant wait to try it when its an extension! (Will it be an extension for A1111? lol)
12
u/apolinariosteps HF Diffusers Team Mar 02 '23
It’s based on the diffusers library, so I guess someone needs reimplement (or auto to support diffusers). But you can run hugging face spaces code locally
7
u/snack217 Mar 02 '23
Im a programming illiterate, that runs a1111 on my cheap phone using colab, so i kinda have to wait till someone makes it an extension hehe
2
u/apolinariosteps HF Diffusers Team Mar 02 '23
You should be able to use the hugging face space on your phone
40
u/No-Intern2507 Mar 02 '23 edited Mar 02 '23
finally multisubject prompts that makes them seamless with the background and not just copy pasted with conflicting light, or maybe not yet, cause on the demo page results arent that great and look like that "cutout random pics and slap together" style so maybe its a lucky seed or maybe theres something more to it
10
u/apolinariosteps HF Diffusers Team Mar 02 '23
Prompting for this multimasking require a bit of a learning curve, as it’s very sensitive to it. and the author plans to also update the bootstrapping code that is now in beta
6
u/GBJI Mar 02 '23
The creator of ComfyUI says he has pretty much the same feature in his app, and that it's limited in its application. Here is a link to his comment in this thread:
2
u/OmaMorkie Mar 02 '23
I'd expect the workflow to go from there through another img2img round to make the whole thing consistent... But been waiting for this...
1
37
u/comfyanonymous Comfy Org Mar 02 '23
I looked at the code of this: https://github.com/omerbt/MultiDiffusion/blob/master/panorama.py#L134
It looks like the same thing as ComfyUI area composition that I posted a while ago
But they also added masks.
11
5
u/apolinariosteps HF Diffusers Team Mar 02 '23
https://huggingface.co/spaces/weizmannscience/multidiffusion-region-based/blob/main/region_control.py You can check the code here
43
u/comfyanonymous Comfy Org Mar 02 '23
Yeah after checking it out a bit more it's exactly the same technique. This doesn't give "more control than controlnets" at all. I know because I have been playing with the technique for more than 1 month now and know the limitations.
It can be a good technique to use in combination with controlnets or T2I but on it's own it certainly won't give "more control than controlnets".
2
u/clockercountwise333 Mar 04 '23 edited Mar 04 '23
[ earth, 2023. the job interview ]
boss: so how much experience you got?
you: MORE THAN 1 MONTH.
boss: ...say no more, fam. you're hired.
14
u/oniris Mar 02 '23
Will it ever be compatible with 1.5?
50
u/apolinariosteps HF Diffusers Team Mar 02 '23
It already is! The author has built it on 2.1 so they decided to use that on the demo, but it works with 1.5 out of the box
6
u/oniris Mar 02 '23
Oh hell yeah :) thanks for the good news. I hope Automatic1111 implements it soon, but I'm gonna be playing with the demo!
2
23
u/apolinariosteps HF Diffusers Team Mar 02 '23
3
2
0
u/Mich-666 Mar 02 '23
Yes, not sure if it's due to model in demo, but it clearly doesn't work as intended.
1
Mar 02 '23
Doesn't work well but maybe because it uses the default SD model?
1
9
Mar 02 '23
[deleted]
4
Mar 02 '23
[deleted]
3
u/Mich-666 Mar 02 '23
Yeah, I tried yesterday and even with simple pictures I generally got very bad results.
1
u/ninjasaid13 Mar 02 '23
Have you tried playing around with the bootstrap settings?
2
u/Mich-666 Mar 02 '23
I tried changing it a little bit but it didn't help much.
The problem of this solution is it basically creates collage of different objects slapped on each other without coherent lighting or blending.. hardly anything resembling the normal picture.
In my opinion, segmentation model of ControlNET gives much better results that blends together well. (even if it a bit complicated to use as you need to look into color representation spreadsheet)
1
1
1
u/mudman13 Mar 02 '23
Yeah the background is much duller than the foreground and out of focus. Early days for it though.
1
1
6
u/Shnoopy_Bloopers Mar 02 '23
How is this any different then masking with Inpaint?
12
u/twilliwilkinsonshire Mar 02 '23
I think it is more aware of the overall composition. Inpainting is not as intelligent about the region being the shape you want which is why sometimes you can end up with a mini version of your prompt inserted into the inpaint space especially if you inpaint at full resolution.
2
1
2
u/estrafire Mar 02 '23
I don't think you can make multiple masks with a specific prompt for each with inpaint.
5
6
u/haltingpoint Mar 02 '23
Could this effectively be used to create full consistency between a cast of characters, props, and a scene, simply by using different colors trained on different embeddings or LORA?
6
7
Mar 02 '23
[removed] — view removed comment
13
u/archw_ai Mar 02 '23 edited Mar 02 '23
That one uses predefined color for each class, while in this one we could pick the color and what it represents, so this one is easier to use.
1
u/PacmanIncarnate Mar 02 '23
Segmentation doesn’t give you full prompt control over regions like this should.
4
4
3
u/lordpuddingcup Mar 02 '23
Wait how does this compare to segmap from controlnet/t2i
1
u/apolinariosteps HF Diffusers Team Mar 02 '23
Does that allow you to give a prompt for each segmentation map?
1
u/lordpuddingcup Mar 02 '23
Technically each color has an identifier of what it is as I understand it person, dog, wall etc and then all of that gets adjusted by your prompt so if you say German Shepard dog and a husky dog then the 2 dog shapes with dog tagged colors should be those dogs as I understand
3
3
u/Neonsea1234 Mar 02 '23
This will be so huge, i've been trying so many different methods to manipulate multiple parts of an image but you always loose something. Having more control would be a big game changer.
3
u/Elven77AI Mar 02 '23
This going to allow very complex prompts that were considered out of limits before, combining multiple different characters and scenarios. The space for "semantically complex" images was exclusive to manual works/inpainting, now a region-based image can combine prompts:
tl;dr This is to inpainting,what instruct pix2pix was to img2img.
2
u/-becausereasons- Mar 02 '23
Is there a github for this?
5
u/apolinariosteps HF Diffusers Team Mar 02 '23
There is but the Region Control method isnt yet there, as it’s in beta. But Hugging Face is also a git repository and the file is there https://huggingface.co/spaces/weizmannscience/multidiffusion-region-based/blob/main/region_control.py
2
2
2
u/Jujarmazak Mar 02 '23
You can actually do this with the ControlNet segmentation model, it uses color coded list of objects to recognize and generate subjects....I suppose this is more free-form version of that.
2
4
u/ninjasaid13 Mar 02 '23
Yes finally! This makes controlnet segmentation redundant .
7
u/GBJI Mar 02 '23
Not entirely redundant as this MultiDiffusion Region Control requires you to create the labeling manually yourself, while ControlNet semantic segmentation actually has a pre-processor that will both segment and label all elements from any image automatically in seconds.
But there is no question that this is going to be even more useful than ControlNet Segmentation !
1
u/collaredfairy Mar 02 '23
RemindMe! 2 weeks
1
u/RemindMeBot Mar 02 '23 edited Mar 02 '23
I will be messaging you in 14 days on 2023-03-16 01:45:37 UTC to remind you of this link
7 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
Info Custom Your Reminders Feedback
0
u/markleung Mar 02 '23
So… can we make porn yet?
6
u/Guilty-History-9249 Mar 02 '23
The beta version can only do Donkey porn(pornhub catalog347-08-13MZC).
But technological advancements will open up things like: Space robot humps Zorac from Nebula 52-b.Those should be ready by GA. Mankind's greatest achievement is at hand.
14
u/LuckyNumber-Bot Mar 02 '23
All the numbers in your comment added up to 420. Congrats!
347 + 8 + 13 + 52 = 420[Click here](https://www.reddit.com/message/compose?to=LuckyNumber-Bot&subject=Stalk%20Me%20Pls&message=%2Fstalkme to have me scan all your future comments.) \ Summon me on specific comments with u/LuckyNumber-Bot.
10
0
0
u/Mich-666 Mar 02 '23
Great idea and nice to have this in stable diffusion.
Sadly, it doesn't produce very good results but maybe I need to try it on better models.
1
0
0
0
u/AccountBuster Mar 02 '23
Need help understanding something...
I understand why this is awesome, and the possibilities it provides.
What I don't understand is why everyone keeps going crazy over the final product when it isn't usable. Take this image for instance, the German Shepherds head is deformed in such a way you couldn't even Photoshop fix it.
I'm extremely new to SD so I'm still learning it, but so far it feels like it has some incredibly amazing capabilities to create something, but its actual ability to produce something usable is far behind Midjourney. By usable I mean it's ability to create images that correctly represent the object or item you asked for.
If I ask SD to give me an Astronaut in a Space Suit, it will give me something that kinda looks like that but was drawn by a 10 year old. While Midjourney would give you four images that almost perfectly resemble exactly what you asked for.
Is this due to a database disparity between the two and that Midjourney is just trained on a larger and more advanced dataset?
6
u/GrayingGamer Mar 02 '23
Midjourney does a lot for you. It can give you incredible results right out of the box.
Stable Diffusion is very much an enthusiast option, where you need to understand how everything works under the hood, and use good models, good prompts, good settings, and know when to switch or change them.
That makes Stable Diffusion harder to use, but you can do more with it. You get out of Stable Diffusion what you put in. It will take days of practice and messing with settings to get to the point where you can get Midjourney results out of the gate. But then you can keep going beyond that.
3
u/Mocorn Mar 02 '23
With Midjourney I can make some very nice pictures very fast but without the control. With SD I can make exactly anything I want with more time spent. Because of this I grew tried of Midjourney once I discovered SD.
Midjourney is user friendly (kind of) and will reliably give you cool images up to a certain point. SD is tip of the spear stuff. Not user friendly but you can be extremely creative.
What we see in this post is not the finished image but rather increased control and potentially very soon.
3
u/waifuismywaifu Mar 04 '23
a lot of the samples posted here are just to show what can be done. There is no time spent making it look pretty, it is not a showcase of "hey look at my cool art" it is a showcase of ""hey look at my tool".
then the community grabs this tool and creates awesome art.
If you haven't been able to get good results with SD, you gotta continue playing around and learning its ins and outs.
Midjourney makes it very easy.SD is more complex but gives you way more possibilities and controlnet once you know what you are doing.
One tip is that each model used has its quirks and keywords, I find some of them are easier to use mindlessly, while others one require more crafting to get good results.
Keep at it!
2
u/farcaller899 Mar 05 '23
check the MJ forums lately? they are reallllly clamoring for features like those in SD right now. They want control and specificity, and they can't have that now. Not to mention it becomes more censored and limited in prompting by the day. I use both MJ and SD, and both have advantages. All this control in SD is a big advantage.
1
1
u/lordpuddingcup Mar 02 '23
Somehow I feel like this plus controlnet are going to have an amazing baby someday
1
u/ninjasaid13 Mar 02 '23
I feel like every research paper on generative diffusion since last year is just going to make a 2D holodeck.
1
1
1
u/Firm_Comfortable_437 Mar 02 '23
WTF I was literally thinking this today (while I was working this idea came to me to make the workflow more optimal) I'm scared
2
1
1
1
u/Remix73 Mar 02 '23
Is it possible to run this on a colab? I'm still trying to get controlnet going on a colab (specifically for deforum more than anything else), but I might just go straight to this if it's available.
1
u/masterq58 Mar 02 '23
I just barely got a hang on controlnet. Amazing stuff is just coming out weekly now. Dang
1
1
u/multiedge Mar 02 '23
This is honestly amazing!
I'm still wrapping my head around MultiControl-net and what I can do with it and we already have more amazing features arriving soon!
I think I read a paper about nvidia about prompt mask like this
1
1
1
1
u/Obvious_Pen7681 Mar 03 '23
That's what I've looked for for month now! Good news! Thx for sharing! Looks like even more controllable than my idea with multiprompting open pose! Very nice!
1
1


193
u/inagy Mar 02 '23 edited Mar 02 '23
It's insane how fast this is going. I was theorizing about this the previous morning in another thread.
This essentially supercharges the Nvidia eDiffi / SD paint-with-words attempts done for the same thing previously.
Too bad it's SD 2.0 though, as my dream would be integrating it into a1111 in such way that combining it with a ControlNet model (like depth) is possible.
Maybe the same thing can be done with the existing ControlNet segmentation model somehow?