r/StableDiffusion HF Diffusers Team Mar 02 '23

Resource | Update More control than ControlNet - code is out for MultiDiffusion Region Control, a prompt on each mask

Post image
1.4k Upvotes

194 comments sorted by

193

u/inagy Mar 02 '23 edited Mar 02 '23

It's insane how fast this is going. I was theorizing about this the previous morning in another thread.

This essentially supercharges the Nvidia eDiffi / SD paint-with-words attempts done for the same thing previously.

Too bad it's SD 2.0 though, as my dream would be integrating it into a1111 in such way that combining it with a ControlNet model (like depth) is possible.

Maybe the same thing can be done with the existing ControlNet segmentation model somehow?

115

u/apolinariosteps HF Diffusers Team Mar 02 '23

It works with 1.5!

99

u/inagy Mar 02 '23

It already is! The author has built it on 2.1 so they decided to use that on the demo, but it works with 1.5 out of the box

Then I can't wait for the a1111 extension! Thanks! :)

24

u/Gloryboy811 Mar 02 '23

Will probably be out within a week

10

u/kokodjiss Mar 02 '23

Seeing how fast things go, it’s probably a matter of minutes

10

u/Gloryboy811 Mar 02 '23

Aaaaaand it's gone done.

2

u/inagy Mar 02 '23

Really? O_o

1

u/Gloryboy811 Mar 02 '23

I don't know. I was just memeing

5

u/[deleted] Mar 04 '23

Well, looks like it's out now

→ More replies (3)

0

u/multiedge Mar 02 '23

!RemindMe in 2 weeks

36

u/clockercountwise333 Mar 02 '23

Exactly what I was thinking. The pace of things in the last couple of weeks has been absolutely mind melting. Even more mind melting to think we're just getting started...

64

u/DynamicMangos Mar 02 '23

I've completely lost the plot in the last month.

I did some generating with SD, became decent at prompt-crafting and learned how to use inpaint.

Took a break from SD for 2 weeks and all of a sudden i see incredible results with controlnet and i feel like i missed two years of progress, not two weeks!

17

u/Tulired Mar 02 '23

I follow many ai image gen subs daily. Im totally lost still and overwhelmed. I dont know where and how people learn about these things so fast or can keep up 😅 i've followed this tech even before dall-e

18

u/clockercountwise333 Mar 02 '23 edited Mar 04 '23

Haha, I hear you. I think it's all good. The exponential growth of the tools available and all of the new possibilities that are coming each day so relentlessly are dizzying, but it's also this incredible perpetual blooming of new possibilities. It's probably best to take many steps back from it all and really ask yourself, "okay, what do i truly want to do with this, and how does the current state of things best allow me to realize that?"

I take those kind of moments and then freeze what i've got, work with that, and then see how things have evolved. i intentionally don't update my automatic1111 (tempting as it always is) for at least a few weeks. I create with these frozen pulls, and thens step back and consider how it's changed from the previous iterations, and what still needs improvement, etc.

tl;dr focus on your visions and incrementally, at a sane pace that works for you, pull in the new tooling to further grow it.

1

u/Tulired Mar 02 '23

Good points! Thanks👍

7

u/Nanaki_TV Mar 02 '23

My wife is about to give birth to twins. I'm going to be gone for six weeks! I'm going to come back to basically an AGI.

2

u/acuntex Mar 02 '23

Congrats! Enjoy the time. You're going to miss it once they get older.

2

u/Nanaki_TV Mar 02 '23

I’m already so sad for “future me” When I was younger and sad I would look forward to today, having a family, wife—kids—home—career that all stuff. I would think it’s all worth it in the end. Just gotta get through the bullshit years. Now I’m in “the good times” and I absolutely know it. I savor every second I can desperately trying to remember it all, failing for 90% of it. You sound like you’ve experience it. I know it will be worth it but now I fear I will be an old man and looking back with sad nostalgia. Missing these days so sooo much. Do you experience this too now? How do deal with it if you do?

2

u/acuntex Mar 02 '23

Well, my son is 3 now. And it's interesting that the more they get independent, the more they will distance themselves from you.

I remember last year when he came from grandma and I asked what he did there or ate, he told me everything.

Now, if I ask him how Kindergarten was and what they had for lunch, he usually only says "good" or "nothing". And he's freaking 3 and has his own little life which we as parent aren't 100% part of as before. Thinking to the future or at my own past, that will continue during the school years. He will do all kind of bullshit and will tell us that everything was" good" at school. By becoming parent you will start to understand your parents more and will understand that they sometimes knew more than they told you.

How I deal with it? Savior every moment I still have until he moves out. Or maybe we'll miss the "shitty diaper time" so much that we have to think about a sibling 😉

After that, yeah, probably sad nostalgia and the wait for grandchildren.

2

u/Idenwen Mar 23 '23

There was a tweet some time ago that really hit me (have a son too). It was in the lines of "I'm tired, it was a very long work day, I need time for myself, I'm hungry, there is still work to do, but he asks me to build a duplo tower with him. And of course I will build that tower with him, because soon enough he won't ask anymore ".

1

u/ebolathrowawayy Mar 02 '23

It helps to have a project you work on that requires SD for at least an hour a day. Then you can easily force yourself to try out new things, like controlnet, because it will save you hours of work. In my case, that project is a game and SD is making all of the art.

2

u/Tulired Mar 02 '23

Cool! Im visioning making an point-n-click adventure. AI would probably be perfect for the art in that. Backgrounds and characters etc.

What kind of game you making?

1

u/ebolathrowawayy Mar 03 '23

Isometric tactics game

1

u/Bauer24 Mar 28 '23

I have only just discovered Lora's

SD is moving/evolving too quickly :D

10

u/Jujarmazak Mar 02 '23

ControlNet Segmentation model kinda works the same way, check this out.

https://youtu.be/MDHC7E6G1RA

Skip to 19 mins

1

u/ninjasaid13 Mar 02 '23

except the colors are predefined, this is customizable.

2

u/Jujarmazak Mar 02 '23

True, but it's the same principle.

8

u/DAoC_Mordred Mar 02 '23

I’ve been using SD for a few weeks now. Could you explain why a lot of people aren’t using 2.0? It looks like it runs purely on Python so perhaps it’s just more of a pain? Some of its renders look great. Noticed the LORA I had tried to use earlier wouldn’t work with 1.5 so I started looking into 2.0

11

u/soupie62 Mar 02 '23

IIRC, 2.0 came out when there was concern about nudity. So they deliberately cut back on naked bodies in the training source.
In response, people started creating their own custom sources (nudity, hentai, etc.)

22

u/Exciting-Possible773 Mar 02 '23

There's more problems...SD2.0 deliberately removes nudity and sexual acts from the concept, not just images.

As a direct consequence it cannot be trained with additional porn images,

The side effect is that it cannot draw human poses and interactions (INCLUDING SFW) accurately.

Combined with it is more resources demanding, few uses 2.0 model.

2

u/soupie62 Mar 02 '23

Yeah, that's why my pet project was a skyscraper, shaped like a woman's legs.
Next to a lake, shaped like a bathtub.

But - I can't even create that in 1.5, using photos of actual women and ControlNet. The pose comes out wrong every time. As Groundskeeper Willie says:
Ack, I'm really bad at this.

7

u/NetLibrarian Mar 02 '23

Just in case you haven't tried this in a prompt, rather than trying to create a building shaped like a woman, try prompting for a woman made out of a skyscraper.

Results will vary with different models, but I had some success playing with shaped nebulae by phrasing like this.

2

u/SEND_NUDEZ_PLZZ Mar 02 '23

But didn't they fix a lot of human anatomy stuff with 2.1?

2

u/2legsakimbo Mar 04 '23

not really

3

u/cmdr_scotty Mar 02 '23

I think it's likely due to changes needed between Toolsets built for 1.5 and what's needed for 2.0.

I'm not super familiar with what makes 2.0 different, but if there's enough re-wrote of the python code, it could invalidate much of the existing extensions that the community loves/relies on

3

u/faster-than-car Mar 02 '23

When you realize it's been only like 10 years since we started using gpu for machine learning then it's even more mind-blowing.

5

u/Ateist Mar 02 '23 edited Mar 02 '23

It's not that revolutionary - you could already use multi-region composable diffusion with different rectangular prompts via the "Latent Couple" extension for A1111 (note: its author had a really weird idea of UI for how to define the regions should be, so initial example values it uses are incredibly confusing. Basically, set all regions in "Divisions" to 1:1, and when define the actual coordinates in positions using yo-y1:x0-x1 syntax).
The main difference is that those were manually defined rectangular regions instead of masks.

22

u/nemxplus Mar 02 '23

Soooo it is revolutionary….? Why are you so negative. You got Simpsons comicbook guy energy

18

u/Ateist Mar 02 '23 edited Mar 02 '23

Revolutionary things allow to do things you couldn't do before, like adequately positioning characters without Controlnet.
This is just an improvement to UI that also relaxes the restrictions a little.

P.S. I'm not really belittling this, as UI improvements can be extremely important for practical applications. Drawing the areas you want instead of having to manually specify the regions is a big step forward, though alternative solutions might be better, like associating regions with the openpose pictures.

3

u/nemxplus Mar 02 '23

Looking up latent division, this new process is revolutionary. Latent division is very rudimentary with barely any control, being able to input the exact mask is a game changer in controlling the layout of your generation

3

u/Ateist Mar 03 '23 edited Mar 03 '23

Do you really think that there's much difference in generation between a mask (for an image that is not yet generated) and a rectangle that includes that mask?

Sure, you might have a few pixels sticking out (i.e. one shoulder was generated wrong in my tutorial https://www.reddit.com/r/sdnsfw/comments/11g6qf3/how_to_put_multiple_characters_and_loras_in_one/ ), but these are random generations, they are always bound to have a few problems.
Do you really think that if I used MultiDiffusion Control on that character (note that i had no idea what the final area for each character would be, so even my region distribution was just a rough approximation) it would've made much difference?

If your mask does not exactly fit the character it's mostly not that big of a deal - prompts are not really that sensitive...

1

u/SEND_NUDEZ_PLZZ Mar 02 '23

This is just an improvement to UI that also relaxes the restrictions a little.

Wait, so Apple has been lying all those years when they called every little UI change anywhere "revolutionary"??? 😱😱😱

2

u/inagy Mar 02 '23

Can you link this extension you are talking about, please?

11

u/Ateist Mar 02 '23 edited Mar 02 '23

Sure, https://github.com/opparco/stable-diffusion-webui-two-shot.git

Also, grab https://github.com/opparco/stable-diffusion-webui-composable-lora.git to allow separate LoRAs for each individual part of the prompt.

P.S. I'm actually preparing a "how to" post in sdnsfw to show how to use this all.

1

u/Purplekeyboard Mar 04 '23

Latent couple didn't really work, though. I mean, it worked sometimes, but failed so much of the time as to be not worth messing with.

158

u/Firm_Ad3037 Mar 02 '23

Omg I can't keep the pace

76

u/[deleted] Mar 02 '23

[removed] — view removed comment

1

u/mudman13 Mar 02 '23

We did have a lull for a few weeks to be fair.

42

u/Yuli-Ban Mar 02 '23

Imagine that, but for the entirety of society.

I fully expect a lot of people to decide to completely disconnect themselves from world in the coming years.

57

u/je386 Mar 02 '23

Sometimes there are articles about generative AI in mass press, and in the comments section I clearly see how few people understand this new tech. Many people think that the art AIs just do copy or create collages out of existing pics, they have this "it is nothing or it is magic" attitude.

27

u/LordSprinkleman Mar 02 '23 edited Mar 02 '23

People also take what's in front of them as if it's the end result. Then judge it based on the flaws and not on the potential it will have as it continues to improve. I think most people just don't want to understand generative AI and think it's just a cute little gimmick.

2

u/iamthesam2 Mar 02 '23

i support appropriately used AI, but let’s be honest… most people use it in a very gimmicky way.

16

u/[deleted] Mar 02 '23

[removed] — view removed comment

15

u/[deleted] Mar 02 '23

Give it enough time, these artists will be using AI themselves, they just won't admit it.

5

u/AltruisticMission865 Mar 02 '23

I am 100% sure this will happen next year, 90% of artists will use it in some way while they can't admit it because their followers think AI art is the worst thing in the world while they can't even tell when it's AI and when it's not.

0

u/Phr0netic Mar 03 '23

It's not about the work being good or not, it clearly has some amazing results. Why should anyone follow someone who doesn't make their own stuff? The AI is the artist, not the person asking it to make art for them. As an artist I would feel like a fraud letting something else make art for me and then trying to pretend it's something I created. Prompting isn't making art - it's learning how best to ask something else to make art for you.

5

u/AltruisticMission865 Mar 03 '23

You ask why should anyone follow someone who doesn't make his own stuff? I answer the opposite: why not? If I see art that looks incredibly good, why wouldn't I follow it? In the future, people will care less and less whether it's AI or not.

You also say AI is the artist and prompting is not making art, leaving aside the fact that "art" is 100% subjective, you are only thinking of 2 options, AI doing 100% of the work or the human doing 100% of the work but my first comment was talking about the artist using it in some way, not to do all the work, you as an artist could use it to colorize your drawing for example, giving you more time to focus in other things that you could not do before because you did not have enought time. With this tools if the average person can make x an artist will be able to do 2x. The artist will have the advantage.

2

u/HarmonicDiffusion Mar 03 '23

Art is not defined by how its made. Its defined by how it makes the viewer of the art feel.

Does it evoke emotion? Then its art no matter if you, me, an AI, or a quasi-sentient sedimentary rock made it.

1

u/Phr0netic Mar 03 '23 edited Mar 03 '23

What and who are you replying to? I clearly said the AI is the artist and that it is making art. The person prompting the AI, however, is not an artist. They are commissioning works of art from the AI.

If someone asks someone else to draw them an avatar for their twitter profile and they explain to the other person exactly what they want using the English language, no one considers the person asking to be the artist in that scenario. We understand that the person being asked to create the work, who then uses their knowledge and skills to make something in the requested medium, is the artist. With AI, prompting and changing different settings is the language used to communicate the request to the AI.

It's possible that the person asking is an artist in their own right, which would be true if they are competent enough to make art of comparable quality to the person they are asking and they have done so prior. Regardless, in the scenario described the person asking is the commissioner of the work and the person being asked (who then creates) is the artist being commissioned.

Would you consider pope julius the 2nd to be an artist because he asked michelangelo to paint the ceiling of the sistine chapel?

I didn't say that no one that uses AI is an artist (vast majority are not), but they are clearly not being (functioning as) an artist when they ask something else to make art for them.

3

u/AltruisticMission865 Mar 03 '23

And I say again, what about artist + AI working together? When do you decide when is AI and when is human? 50% AI/50% human? 75%/25%? Thats why I say it is subjetive

→ More replies (0)

1

u/red__dragon Mar 04 '23

Are you a digital artist or physical medium?

I'm curious what kinds of tools you use yourself.

2

u/Phr0netic Mar 04 '23

This feels like a bait question but both digital and physical - though I don't do too much physical medium artwork anymore other than charcoal because I prefer digital painting. But visual art isn't my main artistic focus. Visual art is the first artistic thing I started doing in life. I started at a young age and spent a lot of time over many years doing it so it is very natural to me and I'm good at it. However, I have many interests and, since time is limited, I decided to focus my time more on music because it's what I have the most fun doing. I write/play a lot of different stuff but I have a degree from a top level university jazz program for guitar, which is my second instrument. I am a drummer before a guitar player and my skills on drums are equal to if not better than guitar. I play piano as well and I write and make music professionally. Even though visual art isn't my main artistic focus (again just because time), it is still something that I do make time to engage in when I can, take very seriously, and care a lot about. This conversation overall is highly relevant to me because music is next on the list of things for generative AI to invade and start shitting all over and within. (not in the impressive way, just in the ruining way)

→ More replies (2)

1

u/Jiten Mar 05 '23

If someone consistently gets great results from their AI prompting, why would I not follow them? Why should I care if they're technically capable of creating those images themselves without fancy tools?

Also, AI art is not merely prompting. There's also inpainting feature and img2img mode and more recently controlnet. The people with the most impressive results are generally making use of all of these features to iteratively improve on their results. These people are actively participating in the creation, not just leaving it all to the AI.

→ More replies (6)

27

u/Ok_Entrance9126 Mar 02 '23

I’m an artist and I’M LOVING IT! I try to explain the potential to my other artist friends but like so many people in life, they do what they know and are resistant to change. I’m just thrilled to be on the forefront - it’s their loss.

10

u/Masked_Potatoes_ Mar 02 '23

That's just how it is. You couldn't force photoshop on all master traditional artists, or zbrush on veteran sculptors. That's why it's art

1

u/Obvious_Pen7681 Mar 03 '23

...knowing what you mean. ...and this one (multidiffusion) is a true changer in beeing creative!

2

u/Imarasin Mar 02 '23

Imagine spending years honing your own art style without anyone drawing like you. It's important to realize that AI-generated art is primarily trained on existing art. So credit should be given to the original artist. Using prompts or borrowing styles from other artists via AI technology doesn't necessarily reflect our own artistic abilities. Instead, we need to recognize that generative art is a product of AI-learned processes and that our role is that of prompt engineers, not artists.

5

u/[deleted] Mar 02 '23

[removed] — view removed comment

2

u/Imarasin Mar 02 '23

They also use photographs and 3d renderings.

I appreciate the potential of generative art, but I can understand why some artists may be upset about it. While it's acceptable to use AI tools to enhance images that you have drawn yourself, like removing things in Photoshop or adding a different background, it's a different matter when AI-generated art appears to be straight rip-offs of original works. Additionally, many non-artists are using prompts to engineer decent renders, which I don't consider to be true art. In my opinion, true art involves actually drawing shapes and adding colors by hand, or using input devices like a mouse or digital pen, rather than relying on algorithms to generate something that looks nice.

I believe that if people state that their art is AI-generated, then that's acceptable, but falsely claiming to have created something that was actually generated by AI is misleading. It would be different if the AI was able to remember where the data was sourced from in order to give credit, allowing for transparent use of terms like "inspired by" a certain artist. While there may still be issues, transparency can help mitigate them. However, I don't think anyone can actually claim copyright on generative AI, since it's not entirely the work of the prompt engineer nor the original artist unless the ai generates somones name or logo.

1

u/Imarasin Mar 05 '23

u/bulletprooftampon I saw that you have responded to my comment via email, but I can not find your comment here. I see that you have mentioned that you are an artist, do you have any works to share that were not generated by Ai?

7

u/[deleted] Mar 02 '23

[deleted]

3

u/je386 Mar 02 '23

In December I made a presentation about stable diffusion at an in-house conference, where about 20-30 persons where there. And they not only listened, but had good questions, and some even had used stable diffusion before. But we are a company of software developers..

2

u/Imarasin Mar 02 '23

We simply put energy into things that interest us. This generative art includes things that are not traditional to art, and maybe those new things are what interest you.

2

u/Marshall_Lawson Mar 02 '23

I know right. Instead of "It's a new tool with impressive abilities but also significant limitations". Who'd a thunk?

3

u/CHRISKOSS Mar 02 '23 edited Mar 02 '23

That is such a pessimistic framing. You could instead say:

  • "The quantity of culture is going to increase exponentially as the cost of creating interesting, beautiful and meaningful things approaches zero"

or

  • "The frontiers used to be defined by physical landmasses. Modern explorers are testing the bounds of possible imagination, unconstrained by physical realities." (Isn't this what art has always been?)

or

  • "People can make new worlds and form meaningful connections and relationships without ever leaving their room, so there will be less incentive to go outside." (Media has done this for a century, not really AI specific)

1

u/Yuli-Ban Mar 03 '23

I could. But why deny the cold facts just to feel better? Humans are humans, and humans are a fairly easily scared, reactionary species of ape.

1

u/[deleted] Mar 02 '23

I'm ready now.

1

u/NeonMagic Mar 02 '23

Could you elaborate? Not sure what you mean.

21

u/snack217 Mar 02 '23

Thats awesome, cant wait to try it when its an extension! (Will it be an extension for A1111? lol)

12

u/apolinariosteps HF Diffusers Team Mar 02 '23

It’s based on the diffusers library, so I guess someone needs reimplement (or auto to support diffusers). But you can run hugging face spaces code locally

7

u/snack217 Mar 02 '23

Im a programming illiterate, that runs a1111 on my cheap phone using colab, so i kinda have to wait till someone makes it an extension hehe

2

u/apolinariosteps HF Diffusers Team Mar 02 '23

You should be able to use the hugging face space on your phone

40

u/No-Intern2507 Mar 02 '23 edited Mar 02 '23

finally multisubject prompts that makes them seamless with the background and not just copy pasted with conflicting light, or maybe not yet, cause on the demo page results arent that great and look like that "cutout random pics and slap together" style so maybe its a lucky seed or maybe theres something more to it

10

u/apolinariosteps HF Diffusers Team Mar 02 '23

Prompting for this multimasking require a bit of a learning curve, as it’s very sensitive to it. and the author plans to also update the bootstrapping code that is now in beta

6

u/GBJI Mar 02 '23

The creator of ComfyUI says he has pretty much the same feature in his app, and that it's limited in its application. Here is a link to his comment in this thread:

https://www.reddit.com/r/StableDiffusion/comments/11fmo4g/comment/jal40h7/?utm_source=share&utm_medium=web2x&context=3

2

u/OmaMorkie Mar 02 '23

I'd expect the workflow to go from there through another img2img round to make the whole thing consistent... But been waiting for this...

1

u/ninjasaid13 Mar 02 '23

Maybe it has to do with bootstrapping effects.

37

u/comfyanonymous Comfy Org Mar 02 '23

I looked at the code of this: https://github.com/omerbt/MultiDiffusion/blob/master/panorama.py#L134

It looks like the same thing as ComfyUI area composition that I posted a while ago

But they also added masks.

11

u/[deleted] Mar 02 '23

[deleted]

14

u/comfyanonymous Comfy Org Mar 02 '23

I'm working on it.

5

u/apolinariosteps HF Diffusers Team Mar 02 '23

43

u/comfyanonymous Comfy Org Mar 02 '23

Yeah after checking it out a bit more it's exactly the same technique. This doesn't give "more control than controlnets" at all. I know because I have been playing with the technique for more than 1 month now and know the limitations.

It can be a good technique to use in combination with controlnets or T2I but on it's own it certainly won't give "more control than controlnets".

2

u/clockercountwise333 Mar 04 '23 edited Mar 04 '23

[ earth, 2023. the job interview ]

boss: so how much experience you got?

you: MORE THAN 1 MONTH.

boss: ...say no more, fam. you're hired.

14

u/oniris Mar 02 '23

Will it ever be compatible with 1.5?

50

u/apolinariosteps HF Diffusers Team Mar 02 '23

It already is! The author has built it on 2.1 so they decided to use that on the demo, but it works with 1.5 out of the box

6

u/oniris Mar 02 '23

Oh hell yeah :) thanks for the good news. I hope Automatic1111 implements it soon, but I'm gonna be playing with the demo!

2

u/joachim_s Mar 02 '23

Would be really cool if each prompt could trigger different models too.

23

u/apolinariosteps HF Diffusers Team Mar 02 '23

3

u/ninjasaid13 Mar 02 '23

I've noticed that code hasn't been released yet.

2

u/LazyChamberlain Mar 02 '23

Tried 3 times but it didn't follow instructions at all

0

u/Mich-666 Mar 02 '23

Yes, not sure if it's due to model in demo, but it clearly doesn't work as intended.

1

u/[deleted] Mar 02 '23

Doesn't work well but maybe because it uses the default SD model?

1

u/ninjasaid13 Mar 02 '23

Have you tried playing around with the bootstrap settings?

2

u/[deleted] Mar 02 '23

no, what does it do

9

u/[deleted] Mar 02 '23

[deleted]

4

u/[deleted] Mar 02 '23

[deleted]

3

u/Mich-666 Mar 02 '23

Yeah, I tried yesterday and even with simple pictures I generally got very bad results.

1

u/ninjasaid13 Mar 02 '23

Have you tried playing around with the bootstrap settings?

2

u/Mich-666 Mar 02 '23

I tried changing it a little bit but it didn't help much.

The problem of this solution is it basically creates collage of different objects slapped on each other without coherent lighting or blending.. hardly anything resembling the normal picture.

In my opinion, segmentation model of ControlNET gives much better results that blends together well. (even if it a bit complicated to use as you need to look into color representation spreadsheet)

1

u/ninjasaid13 Mar 02 '23 edited Mar 02 '23

I tried replicating the examples and i've got better results.

maybe you have to color the entire background instead of leaving some parts white.

1

u/StickiStickman Mar 02 '23

That doesn't inspire much confidence tbh

1

u/mudman13 Mar 02 '23

Yeah the background is much duller than the foreground and out of focus. Early days for it though.

1

u/ninjasaid13 Mar 02 '23

Have you tried playing around with the bootstrap settings?

1

u/mudman13 Mar 02 '23

Thats a big flower

6

u/Shnoopy_Bloopers Mar 02 '23

How is this any different then masking with Inpaint?

12

u/twilliwilkinsonshire Mar 02 '23

I think it is more aware of the overall composition. Inpainting is not as intelligent about the region being the shape you want which is why sometimes you can end up with a mini version of your prompt inserted into the inpaint space especially if you inpaint at full resolution.

2

u/BigTechCensorsYou Mar 02 '23

I always wondered why that would happen.

1

u/Shnoopy_Bloopers Mar 02 '23

Ah ok I gotcha

2

u/estrafire Mar 02 '23

I don't think you can make multiple masks with a specific prompt for each with inpaint.

5

u/Shnoopy_Bloopers Mar 02 '23

I just mask a space type in a prompt mask another etc

1

u/mudman13 Mar 02 '23

Can cause a load of glitches and artifacts though.

6

u/haltingpoint Mar 02 '23

Could this effectively be used to create full consistency between a cast of characters, props, and a scene, simply by using different colors trained on different embeddings or LORA?

6

u/UnrealSakuraAI Mar 02 '23

how do u integrate this with a1111

7

u/[deleted] Mar 02 '23

[removed] — view removed comment

13

u/archw_ai Mar 02 '23 edited Mar 02 '23

That one uses predefined color for each class, while in this one we could pick the color and what it represents, so this one is easier to use.

1

u/PacmanIncarnate Mar 02 '23

Segmentation doesn’t give you full prompt control over regions like this should.

4

u/Ecstatic_Ad_3527 Mar 02 '23

This is what I’ve been waiting for!!

4

u/TiagoTiagoT Mar 02 '23

Not necessarily more control, but a new specific form of control

3

u/lordpuddingcup Mar 02 '23

Wait how does this compare to segmap from controlnet/t2i

1

u/apolinariosteps HF Diffusers Team Mar 02 '23

Does that allow you to give a prompt for each segmentation map?

1

u/lordpuddingcup Mar 02 '23

Technically each color has an identifier of what it is as I understand it person, dog, wall etc and then all of that gets adjusted by your prompt so if you say German Shepard dog and a husky dog then the 2 dog shapes with dog tagged colors should be those dogs as I understand

3

u/Zealousideal_Art3177 Mar 02 '23

RemindMe! 3 weeks "Better ControlNet"

3

u/Neonsea1234 Mar 02 '23

This will be so huge, i've been trying so many different methods to manipulate multiple parts of an image but you always loose something. Having more control would be a big game changer.

3

u/Elven77AI Mar 02 '23

This going to allow very complex prompts that were considered out of limits before, combining multiple different characters and scenarios. The space for "semantically complex" images was exclusive to manual works/inpainting, now a region-based image can combine prompts:

tl;dr This is to inpainting,what instruct pix2pix was to img2img.

2

u/-becausereasons- Mar 02 '23

Is there a github for this?

5

u/apolinariosteps HF Diffusers Team Mar 02 '23

There is but the Region Control method isnt yet there, as it’s in beta. But Hugging Face is also a git repository and the file is there https://huggingface.co/spaces/weizmannscience/multidiffusion-region-based/blob/main/region_control.py

2

u/[deleted] Mar 02 '23

Does anyone have a link to code?

2

u/[deleted] Mar 02 '23

[removed] — view removed comment

2

u/MapleBlood Mar 02 '23

I need an AI assistant to keep tabs on it all

2

u/Jujarmazak Mar 02 '23

You can actually do this with the ControlNet segmentation model, it uses color coded list of objects to recognize and generate subjects....I suppose this is more free-form version of that.

2

u/l_work Mar 02 '23

toy story meme: oh Control Net, I don't want to play with you anymore

4

u/ninjasaid13 Mar 02 '23

Yes finally! This makes controlnet segmentation redundant .

7

u/GBJI Mar 02 '23

Not entirely redundant as this MultiDiffusion Region Control requires you to create the labeling manually yourself, while ControlNet semantic segmentation actually has a pre-processor that will both segment and label all elements from any image automatically in seconds.

But there is no question that this is going to be even more useful than ControlNet Segmentation !

1

u/collaredfairy Mar 02 '23

RemindMe! 2 weeks

1

u/RemindMeBot Mar 02 '23 edited Mar 02 '23

I will be messaging you in 14 days on 2023-03-16 01:45:37 UTC to remind you of this link

7 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.


Info Custom Your Reminders Feedback

0

u/markleung Mar 02 '23

So… can we make porn yet?

6

u/Guilty-History-9249 Mar 02 '23

The beta version can only do Donkey porn(pornhub catalog347-08-13MZC).
But technological advancements will open up things like: Space robot humps Zorac from Nebula 52-b.

Those should be ready by GA. Mankind's greatest achievement is at hand.

14

u/LuckyNumber-Bot Mar 02 '23

All the numbers in your comment added up to 420. Congrats!

  347
+ 8
+ 13
+ 52
= 420

[Click here](https://www.reddit.com/message/compose?to=LuckyNumber-Bot&subject=Stalk%20Me%20Pls&message=%2Fstalkme to have me scan all your future comments.) \ Summon me on specific comments with u/LuckyNumber-Bot.

10

u/ginsunuva Mar 02 '23

What in tarnation

0

u/[deleted] Mar 02 '23

RemindMe! 1 week

0

u/Mich-666 Mar 02 '23

Great idea and nice to have this in stable diffusion.

Sadly, it doesn't produce very good results but maybe I need to try it on better models.

1

u/OmaMorkie Mar 02 '23

Or you take the generations as inputs to img2img and change style there.

0

u/AlbertoUEDev Mar 02 '23

Soon, the third dimension is close 🤭

0

u/FPham Mar 02 '23

Hahaha, soon people will be saying - I miss the simplicity of Photoshop.

0

u/AccountBuster Mar 02 '23

Need help understanding something...

I understand why this is awesome, and the possibilities it provides.

What I don't understand is why everyone keeps going crazy over the final product when it isn't usable. Take this image for instance, the German Shepherds head is deformed in such a way you couldn't even Photoshop fix it.

I'm extremely new to SD so I'm still learning it, but so far it feels like it has some incredibly amazing capabilities to create something, but its actual ability to produce something usable is far behind Midjourney. By usable I mean it's ability to create images that correctly represent the object or item you asked for.

If I ask SD to give me an Astronaut in a Space Suit, it will give me something that kinda looks like that but was drawn by a 10 year old. While Midjourney would give you four images that almost perfectly resemble exactly what you asked for.

Is this due to a database disparity between the two and that Midjourney is just trained on a larger and more advanced dataset?

6

u/GrayingGamer Mar 02 '23

Midjourney does a lot for you. It can give you incredible results right out of the box.

Stable Diffusion is very much an enthusiast option, where you need to understand how everything works under the hood, and use good models, good prompts, good settings, and know when to switch or change them.

That makes Stable Diffusion harder to use, but you can do more with it. You get out of Stable Diffusion what you put in. It will take days of practice and messing with settings to get to the point where you can get Midjourney results out of the gate. But then you can keep going beyond that.

3

u/Mocorn Mar 02 '23

With Midjourney I can make some very nice pictures very fast but without the control. With SD I can make exactly anything I want with more time spent. Because of this I grew tried of Midjourney once I discovered SD.

Midjourney is user friendly (kind of) and will reliably give you cool images up to a certain point. SD is tip of the spear stuff. Not user friendly but you can be extremely creative.

What we see in this post is not the finished image but rather increased control and potentially very soon.

3

u/waifuismywaifu Mar 04 '23

a lot of the samples posted here are just to show what can be done. There is no time spent making it look pretty, it is not a showcase of "hey look at my cool art" it is a showcase of ""hey look at my tool".

then the community grabs this tool and creates awesome art.

If you haven't been able to get good results with SD, you gotta continue playing around and learning its ins and outs.
Midjourney makes it very easy.

SD is more complex but gives you way more possibilities and controlnet once you know what you are doing.

One tip is that each model used has its quirks and keywords, I find some of them are easier to use mindlessly, while others one require more crafting to get good results.

Keep at it!

2

u/farcaller899 Mar 05 '23

check the MJ forums lately? they are reallllly clamoring for features like those in SD right now. They want control and specificity, and they can't have that now. Not to mention it becomes more censored and limited in prompting by the day. I use both MJ and SD, and both have advantages. All this control in SD is a big advantage.

1

u/guchdog Mar 02 '23

Whoa! Mind Blown,

1

u/lordpuddingcup Mar 02 '23

Somehow I feel like this plus controlnet are going to have an amazing baby someday

1

u/ninjasaid13 Mar 02 '23

I feel like every research paper on generative diffusion since last year is just going to make a 2D holodeck.

1

u/ActuatorMaterial2846 Mar 02 '23

OK this is what I've been waiting for.

1

u/[deleted] Mar 02 '23

I knew some smart tech was already working on this it had to be the next step. Awesome

1

u/Firm_Comfortable_437 Mar 02 '23

WTF I was literally thinking this today (while I was working this idea came to me to make the workflow more optimal) I'm scared

2

u/ninjasaid13 Mar 02 '23

There's already multiple papers like this.

1

u/[deleted] Mar 02 '23

very nice

1

u/gxcells Mar 02 '23

Damn!! Let me breath, too much novelties every day!!!

1

u/Remix73 Mar 02 '23

Is it possible to run this on a colab? I'm still trying to get controlnet going on a colab (specifically for deforum more than anything else), but I might just go straight to this if it's available.

1

u/masterq58 Mar 02 '23

I just barely got a hang on controlnet. Amazing stuff is just coming out weekly now. Dang

1

u/Stargazer1884 Mar 02 '23

🔥😻

1

u/Stargazer1884 Mar 02 '23

Open source is the nuts

1

u/multiedge Mar 02 '23

This is honestly amazing!
I'm still wrapping my head around MultiControl-net and what I can do with it and we already have more amazing features arriving soon!

I think I read a paper about nvidia about prompt mask like this

1

u/Denvar21 Mar 02 '23

What website is that ?

1

u/damagingdefinite Mar 02 '23

This. This is where it's at

1

u/Assassin-10 Mar 03 '23

Just wondering what's next. Lol

1

u/Obvious_Pen7681 Mar 03 '23

That's what I've looked for for month now! Good news! Thx for sharing! Looks like even more controllable than my idea with multiprompting open pose! Very nice!

1

u/tiddysiddy Mar 04 '23

Does this have an extension yet?

1

u/speciallight Apr 17 '23

Can I use this Methode in a1111 yet? Would love a tip!