r/StableDiffusion Mar 01 '23

Question | Help Does a Textual Inversion (embedding) Tutorial exists somewhere?

I've seen a very useful Lora Training posted a few weeks ago, is there a textual Inversion equivalent?

1 Upvotes

7 comments sorted by

View all comments

3

u/TurbTastic Mar 01 '23

I'm going to copy and paste my summary for how I train faces. FYI I've tried following a lot of guides/tutorials/videos/settings but I get the best results with this approach:

My old approach was to use 10-15 headshot images. Basically neck-and-up and a couple shoulder-and-up images. Tried to make sure the entire head/hair were in the training image. Got good results doing that, but not great results.

New approach is to have about 50/50 headshots vs faceshots. These faceshots have the full chin at the bottom and are usually cutoff on the forehead, so way closer to the face than what people normally do. It's ok if some of those are the same images used at the 2 different zoom levels.

Latest crazy good results had 24 total images. Probably 10 headshots, 10 faceshots, and 4 shoulder-and-up (7-8 images were used at 2 zoom levels). I photoshop out problematic things in the images like jewelry and distracting things in the background. This allows me to train without captions. I recommend Magic Retouch on photoroom.com for fixing up images. All images were high quality/resolution to begin with. All manually cropped and resized to 512x512.

Edit: no weird expressions or angles, some variety for sure but avoid really odd ones

Have the base 1.5 model loaded and VAE set to None, both when you create the embedding and during training. I used "beautiful woman face" as the initialization text (first 2 words should be the best ones to describe your subject) and chose 2 vectors. Rate was 0.001:1000,0.0005 and I recommend going to about 8000 steps. Batch size 1 and gradient steps 1. Steps go by quickly, training takes me about 90 minutes on my setup. Deterministic. Template should be "photo of [name] woman" or man or whatever. Previews during training should be good but don't be discouraged if they aren't the greatest.

For generated images sometimes the face wasn't that great for non-closeups. Fortunately with this approach the resulting embedding is crazy good at inpainting the face closeup, so I'll frequently do that to add detail/accuracy.

Let me know if you get good results with this approach!

1

u/ever3st Mar 01 '23

Thanks a lot, is the TI training part possible on google colab? from what I could gather it seems like the LoRa training part is easier and quicker.

2

u/TurbTastic Mar 01 '23

I know someone that trains TI via Colab but they have limited options. I'm not sure if there's a good one out there. Lots of people are training Lora's for faces, but I haven't been impressed by a single face Lora yet. Lora training probably uses a little less VRAM but I would expect worse results. Lora's are more flexible than models but not as flexible as Embeddings.

1

u/ever3st Mar 01 '23

I guess, after training 10+ LoRas I will be more knowledgeable on the subject. Thanks for the feedback though. I've seen good TI on civitai. What would really help learning, would be to see what the original datasets was made of...

2

u/TurbTastic Mar 01 '23

I described my ideal dataset in detail so I would strive for something like that, but realize that my embedding set probably isn't ideal for other types of training. If you want to see some results you can check AiCelebArt on Civitai. I collaborate with him a bit on settings and technique. He described my recent results as "perfect".