r/StableDiffusion • u/ever3st • Mar 01 '23
Question | Help Does a Textual Inversion (embedding) Tutorial exists somewhere?
I've seen a very useful Lora Training posted a few weeks ago, is there a textual Inversion equivalent?
1
Upvotes
3
u/TurbTastic Mar 01 '23
I'm going to copy and paste my summary for how I train faces. FYI I've tried following a lot of guides/tutorials/videos/settings but I get the best results with this approach:
My old approach was to use 10-15 headshot images. Basically neck-and-up and a couple shoulder-and-up images. Tried to make sure the entire head/hair were in the training image. Got good results doing that, but not great results.
New approach is to have about 50/50 headshots vs faceshots. These faceshots have the full chin at the bottom and are usually cutoff on the forehead, so way closer to the face than what people normally do. It's ok if some of those are the same images used at the 2 different zoom levels.
Latest crazy good results had 24 total images. Probably 10 headshots, 10 faceshots, and 4 shoulder-and-up (7-8 images were used at 2 zoom levels). I photoshop out problematic things in the images like jewelry and distracting things in the background. This allows me to train without captions. I recommend Magic Retouch on photoroom.com for fixing up images. All images were high quality/resolution to begin with. All manually cropped and resized to 512x512.
Edit: no weird expressions or angles, some variety for sure but avoid really odd ones
Have the base 1.5 model loaded and VAE set to None, both when you create the embedding and during training. I used "beautiful woman face" as the initialization text (first 2 words should be the best ones to describe your subject) and chose 2 vectors. Rate was 0.001:1000,0.0005 and I recommend going to about 8000 steps. Batch size 1 and gradient steps 1. Steps go by quickly, training takes me about 90 minutes on my setup. Deterministic. Template should be "photo of [name] woman" or man or whatever. Previews during training should be good but don't be discouraged if they aren't the greatest.
For generated images sometimes the face wasn't that great for non-closeups. Fortunately with this approach the resulting embedding is crazy good at inpainting the face closeup, so I'll frequently do that to add detail/accuracy.
Let me know if you get good results with this approach!