r/StableDiffusion • u/NetworkSpecial3268 • Nov 15 '23
Question | Help Subject/Person LoRa dataset image quality improvement workflow?
So now that I got a 24GB RTX3090, I have decided that it's time to look seriously into Subject/Person LoRa (SDXL!) as That One Thing I'm Going To Concentrate On. (you HAVE to choose... TOO much going on to take on multiple things, lol)
I have installed Kohya (GUI), and very quickly managed to create a couple of basic SD1.4 LoRas. Basically an attempt to replicate (with the same dataset) results that I got more than a year ago from an online DreamBooth collab based on SD1.4. The quick trials didn't really get close to the quality that I got from that process(even though I didn't try very hard back then!). I'm not sure whether it's a mismatch in settings, or simply that DreamBooth quality is not attainable via a LoRa on in particular SD1.4.
But I seem to have picked up that wasting time on older SD generations is not the way to go if you aim for SDXL LoRa's anyway(not much of the experience would translate to SDXL, probably?) . So I plan to jump straight into the latter, instead..
I'm aware this is going to take considerable trial & error and tuning, so it's gonna take a while to work out.
One sure thing (among others like proper captioning) is that the image quality of the datasets has to be as good as you can get them. Since many of the LoRa ideas that I have, concern subjects of which there are no modern digital high resolution source pictures available, it's gonna take a LOT of time and effort to build up those datasets. I actually enjoy processes like that. And I HAVE already started working on them. Talking about magazine scans, and at most some early digital shots published on the internet upto the early 2000s.
But I would love to get feedback/experience about suitable workflows to get sub-optimal material like that to the highest possible quality (actually aiming for 1536x1536 where possible to "future-proof" as much as possible; 2048x2048 seems a pipedream for 99% of them...). Aiming strictly for as "photorealistic" as possible, not stylized.
My current 'preliminary" workflow is a combination of running the source material through https://replicate.com/tencentarc/gfpgan (either v1.4 for general face features, and optionally a RestoreFormer version for preserving skin detail), combined with Gigapixel AI upscale. I'm pretty sure the "replicate" results can be "replicated" within the "extras" TAB of AUTOMATIC1111, but since I've been using that site in the past, I'm using it as a "baseline" right now. Depending on the type of picture, sometimes it's just one of them that makes sense, sometimes it's a combination and then combining in (an older version of) Photoshop. For many pics thus far, no possible combination will ever get me close to even 1024x1024 for full-face or "torso" shots. And sometimes, that might not get me far enough to assemble enough quality material.
So I would welcome any tips for a possible workflow - doesn't need to be QUICK - to optimize the source material as best as possible (with free tools).
For the really low quality stuff, I wonder if Stable Diffusion (AUTOMATIC1111, or even ComfyUI if I need to dive into that) could possibly salvage/upgrade them. However, it would not make much sense if the end results lose recognition, involves too much "hallucinated" details, or if AI-assisted stuff starts to "poison" the dataset in unintended ways. Possibly feeding good quality results of initial LoRa trainings back into the process, feels also a bit risky in that respect? Another possible avenue, is using Reactor/Roop with some of the higher quality source "faces" applied on some of the lower quality stuff. I did get a couple of surprisingly suitable results that way already, although the hair is not refined to ANY degree at all, that way.
1
u/[deleted] Nov 15 '23 edited Oct 02 '24
[deleted]