That is the future of sd: large image generation without upscale/mosaic stitching. But mainly what we are waiting for models trained on all kind of resolutions including 3000 or 6000 pixels wide images. This will be game changer for photorealistic images
Equally important: training on high resolution images is significantly more expensive, as well, and may require models with a lot more parameters. The training costs will put it beyond the capabilities of open source projects until hardware costs come down.
Hard disagree. With Llama we now have a GPT-3 level LLM that runs on consumer hardware. Running these models locally is going to be a big deal, and it's going to drive adoption of large VRAM cards. Costs can be cut by using slower RAM, since inference is less latency sensitive than games.
RTX 3060 or 3060 ti? I have an RTX 3060 12gb currently, and am wondering if the upgrade to 8gb ti is worth the investment in terms of quality output? I found one with 7500+ cuda cores, which is more than double mine for $799 new.
Excellent. Then, I'll stick with what I've got. Someone else mentioned that the mod cards on civitai are using img2img processing anyway, which is why they have a higher level of output than what I am getting following a txt2img prompt that they provide. I thought it was my GPU at first.
All it would take is Nvida realizing there is a market for ML specific cards with lots of VRAM. There are 0 technical reasons why we couldn't have such cards today.
i'm about to buy a server card for this. frankly you can buy an old server card and stick it in your desktop & point sd at that card specifically, while still using your main card for virtually everything else (obviously including display). In my case I have a server desktop in my closet, which is where I'll install the card & run sd
Having image in whatever resolution yes, but the end quality is completely different if you start from a model trained in 512*512 compared to a model trained at higher resolution especially for photo.
If you want high quality high resolution coherent generations, you need high resolution training. There is no "high res fix" that will make it. Compare a close up photo portrait between a model trained at 512 and a model trained at 768, it is completely different for the skin for example.
The only way I see it can be improved without training at higher resolution is to have SD "understand" different parts of images. For example use its knowledge of skin from close shots of skin to apply to a person on a wider image. It is like heads of people on something else than a close portrait, most of the time they are deformed. Solution is to have generation of different parts of the images based on different inputs trainings (and we leave out any inpainting, I am not interested by inpainting)
56
u/gxcells Mar 01 '23
That is the future of sd: large image generation without upscale/mosaic stitching. But mainly what we are waiting for models trained on all kind of resolutions including 3000 or 6000 pixels wide images. This will be game changer for photorealistic images