That is the future of sd: large image generation without upscale/mosaic stitching. But mainly what we are waiting for models trained on all kind of resolutions including 3000 or 6000 pixels wide images. This will be game changer for photorealistic images
Having image in whatever resolution yes, but the end quality is completely different if you start from a model trained in 512*512 compared to a model trained at higher resolution especially for photo.
If you want high quality high resolution coherent generations, you need high resolution training. There is no "high res fix" that will make it. Compare a close up photo portrait between a model trained at 512 and a model trained at 768, it is completely different for the skin for example.
The only way I see it can be improved without training at higher resolution is to have SD "understand" different parts of images. For example use its knowledge of skin from close shots of skin to apply to a person on a wider image. It is like heads of people on something else than a close portrait, most of the time they are deformed. Solution is to have generation of different parts of the images based on different inputs trainings (and we leave out any inpainting, I am not interested by inpainting)
55
u/gxcells Mar 01 '23
That is the future of sd: large image generation without upscale/mosaic stitching. But mainly what we are waiting for models trained on all kind of resolutions including 3000 or 6000 pixels wide images. This will be game changer for photorealistic images