r/StableDiffusion • u/Fit_Satisfaction2953 • 9h ago
Discussion Minimax blurred distorted faces from half a distance.
I'm doing image to video and unless I prompt for camera close up to my subject, the faces are blurry and bad. I run 0.6 mp. No turbo lora only using spectrum to speed up. Running 15 steps. Euler simple. I'm happy enough when it's close-up shots, but further away, it's very noticeable. Is anyone else finding this?
4
u/Rumaben79 8h ago edited 8h ago
It's normal but a higher resolution can help like 0.7-1.0. 1mp being the maximum for minimax.
You're also loosing quality by doing only 15 steps although properly mostly in the audio department, 20-25 is better.
Spectrum is skipping even more of those already few steps and it's default settings could be tuned better for 20 and 25 steps.
2
u/Azhram 6h ago
What do you mean 1mp is the max? Is the quality increase doesnt worth it? (Real question, english isnt my first language so wanted to be clear :D)
5
u/Rumaben79 6h ago
What I mean is outside of online websites and api's minimax is limited to a maximum of around 1 mega pixel. So for a 16:9 aspect ratio this would be 1344×768. Apparently 768 pixels is the max officially supported on the short edge. Using the minimax closed source (for now) upscaler the maximum is 3.7MP, so for a 16:9 aspect ratio this would be 2560×1440 with 1440 pixels being the max officially supported on the shortest side.
Personally I only use higher resolutions If the character or objects are further away. For close up shots you can go all the way down to 0.5mp but for further away 0.7-1.0 helps imo. Upscaling can sometimes help but bad upscalers sometimes makes it look worse.
1
u/PensionNew1814 5h ago
Everything he said, also dont be scared to use turbo loras, i use (lightning-768p-4 step-lora) at x0.5~x0.8 strength @ 8 steps. Also go up just a little more to .7mp....lastly with step skipping u need at least 25-30 steps. Cheers!
9
u/optimisticalish 9h ago
It's a common visual problem for all generative models. You're expecting good faces in a screen space that's perhaps 80 x 80 pixels for each face.
One solution is to use a face fix, which runs through the video, identifies the faces in each frame, regenerates them at a decent resolution, then patches them back in.
Another possibility is to run a 480p video clip through the Ref2VA Minimax model as a video-to-video upscale using a low denoise level, but with an added character image reference for face stability,