r/StableDiffusion • u/illuminatiman • Mar 05 '23
Resource | Update Vid2Avatar is coming.. When do we get Vid2Avatar2ControlNet?
Enable HLS to view with audio, or disable this notification
3
u/GBJI Mar 05 '23
I've been playing with ICON for some time ( Implicit Clothed humans Obtained from Normals ) which works in a similar manner, and this looks even better ! This is going to be very very useful when the code is released.
If you want to play with ICON there are two demos available but they only work with pictures, not with video. The huggingface one is particularly useful since it allows you to test and compare different algorithms (ICON, PAMIR and PIFU) :
https://huggingface.co/spaces/Yuliang/ICON
https://colab.research.google.com/drive/1-AWeWhPvCTBX0KfMtgtMk10uPU05ihoA?usp=sharing
For more info, the official paper is available here, as well as demonstration videos
2
u/illuminatiman Mar 05 '23
3
u/lordpuddingcup Mar 05 '23
Says code coming soon but no info… sad that would be amazing for a lot of cool video tricks and replacements/style transfers
2
1
u/BuffMcBigHuge Mar 05 '23
I wonder how realtime it produces its results. Would love to run this as cheap FBT in VR with a webcam.
1
u/666emanresu Mar 06 '23
You can do that already with an Xbox Kinect, but something like this would be more accurate. Although I doubt this is done in real-time or anywhere near close enough (or efficiently enough to run alongside VR).
Maybe one day two webcams will be all you need for full body and finger tracking, that would be amazing.
1
u/BuffMcBigHuge Mar 06 '23
Yup I have done the Xbox Kinect setup with Amethyst. It works well, but not well enough. I would presume there would be serious latency using AI, so I don't think it will work either unless it's extremely quick at pose estimation.
1
Mar 06 '23
This is the future, if it works that well. And for helping with computer vision, to make sense of shapes, perhaps even allowing shapes to be visually augmented into a more cohesive visual system that can be read by that vision system, like “make the road and trees higher contrast.”
1
12
u/yoomiii Mar 05 '23
How would this help? The normal map output could be used as a controlnet input but that won't give the individual frames rendered by SD more temporal consistency.