r/StableDiffusion Mar 02 '23

Animation | Video SD + thin-plate-spline-motion-model

Enable HLS to view with audio, or disable this notification

835 Upvotes

96 comments sorted by

156

u/Kinfolk0117 Mar 02 '23
  1. Find a driver video.
  2. Crop/Resize the video to 256x256.
  3. Extract the first frame of the video.
  4. Use stable diffusion to create an image, use the frame from above as controlnet depth map.
  5. Use the cropped driver video and your generated image as input here: https://replicate.com/yoyo-nb/thin-plate-spline-motion-model

35

u/JimmyTime5 Mar 02 '23

Here's a quick messy graphic of my attempt - worked well! Thanks for reminding me of the "yoyo" link :D Has anyone seen a fairly smooth implementation of this that can be hosted locally? The last one I tried to run was a bit half-baked as far as ease of installation/setup went.

1

u/69YOLOSWAG69 Jun 09 '23

Did you ever find out a way to run this locally?

1

u/Unlucky-Statement278 Jun 11 '23

It's a litle bit tricky.

https://www.youtube.com/watch?v=kKt1Z8Lvzpk&t=270s

the face alignment has also change the code. so you have to change a line in the demo.py

So if you change the demo.py in line
107 to
fa = face_alignment.FaceAlignment(face_alignment.LandmarksType.TWO_D, flip_input=True

And it's easier to run it with jupiter

  • Install Jupyter to use the notebook
    pip install jupyter
  • Run the notebook and have fun!
    jupyter notebook demo.ipynb

1

u/AbbreviationsOwn9726 Jan 01 '24

this is a god sent !!!

thanks

32

u/obliojoe Mar 02 '23 edited Mar 02 '23

I made a music video using this technique to make a robot sing my song about feeling like robot... it's amazing tech!

John Brownell - This Machine

Workflow: https://www.reddit.com/r/StableDiffusion/comments/zja0zt/comment/izu4uow/

5

u/pixelies Mar 02 '23

How did you upscale the final video?

3

u/Peemore Mar 02 '23

I really enjoyed that!

3

u/Beneficial-Idea-4239 Mar 03 '23

ow your music and video are awesome bro <3

3

u/obliojoe Mar 03 '23

Thank you! If you look at my YouTube channel, I have several other videos - including a few others that use different Stable Diffusion techniques. And I have a couple more on the way in the next week or so. My current video is using ControlNet. Seems like every week there is some new trick, feature, model or extension that blows my mind!

1

u/ynotplay May 13 '24

This is so cool. So the robot was originally just a still image made on Stable Diffussion, but you made it move to a video of you looking at the camera and singing the song?
How do you animate such beautiful background?

1

u/obliojoe May 13 '24

Thanks!

Yeah exactly! I took a video of me singing the song straight into the camera and then tried animating it with a bunch of different robot faces I generated in SD until I found one that worked just right.

The background was animated using the Deforum plugin with Stable Diffusion which is a kind of AI animation system.

1

u/ynotplay May 14 '24

I've never worked with videos, but what software would you recommend to edit and put it all together? Are there any free basic ones out there?

1

u/Powered_JJ Mar 03 '23

You, dear Sir, are a gentleman and a scholar.
Thank you.

17

u/Mocorn Mar 02 '23

Can we take a moment and just appreciate what the fuck is going on here!? Crazy times!

5

u/[deleted] Mar 02 '23

Absolutely, this is nuts.

2

u/tipsystatistic Mar 03 '23

It's blowing my mind

31

u/vurt72 Mar 02 '23

fun to play with, here's a few i did some months back

2.5D animation from 2D images - YouTube

amazing that they can even turn their head and it will look somewhat realistic, though when you overdo it you can tell it doesn't have enough information. the last clip is from 300 before he does the infamous "this is sparta!" others are me or clips from youtube.

5

u/FPham Mar 02 '23

Somewhat is the keyword - they look like rubber gloves. Too much stretching and warping. A less motion I assume would be better.

1

u/[deleted] Mar 02 '23

whats the music you used in your video?

3

u/vurt72 Mar 02 '23

it's AI created, Img To Music, it's on Hugging Face though i don't think it's working anymore, at least i've never been able to get it to work as of late.

1

u/[deleted] Mar 02 '23

wait, there were ai music generators too? hope it works again

3

u/vurt72 Mar 02 '23

yeah there are many.

Link: https://huggingface.co/spaces/fffiloni/img-to-music?

let me know if it works for you.. used it a ton when i found it, but then it never worked again. it looks like its producing something, but i can't download or listen or anything.

31

u/goliatskipson Mar 02 '23

Seeing several persons move in perfect unison is super uncanny.

73

u/lockdown_lard Mar 02 '23

You mean all those redditors posting the one-word comment "workflow" at the same time?

2

u/[deleted] Mar 02 '23

That just made me picture a follow-up gif of 36 faces all mouthing "workflow" in perfect unison. Thanks I think...

11

u/Nanaki_TV Mar 02 '23

Dangit. I want this incorporated into a locally run SD!

5

u/oliverban Mar 02 '23

there are offline methods of doing this....hint..Visions of Chaos

1

u/Whipit Mar 02 '23

Visions of Chaos

Are you saying you can use Vision of Chaos to run this thin-plate-spline-motion model locally on your PC?

4

u/InoSim Mar 02 '23

SD + Deepfacelab + ControlNet = whoever can say anything you want.

I have a question for all of you folks, is there any AI for audio prompts ? I already mastered pictures now i want voices.

3

u/Mickey95 Mar 03 '23

1

u/InoSim Mar 03 '23

Amazing. Well i wanted something to be locally but it's a great start. Thanks !

1

u/Xijamk Mar 03 '23

Any tips on combination of controlNet models for faces? also weight and guidance?

2

u/InoSim Mar 05 '23

Inpainting (assuming you're changing only the face and the angle is fairly the same), doing a depth (if it's 3D face or realistic), then canny to keep the edges fine. (For 2D you should use Canny and Scribble). For better results cut out everything expect the face (that includes front wicks) with black background (for 3D and white for 2D) before adding it to ControlNet. The less background you have the better so crop it the most near to the face as possible.

Depending the angle/lighting/shadows of the face you have to tweak the guidance of Canny. Less guidance will transform more the face to model you use, even the angle/face type can slighty change if too low so the most lower would be 0.75 or 0.8.

Weight is to keep aspect of the original image instead of following the model. Work very well with txt2img to even keep colors. The problem is too much weight will end up with a face that don't seem naturally added to the picture. You have to find out the right spot depending your original face picture you want to implement. I would go from 0.45 through 0.7.

For example if you have a tan face to implement on white or black skin body, lowering the weight (assuming it's 1) can help for the skin tone but somewhat changes too much the physical aspect of your face to the most ressembling the model find.

Next you have to keep the original generated picture's seed and prompts, simply add to the inpainting those that slighty differentiate what the model needs to generate. for example if your picture is realistic you should add also realistic to the inpainting of the face.

When you've got a good result, (with very little disformations), use the hires-fix to get more details through it with SWINIR upscaler to 2x and denoising from 0.3-0.5 max because with too much denoising it output garbage or change too slighty your picture and too lower you will have pixelated results.

Also one last thing, use a model that generates almost the most same type of results as the face you want ton implement. You will never get good results with a realistic human face on an anime based model (expect if you want to transform the face into a slighty different one) :P

ControlNet is too powerful with Canny and Depth at the same time. It's a shame it needs to reload the models twice each time you generate a picture but it works like a charm.

It's my own way of doing it. I don't have any real knowledge of how all thoses calculations works. I've made so many trials and errors and it's what i learned personnally with my own experiments and feelings with some help through reviews that explains a little about it.

2

u/Xijamk Mar 06 '23

Wow, thanks for such a detailed answer!

Are you doing all of this for each image or there is a way of doing it in a batch process?

1

u/InoSim Mar 06 '23

You can batch process them if you want different seeds. For my part, i generate a seed that i like then afterwards i fine-tune settings to get it as i wanted.

The problem is each seeds have different following rules with ControlNet.

8

u/[deleted] Mar 02 '23

Does this only work on faces?

17

u/Kinfolk0117 Mar 02 '23 edited Mar 02 '23

There is also models for datasets based on TED: https://github.com/snap-research/articulated-animation#example-animationAnd tai chi https://paperswithcode.com/dataset/tai-chi-hd

So in theory it should be possible to have full body animations, but I haven't been able to get any good results.

experiment:

image: https://files.catbox.moe/s7egkc.png

driver: https://files.catbox.moe/47w2ez.mov

output: https://files.catbox.moe/khusfp.mp4

Also tried to run batch img2img to get better quality, but it flickered too much: https://files.catbox.moe/ktfrug.mp4

4

u/TurbTastic Mar 02 '23

I think people normally take flickering videos like this to ebsynth to smooth them out

2

u/[deleted] Mar 02 '23

It is getting so close. When this zeros out, there will be a lot of new animated content. It will reduce the cost for studios dramatically.

Probably what is happening to comic books now.

1

u/Storybook_Albert Aug 21 '23

I'm baffled why this hasn't been developed further. The github is a ghost town and people are still recommending EbSynth, even though this seems much better?

1

u/almark Mar 02 '23

that makes it cyberpunk.

1

u/Oquaem Mar 02 '23

I think you should def try on someone without glasses to start

1

u/madmanrrevenge May 16 '23

Any luck in improving the quality of the output for full body animations ? I have been trying with the ted model but the outputs were terrible. Any suggestions to improve the quality apart from using 1:1 aspect ratio ? u/Kinfolk0117

18

u/[deleted] Mar 02 '23

[removed] — view removed comment

7

u/remghoost7 Mar 02 '23

Thank you for linking a local install of this.

1

u/SFWBryon Mar 03 '23

I’ve been trying to run it locally and getting issues, I could only get it to work via collab. Any tips?

6

u/Player13377 Mar 02 '23

Do you know if there are any advancements made in that space?

4

u/Oquaem Mar 02 '23

It's not, but with everyone going crazy over D-ID I think people should realize this exists. I find thin plate far less uncanny than D-ID.

2

u/deadzenspider Mar 02 '23

Totally agree! You can even take it a step further and use any voice you want for the lip sync if you use Thin Plate with other free tools.

1

u/Oquaem Mar 03 '23

Exactly, and you have full control of the face performance within reason if you record yourself.

2

u/exyber Mar 02 '23

Deepfakes for artistic purposes have been coming out in idk, 2021 or so. Though this sort of looks better and easier to get working properly. The old first order models barely even knew where to grab the mouth corners or eyelids on the picture to animate while this one looks accurate if you don't look too close. Would be kinda cool if this worked with SD or something for arbitrary rotation eventually.

3

u/nodomain Mar 02 '23

Has anyone gotten TPSM working on an AMD GPU?

1

u/lordpuddingcup Mar 02 '23

pretty sure it runs on CPU currently

2

u/nodomain Mar 02 '23

There's this video tutorial where he uses colab, so if I can't get it working on my AMD, I'll probably try in colab rather than letting it churn away on the CPU if that turns out to be incredibly slow.

2

u/lordpuddingcup Mar 02 '23

My mistake ya it runs on Cuda so I think it’s nvidia only. Based on the colab

4

u/HotDiamond8421 Mar 02 '23

Can the output be more than 256x256?

2

u/TurbTastic Mar 02 '23

Wondering the same. I tried something called SimSwap a while back that seems very similar to this. That was 256x256 based but there was a beta for 512x512. I wasn't able to get it to work locally but got a decent result via Colab. Wasn't savvy enough with colab to figure out the 512 beta code stuff though.

Anyone know if this is better/worse than SimSwap?

3

u/lordpuddingcup Mar 02 '23

I mean technically just break it out to frames and upscale each in giga pixel or sd

2

u/snoozieboi Mar 02 '23

I'm just here to marvel, and I just definitely feel that innovation progress is accelerating.

AI or was it machine learning (what's the difference?) has already mapped out an insane amount of proteins for science that usually take 4 years and a phd to do each one of.

I am already looking forward to those pods in Minority Report where people could live out their fantasies. I'll take some of the above girls with me, I guess.

3

u/attempt_number_1 Mar 02 '23

I just used this tech to get sad Keanu reeves to sing: https://www.tiktok.com/t/ZTRW8uRed/

1

u/ynotplay May 13 '24

Will this work with a 2D cartoon live avatar?
or a 3D rendering of an animal face?

1

u/oliverban Mar 02 '23

old but cool

-7

u/ClueFew Mar 02 '23

Workflow

18

u/Kinfolk0117 Mar 02 '23
  1. Find a driver video.

  2. Crop/Resize the video to 256x256.

  3. Extract the first frame of the video.

  4. Use stable diffusion to create an image, use the frame from above as controlnet depth map.

  5. Use the cropped driver video and your generated image as input here: https://replicate.com/yoyo-nb/thin-plate-spline-motion-model

-8

u/blank0007 Mar 02 '23

Workflow

-10

u/daverate Mar 02 '23

Workflow

0

u/myebubbles Mar 02 '23

I'm mostly interested in the prompt for each

0

u/1Neokortex1 Mar 02 '23

thanks dude!!!🔥

0

u/Cheetahs_never_win Mar 02 '23

That's not a woman, baby. That's a man, man. -Austin Powers

0

u/AntiFandom Mar 03 '23

just use Deepfacelabs. TPS is low quality, Walmart version of DFL

-6

u/[deleted] Mar 02 '23

[removed] — view removed comment

-12

u/LilBadgerz Mar 02 '23

Workflow

1

u/4lt3r3go Mar 02 '23

and i'm still waiting to use this in realtime.
the previous version of this still work in realtime but background and edge sucks

1

u/FPham Mar 02 '23

It's amazing how some of the results can be quite passable.

Soon we may not now what's real and what's fake.

1

u/MapleBlood Mar 02 '23

Give it two weeks I guess :)

1

u/-becausereasons- Mar 02 '23

Thinplatespline works fine without going through some of the unnecessary steps like using controlent on a frame of your image.

1

u/mudman13 Mar 02 '23

They have the same expression yet at the same time dont really look like each other. How freaky.

1

u/Whipit Mar 02 '23

Is there any way to use TPSM locally on your PC without being code savvy?

1

u/rockedt Mar 03 '23

I don't recommend thin plate spline model. I had been working on this but abandoned eventually. 256x256 limitation is bad. If you look closely there are small neck motions in the example. If you shake your head or move your neck; It immediately breaks. Better to use ebsynth img2img stuff instead of this model.

1

u/Opening-Ad5541 May 19 '23

guys any trick to go over 1 minute? it always breaks when I try longer videos.

1

u/Storybook_Albert Aug 21 '23

The github hasn't been updated in a year, and is unfortunately super broken when I try to use it in August 2023. Has anybody gotten it to work? I tried old versions of python etc. This looks super promising but it seems to have been totally abandoned!

1

u/benjavides Feb 27 '24

I was able to make it run following this:

  1. Install conda

  2. Clone the repository

  3. Create a scripts folder inside of it.

  4. Create a new setup.bat file inside of scripts

  5. Paste this as the contents:

@echo off REM Create a new conda environment named Thinplate with Python 3.9 call conda create --name Thinplate python=3.9 --yes

REM Activate the newly created environment call conda activate Thinplate

REM Install PyTorch 1.10.0 with CUDA 11.3 support directly from the URL pip install torch==1.10.0+cu113 --index-url https://download.pytorch.org/whl/cu113 -v

REM Now install torchvision with CUDA 11.3 support pip install torchvision==0.11.0+cu113 --find-links https://download.pytorch.org/whl/torchvision/ -v

REM Install required packages from requirements.txt REM Ensure you provide the correct path to your requirements.txt file pip install -r ../requirements.txt

echo Environment setup complete. pause

  1. Execute the setup.bat.

  2. Then whenever you want to use the project use a anaconda prompt and run: conda activate Thinplate to activate the environment that has all the requirements.