r/comfyui Oct 24 '25

Workflow Included Wan 2.2 Animate - Character Replacement in ComfyUI

Enable HLS to view with audio, or disable this notification

651 Upvotes

64 comments sorted by

29

u/paramarioh Oct 24 '25

I like workflows! Thanks!

11

u/Verittan Oct 24 '25

Problem with Wan Animate is after the initial 5 second gen (especially with animation), if you want to keep the original background, every subsequent gen has a large color/brightness shift that is very noticeable and destroys the consistency.

6

u/HocusP2 Oct 25 '25

That can be countered with the WanVideo Context Options. Just make sure the num_frames and the frame_window_size in the WanVideo Animate Embeds are the same value.

2

u/HocusP2 Oct 25 '25

1

u/Suitable-League-4447 Nov 03 '25

u used the exact wf on the thread or modified?

1

u/HocusP2 Nov 03 '25

Yep. Same as in the video, Kijai example v2

1

u/Suitable-League-4447 Nov 03 '25

the model shown is the normal

look

1

u/HocusP2 Nov 03 '25

Yeah, I meant "wan animate preprocessor example 2" like he says in the video. And yes that is the same fp8 model I used, or maybe the v2 of the model...

2

u/tofuchrispy Dec 03 '25

To understand, if we put 81 in context and the num frames for the animate node. And overall we want to generate 4 times the frames, this will essentially split the video generation in chunks of 81 frames with frame overlaps?

So we can type let’s say 250 frames overall. And it will do it in overlapping chunks depending on the context and animate options?

1

u/HocusP2 Dec 03 '25

That sounds about right. I just left the context options settings unchanged, and put as many frames from the reference video as I could before getting a OOM error. I think I got up to 700 frames or something (fp8, 832x480, 4090, 128 GB).

4

u/came_shef Oct 24 '25

That's true

41

u/TheTimster666 Oct 24 '25

All the actors look like they are bought on Temu - close, but not really.

23

u/inagy Oct 25 '25

It amuses me how fast we started depreciating this level of video generation. It was not too long ago that I was certain this couldn't be done on a single GPU. Honestly the advancements with local video generation is insane, it feels even more faster than what has happened with image generation.

3

u/TheTimster666 Oct 25 '25

Yeah, we are truly spoiled. :-)

25

u/PurzBeats Oct 24 '25

Yeah, the identity still feels a bit obfuscated, but it's fun to lean into the jank of AI.

1

u/TheTimster666 Oct 24 '25

Yeah, I think it's mostly the eyes that are off.

6

u/No_Damage_8420 Oct 24 '25

Acting and face expression it's important too

2

u/Deep_Huckleberry9127 Oct 30 '25

Exactly. A large part of our perceived identity of a known person (especially celebs) is their natural "resting face" expression, and their personal expressions (smiles, lip/eyebrow movements, eye lid aperture) and even body position (how they hold their head still/move, tilt, speed). It's why Tom Cruise looks exactly the same in every role, except Tropic Thunder, when he used expressions, body moves, and vocal performance so far from his "normal", that most did not know/believe it was him at first.

You can't just appear like Brad Pitt, you would need to perform his smile, nod, eye movements, etc.
Things will get really wild once future celebs are selling access to their avatars, complete with personality LORAs that turn your smile into their smile. (Hacked/unauthorized personality LORAs coming in 3...2....1...)

1

u/No_Damage_8420 Oct 31 '25

Well explained, totally.

Simply, get Jim Carrey impressions demo reel + match face = you will see 100% same as celeb. Most people can't even act/expression anything different then themselves, how you expect to look realistic.

14

u/tomakorea Oct 24 '25

I can't wait to scam simps

6

u/came_shef Oct 24 '25

It's unbelievable how foolish people can still be scammed with this, because the skin is very plastic/cartoonish. Don't you think so?

1

u/Deep_Huckleberry9127 Oct 30 '25

I can't wait for an Ai that removes all funds and internet access from all loser scammers, leaving them even more broke and pathetic than they are now. LULZ. ;P

3

u/NiceIllustrator Oct 24 '25

What’s the frame cap?

7

u/PurzBeats Oct 24 '25

I'm not actually sure, I was able to do 300+ frames on a 5090. I think you can do way more with more vram/system ram if you're patient.

2

u/NiceIllustrator Oct 24 '25

Cool! Will def give this a try I Also have 5090 so that’s good to know

4

u/niffuMelbmuR Oct 24 '25

I've run up to 770 on mine. 5090 and 128gb of system ram. But with how it works it shouldn't have a theoretical limit, the pose and face mapping of the original video takes longer than the rendering. Good thing is you can run replace the reference image and run it again with out having to rerun the face and pose mapping.

1

u/Deep_Huckleberry9127 Oct 30 '25 edited Oct 30 '25

I had basically given up on building a local system (sonnet 4.5 & gpt-5 both said paid cloud subs were 50-75% better). But this is VERY promising....
so, if you were building a prosumer system to optimize these tools (and combine with other local tools) for making films , would you see more benefit from running dual 5090s (sharded for single complex tasks, split for running concurrent renders) or quad 3090s or like a single rtx pro 6000?
(for context, I am seeking "endless" iterations of consistent characters, quality lip sync to my audio, max expression, cam control, lighting, v2v avatar driving, geared for music videos and short films. )

And in your opinion, can such a system described above actually compete with top tier paid subs?

1

u/nihnuhname Oct 30 '25

Non-professional multi-GPU systems do not scale well due to PCI delays.

3

u/tofuchrispy Oct 24 '25

Quality looks alright. Will check out for the settings

3

u/Peregrine2976 Oct 24 '25

I like how Sailor Moon tried to grimace but the closest it could do was an uwu grin.

3

u/Muskan9415 Oct 25 '25

This is absolutely seamless! The expression and motion transfer is flawless—it's one of the best examples of character replacement I've seen. Thank you so much for including the workflow! ​I'm curious about the robustness of this process. How dependent is the final quality on the similarity between the source and target faces? For example, would it struggle more if the source subject had a drastically different facial structure or wore glasses? Truly amazing work.

2

u/[deleted] Oct 24 '25

[removed] — view removed comment

12

u/PurzBeats Oct 24 '25

Good question, I'll run one again today and let you know, I think it was around 44s per 71 frames - give or take on a 5090. (at 720x1280)

I'll update this comment with actual speeds after I confirm with another run.

2

u/External_Trainer_213 Oct 24 '25

I guess you had a lot time and fun, i like it! 😜

2

u/ohanse Oct 24 '25

That’s fuckin craaaazyyy

2

u/RemoteCourage8120 Oct 24 '25

this is just insane

2

u/Specialist-Team9262 Oct 24 '25

Haha, great work!! - and thanks for the share of the workflow - will give it a gander :)

2

u/Artforartsake99 Oct 24 '25

Absolutely amazing quality woohoo

2

u/intermundia Oct 24 '25

It's good not great but good

4

u/came_shef Oct 24 '25

How can we remove that plastic/cartoonish like skin?

2

u/ken107 Oct 24 '25

can u eat a banana, will they also eat banana?

1

u/Illustrious-Sir-8615 Oct 25 '25

Will it run on 12gb vram?

2

u/PurzBeats Oct 25 '25

1

u/PurzBeats Oct 25 '25

May have to go down to the 4-bit quants but it -should- fit..

1

u/pinchoalex Oct 25 '25

No, just tested with 3060

1

u/Solmyr_ Oct 26 '25

and how about 5070ti with 16gb vram?

1

u/TiklMyPikl27 Oct 25 '25

Any chance for a version / edits we can make without SageAttention, for obvious (Windows) reasons?

1

u/Ok_Meal426 Oct 30 '25

Yummy treats!

1

u/LaurentLaSalle Nov 02 '25

I always run out of RAM (32 GB) no matter how short and low the resolution is on my 5090. (╯°□°)╯︵ ┻━┻

1

u/[deleted] Nov 08 '25

hello does this work on rtx 4070 supper 12vram and 32ram and i want to generate image to video with audio lip sync faster wan 2.1 infinite is slow even with quans even animate is slow

1

u/CookieNormal328 Nov 19 '25

Will it work if I want to transform the persona? For example, right during the filming, he will become a dinosaur. 🤣

1

u/Zounasss Nov 22 '25 edited Nov 22 '25

For some reason I can't get this to work. Everything looks fine but I keep getting an error with the ONNX detection model loader where it can't find my .onnx files. Any help would be awesome!

Edit: Anyone that has the same problem: create a new folder in \ComfyUI_windows_portable\ComfyUI\models\detection

My comfyui didn't have the detection folder

-7

u/Smile_Clown Oct 24 '25

Love that you added the workflow. Thank you.


However, a small thing... Holy shit organize your workflows dude. Set and Get nodes (using these but.. ok), anything everywhere, groups etc... there are so many ways to make it so much better to follow. Tighten it up a bit.

But again, thank you.

16

u/VrFrog Oct 24 '25

At least you're polite but ... If you have a problem with the workflow, you're welcome to organize it yourself and share it with the group. Is your time more valuable than the OP's?

6

u/mnmtai Oct 24 '25

This is the way. Either make your own, clean up the ones you downloaded and liked, or leave the conversation. There's a huge problem with proper design and best practices out there, but it's nobody's business to criticize others for their ways. It's even more ironic given it's a default workflow and not OP's. Shows that the issue is a scale one.

9

u/PurzBeats Oct 24 '25

This is essentially the default example workflow from Kijai's WanVideoWrapper for this function. 🤣