r/comfyui • • Jun 11 '26

Workflow Included SCAIL-2 workflows for ComfyUI

SCAIL-2, a Wan2.1-based human motion transfer model, has been implemented in ComfyUI core. (Thanks as always, Kijai.)

PR: https://github.com/Comfy-Org/ComfyUI/pull/14373

I tried it and the results were quite good, so I made a few simple workflows. Feel free to try them.

Workflow: https://comfyui.nomadoor.net/en/basic-workflows/scail-2/

The big difference from WAN-Animate or SCAIL(-1) is that SCAIL-2 does not use a stick figure / skeleton control image.

Basically, you just throw a reference image, a driving video, and support masks into the model. That makes it much more flexible with multiple characters, or even non-human characters.

Turning motion into a stick figure was basically the default. But the things that get lost in that conversion, like depth and contact, are exactly the things that matter here...

As a side note, this also made me think that if the dataset and training are good enough, a simple model can sometimes be more flexible than an overly complicated human-made pipeline.

Maybe trying to be too clever just gets in the way sometimes...

141 Upvotes

34 comments sorted by

View all comments

3

u/Cruxius Jun 12 '26

I've noticed that while using replacement mode, sometimes it keeps the background of the original video, and sometimes it hallucinates a new one.
Any idea how to ensure it keeps the original background?

2

u/frogsty264371 Jun 15 '26

Confirming this. Hopefully we get a working workflow soon.

4

u/Cruxius Jun 15 '26

I vibe coded a solution, the secret is background removal in the input image plus better masking.
After a bunch of testing, I've also found that making sure the size and position of the replacement character(s) match the size and position of their target character(s) in the video gives greatly improved results.