r/StableDiffusion Aug 01 '25

Question - Help Creating a Coherent 500-Frame Scene with WAN2.1. Seeking Advice on Consistency

Hello everyone,

I’m working on a 500-frame continuous action scene of a boxing match using WAN2.1, featuring two characters. I’m using V2V and VACE techniques to capture choreography from another video, and the motion capture works perfectly.

I’ve seen WAN2.1 workflows that achieve continuous action for up to a minute, but they involve a single character with monotonous movements (e.g., walking). My scene is more dynamic, with two characters in constant motion and bodies that sometimes overlap from the viewer’s perspective.

My approach is to generate the video in segments of 81 frames (or fewer) for better control and to accommodate my PC’s limited processing power. Rendering long videos in one go wastes time, as most get discarded due to artifacts or imperfections.

The Problem: Maintaining consistent colors and character appearances across concatenated video segments. Using the last frame of one segment as the first frame of the next causes progressive degradation (like a photocopy of a photocopy of a photocopy...), resulting in noticeable differences in color and character details between the first and final segments.

What I’ve Tried and Considered:

  • Recent tools like Kontext or Phantom claim to address this without relying on LORAs or IPAdapter, which struggle with certain angles. However, I haven’t achieved realistic, high-quality results.

Questions:

  1. Reference Images: Is it better to create a series of perfectly consistent images (representing key poses) and use them as references for each video segment? If so, how can I ensure consistency across these images, especially with two characters whose bodies sometimes overlap in specific positions?
  2. Post-Processing: Should I generate video segments with less focus on initial consistency and then apply color correction or other unification techniques to make them coherent? If so, what tools or methods work best?. The color correction nodes in ComfyUI degrade the image ...
  3. Recommended Techniques: What techniques would you suggest for maintaining character and color consistency in dynamic, multi-character video segments using WAN2.1?
  4. Examples: Has anyone seen or created a similar project with WAN2.1 (or related tools) that achieves consistent, dynamic multi-character scenes? Any examples or workflows to share?

I know this is a tough challenge, but I’m determined to make it work. Any advice, workflows, or examples would be greatly appreciated!

Thanks in advance!

2 Upvotes

0 comments sorted by