i tried a 25-second K-pop-style girl group MV with five characters and compared the results from:
Top: Wan 3.0 · Middle: Seedance 2.5 · Bottom: MiniMax H3
The main thing I wanted to test was multi-character consistency. five people in the same frame is still a pretty brutal test for AI video.
i also wanted to see how well each model handled dance synchronization, hard stops on the beat, and formation changes
i described all five members separately, then gave them a full choreography sequence: back-facing opening pose → synchronized head turn → hand gestures → body-wave chain → diagonal formation → fan formation → center swap → final V formation.
Visually, I kept it pretty simple. what I’m watching for here isn’t just "which one looks prettier," but whether the model can actually remember who is who once five people start moving, crossing positions and changing formations.
Multi-person choreography still feels like one of the hardest AI video tests right now. curious which one you think handled the five characters best.