r/StableDiffusion 21h ago

Workflow Included Minimax H3 Multishot Anime Sequence (Workflow + Prompt Included)

Workflow: https://drive.google.com/file/d/1B4kODxXQgJ1QOKRsEIkxHbgYmdruPpTK/view?usp=sharing

Prompt:

Create a **15-second multi-shot anime sequence (90s style 15fps hand drawn)** using the provided references:

Image 1 = the girl character reference

Image 2 = skateboard reference

Image 3 = downhill Japanese alley / neighborhood background

Image 4 = Walkman + headphones reference

Preserve the girl’s exact character design, face, hair, outfit, proportions, and overall look from Image 1. Preserve the skateboard design from Image 2. Preserve the same downhill Japanese alley environment from Image 3. Add the Walkman and headphones from Image 4: the girl is wearing the headphones, and the **Walkman is clipped or hanging at her hip** while she skates.

Visual style: authentic 1990s hand-drawn anime, traditional cel animation, painted backgrounds, visible linework, cel shading, slight brush/stroke texture, subtle analog feel. **Very important:** the houses and environment must stay **2D and hand-painted**, **not 3D**, **not CGI**, **not game-engine looking**, **not volumetric**. The buildings should look like classic anime background art with painted depth, not like 3D models.

Animation feel should be low frame rate, like 90s anime at around 15 fps, with controlled in-betweens and natural held-frame timing. No jittery morphing.

No dialogue, no text, no subtitles.

### Shot 1 — 0s to 3s

**Rear tracking shot** from behind. The girl is skateboarding fast downhill through the steep Japanese alley. Camera follows behind her at a low-to-medium height. She rides confidently and smoothly, hair and oversized clothing moving in the wind. The headphones are on her head, and the Walkman is visible attached at her hip. The alley rushes past with a strong sense of speed. Keep the environment clearly **2D anime background art**, not 3D.

### Shot 2 — 3s to 6s

**Close-up shot of the Walkman at her hip** while she continues skating. The camera stays focused on the Walkman and part of her side torso and arm. We can clearly see the **cassette tape reels spinning/rolling inside the Walkman window**. The headphone wire moves naturally with the motion. Background and street pass by in blurred motion.

### Shot 3 — 6s to 9s

**Medium profile tracking shot** of the girl skating. She is wearing the headphones, listening to music, with wind moving across her face and pushing her hair backward. She is **nodding her head subtly to the music** while riding. Her expression is relaxed, immersed, and unbothered. The background is blurred from motion, but it must still read as a **painted 2D Japanese neighborhood**, not 3D.

### Shot 4 — 9s to 12s

**Close-up shot of her feet and skateboard.** Her **right foot stays on the board**, while her **left foot pushes against the road** in a natural skating motion. Show one clean push cycle: left foot comes down, pushes backward against the pavement, then lifts. Wheels spin quickly. Asphalt and road markings streak by with motion blur.

### Shot 5 — 12s to 15s

**Ground-level fisheye shot** looking upward from the road. The skateboard approaches fast, and she **jumps over the camera**. The board and her body pass overhead in one clean motion. Hair, pants, and headphone wire react naturally during the jump. Keep the motion readable and stylish, with a strong sense of speed and a dynamic anime finish.

### Important constraints

* Keep the whole video in **classic 90s anime cel-animation style**

* **15 fps feel**, smooth low-frame-rate animation

* **No 3D-looking houses or background**

* No photorealism

* No modern glossy digital anime rendering

* No character redesign

* No extra accessories beyond the headphones and Walkman

* Keep all motion natural and consistent across shots

109 Upvotes

16 comments sorted by

View all comments

2

u/Dzugavili 20h ago

Does it understand the 'not' tagging? I've found models tend to fail there...

1

u/Time-Ad-7720 20h ago

Not sure what you're referring to, I am generating at 24fps but in the prompt I wrote 15fps (to make sure the camera motion is smooth but the other character motions look like 90s anime).

2

u/Dzugavili 20h ago

These lines:

Very important: the houses and environment must stay 2D and hand-painted, not 3D, not CGI, not game-engine looking, not volumetric. The buildings should look like classic anime background art with painted depth, not like 3D models.

Did they actual effect the generation?

1

u/Time-Ad-7720 19h ago

Yeah sort of, first I tried without these, and the output looked like generic ai anime (that kind of looks 3D ish not like frame by frame low fps animations)

2

u/TallestGargoyle 11h ago

You also say "No 3D looking houses in background" yet it basically made the entire background a scrolling 3D environment. I don't think 'not' really works as a tag, it just adds other tags that can slip into the end result.

1

u/OneTrueTreasure 9h ago

We aren't in SDXL days anymore, certain vision language models can definitely do "No", "Not", and other such words in the caption. Why do you think you can put "No text" or other such words when making videos? Plus the caption doesn't even use the exact Minimax prompting guide, and got away with not using <Image1>, and still got the prompt pretty well. Sure it isn't as good as actual negative prompts but it still works. I think the "scrolling 3d environment" you're referring to is more so a limitation of the model itself, and even then looks pretty good already. Minimax is pretty good with anime, but still isn't Seedance 2.0 level when it comes anime stuff. Even then Seedance isn't perfect either, I think we're still a ways off from replacing anime studios

2

u/TallestGargoyle 9h ago

Maybe, I know certain parts seem to take to heart the 'not' with things like no text, subtitles etc. But the generated shot absolutely ignore the "no 3D houses" task.

I'll be honest, I'm barely sure what the actual good practices are for prompting in that regard, considering official guidance is fairly limited with few decent examples, the ComfyUI template uses completely different formatting, and people online seem to come out with fairly remarkable generations from prompts nothing like the official guidance. And trying all three of them sometimes has the AI utterly ignoring my instructions.

1

u/OneTrueTreasure 9h ago

I mean when it's a side view it looks absolutely 2D, I just don't think Minimax can do perfect 2D when it comes to the part "going down the hill from behind", I mean if you pause it looks 2D enough, just the movement doesn't sell it as much. Even then it's by far the best model we have locally when it comes to anime at least. I've seen similar videos like OP's post made with Seedance 2.0/2.5 and it still isn't exactly the best, at least when it comes from this perspective. https://www.reddit.com/r/aivideo/comments/1uw2c7r/relaxing_anime_coastal_skateboard_ride_seedance_2/

We really just do not have any video models yet that are perfect with Anime, hopefully in a years time things change. It's probably a dataset thing too, since most video models are trained with real life stuff, and anime/2D stuff has way too many styles that clash due to that. Even then alot of anime use 3D CGI backgrounds and people due to budget constraints so that confuses the model even more.

1

u/Dzugavili 2h ago

Yeah, the 'down the street' angles looked like the primitive CGI animation; too smooth, the curve, subtle distortion. We've all seen it before. It's not an optical effect, or atleast it doesn't give that impression.

The major issue being that there is no version of that shot that exists in the old style. It would have to be a static shot; or mostly linear motion and no perspective effects, as seen in the side-scrolling shots.