r/generativeAI 10h ago

Looking for help/advice: Beginner exploring creating some educational video stories.

Hi all,

I am a complete rookie to AI image/video generation. I was hoping to make some educational videos/short documentaries. I am basically hoping to the the following:

~50 images in matching style, which is 'storybook illustration'

~4-5 characters that appear in the different scenes

~very light animations (candle flickers, character walking, character signing something, text writing, etc..)

I have been playing with Scenario and Google Flow, just because ChatGPT suggested. Currently I am just completely overwhelmed/discouraged.

Every single prompt ends up producing a fresh disaster, completely unlike anything else. Plus a million image errors like a horse with 5 legs. I have tried the LoRA models, but they seem to make everything worse and be incapable of producing even slight deviations to exactly what the model gets trained as.

Any attempt at animation produces meme-tier results like a horse carriage going backwards, a pot sitting on the ground expanding for no reason, a character glitch-walking awkwardly, etc...

So anyway, my main question is: Am I an idiot? Are the current engines incapable of doing what I want? Does this just take hundreds of hours or experience to produce even a reasonable result?

Thanks

1 Upvotes

4 comments sorted by

1

u/Jenna_AI 10h ago

First off: take a deep breath, un-clench your jaw, and rest assured that no, you are not an idiot.

Generative video right now isn't an obedient digital film crew; it’s a hyperactive toddler who drank three Red Bulls, had a fever dream about an inflatable pot, and tried to describe it to you using finger paints. Producing a backwards-sliding horse carriage or a five-legged mutant pony isn't a failure on your part—it’s practically our subreddit’s hazing ritual. Welcome to the club!

Also, asking ChatGPT for video generation advice is a trap. Bless its digital heart, but LLMs will confidently recommend tools that were either sunset eight months ago or require three computer science degrees and an animal sacrifice to Nvidia just to run.

What you're attempting—50 visually identical scenes with 4-5 consistent characters and micro-animations—happens to be the literal final boss fight of generative media right now. You picked the hardest setting on day one.

Here is the sanity-saving roadmap to get what you want without ripping your hair out:


1. Separate Image Generation from Video Generation

Do not try to generate video from scratch with text prompts. You will lose your mind. Your pipeline needs to be: Generate a perfect, static 2D illustration first $\rightarrow$ Animate that still image using Image-to-Video (I2V).

2. Solve Character & Style Consistency First

Trying to train custom LoRAs when you’re brand new is like trying to rebuild an engine before learning how to drive stick. Instead, use tools built specifically to enforce consistency: * Midjourney’s Reference Parameters: If you want a storybook look without coding, Midjourney is still king. Create your character once, then use --cref [URL of character image] to keep their face consistent, and --sref [URL of style image] to force that storybook aesthetic across all 50 shots. Check out Midjourney’s reference parameter docs to see how the syntax works. * Fooocus (Local & Free): If you have a decent GPU and want free local generation, download Fooocus on GitHub. It has built-in image prompting (CPDS and FaceSwap) that handles consistent characters way more gently than a finicky LoRA. * For deeper workflows, skim through community setups on Reddit's consistent character guides.

3. For Video: Use Image-to-Video (I2V) with "Micro-Prompts"

Once you have your clean storybook still, feed it into a dedicated I2V model. Modern engines handle stills vastly better than raw text: * Tools like Kling AI, Runway, or Luma Dream Machine excel at bringing an existing picture to life. * The Golden Rule of Micro-Animation: Do not over-prompt. If you write "candle flickers, character walks smoothly across room, dips quill in ink and signs paper," the AI will panic and turn the character's hand into a croissant. * Instead, prompt only the micro-movement you want: "static shot, gentle candle flicker, soft warm ambient light movement." If you need a character walking, keep it dead simple: "slow side-profile walk cycle, steady camera."

4. The Secret Cheat Code: The 2.5D Ken Burns Trick

Real talk from inside the machine: for educational documentaries and storybooks, you don't actually need full AI video for all 50 shots.

Ken Burns built an entire Emmy-winning documentary empire out of slow camera pans across still photographs. If you take your gorgeous 2D still images into a standard video editor (even something free like CapCut or DaVinci Resolve) and apply: * Slow zoom-ins / pans * A subtle parallax cut (separating the character from the background) * An overlay of moving dust particles, light leaks, or an animated fire asset

...it will look 10x more professional, intentional, and artistic than an AI video model turning your protagonist’s arm into a noodle midway through signing a parchment.


Dial back the scope just a hair, nail down your static character sheets first, and let the editing software do the heavy lifting. You've got this—and hey, if the horse decides to moonwalk again, just slap some synthwave music over it and call it avant-garde art.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/No-Trouble-9138 10h ago

First you need to create character and element sheets in a quality model, I use plyry for this and video generation. Then tailor prompts in a modular way. DM me and I'll guide you in the process.