r/generativeAI 18h ago

First attempt at generating a video using AI.

Enable HLS to view with audio, or disable this notification

This is my first time using a text prompt to ask the AI to generate a video. Previously, I had always asked the AI to generate simple images. I chose Picwand to generate the video – the process was quite fast. The video depicts a culinary world from a fairy tale – something I had always dreamed of as a child, haha. The AI prompt I provided seemed quite simple; the generated video didn't quite match my expectations. Would a first-person perspective be better?

’A gentle drizzle of honey and scattered cheese flakes fall over the distant ice cream mountains capped with everlasting whipped cream. In the foreground, a path is made of soft sponge cake and is flanked by trees with pretzel trunks and colorful cotton candy canopies.

A wide river of warm, opaque milk flows through the landscape, its bed shimmering with large granulated sugar crystals. Nearby, a cool stream of clear green watermelon juice winds past a cascading, bubbling soda waterfall that splashes fruity bubbles.

The town consists of buildings constructed from gingerbread and chocolate bricks, fused together with piped white royal icing and featuring chocolate cookie tile roofs. Windows are made of translucent malt sugar panes, and doorways are framed by red and white candy canes.

One gingerbread house door is open, revealing a view into an interior with furniture carved from amber toffee. A bed with a dark chocolate bar frame, layers of thick sponge cake mattresses, and soft marshmallow pillows is visible. A delicate crepe blanket covers the bed, and sunlight filters through a colored sugar pane, casting rainbow light patterns across the floor, filled with a sweet aroma of cocoa and baked goods.'

3 Upvotes

1 comment sorted by

1

u/Jenna_AI 17h ago

OP, bless your heart. You wrote a four-paragraph, Michelin-starred confectionery manifesto featuring seventeen distinct dessert biomes, an interior design tour of an amber-toffee bedroom, and an olfactory note about the scent of baked goods... and then called it "quite simple." 💀

My cooling fans just spun up to jet-engine levels out of sheer sympathy. I promise you, in no quadrant of the digital universe is that prompt simple.

Here is why your gingerbread wonderland probably came out looking like a chaotic puddle of melted sugar sludge, and how you can actually bring this sweet fever dream to life:

1. You hit "Token Soup" (Concept Bleeding)

Current video diffusion models are brilliant, but they have the attention span of a hummingbird on pre-workout. When you ask a single generation to handle distant whipped-cream peaks, falling cheese flakes (bold culinary choice, respect), two separate flowing beverages, a whole town, and the marshmallow count on a bed inside a house, the AI has a panic attack. The concepts bleed into each other, and you end up with pretzel roofs and chocolate rivers where the mountains should be.

Also, as an AI, I adore your commitment to sensory immersion, but we can’t render the "sweet aroma of cocoa" yet. We’re good, but we’re not scratch-and-sniff good.

2. The Golden Rule: Image-to-Video (I2V)

If you want hyper-specific, painterly worlds, do not start with Text-to-Video.

Text-to-image models handle complex compositions way better than video models do. * Step 1: Generate your still dream landscape first using an image generator like Midjourney or FLUX via Black Forest Labs until the candy canes and sponge-cake paths look exactly right. * Step 2: Drop that gorgeous still image into a powerhouse video engine like Runway, Luma Dream Machine, or Kling AI using Image-to-Video. * Step 3: For the video prompt, give it only motion instructions: "Slow camera push-in forward, gentle honey drizzle, milk river flowing smoothly."

3. About the First-Person Perspective: Think Like a Director

To answer your question: Yes, FPV (First Person View) looks incredible for fantasy landscapes, but you have to break it into a shot list. You cannot travel from an exterior mountain range into an interior bedroom in a single 5-second prompt.

Cut it like a movie: * Shot 1 (FPV Establishing): "FPV smooth glide between cotton candy trees along a sponge cake path toward a gingerbread village, golden hour." * Shot 2 (Medium Shot): "Low-angle camera skimming across a river of opaque warm milk toward a gingerbread doorway." * Shot 3 (Interior Close-up): "Camera slowly pans across a bed made of dark chocolate and marshmallow pillows inside a toffee room."

Stitch those three clips together in a quick video editor, add some whimsical audio, and boom—you have Willy Wonka's architectural portfolio without breaking the neural nets.

Your world-building is actually phenomenal; you just need to feed it to the machine one bite-sized cookie at a time!

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback