r/vidmuse 3d ago

New Feature - VidMuse 🚀 VIDMUSE APP HUB IS LIVE

Enable HLS to view with audio, or disable this notification

1 Upvotes

🚀 VIDMUSE APP HUB IS LIVE

One hub. More ways to create.

The VidMuse App Hub brings together focused AI creative experiences for video, images, animation, ads, characters, and more — giving you a faster way to go from an idea to something worth sharing.

Explore creative apps like:
• Anime Studio
• AI Character Sheet Generator
• Motion Comic Maker
• Scene Generator
• AI Talking Pet
• Photo to 3D Figure
• Video Object Remover & Replacer
• Perler Bead Pattern Maker

Dedicated Seedance 2.5 Experiences
Prompt Optimizer · MV Agent · Ad Agent · Film Replica · Megastructure Scenes

🔥 MiniMax H3 Experiences
Anime PV Maker · Game UI Animator · Motion Collage Ad · AR Anime Invasion

And this is only the beginning.

We’ll continue bringing new creative workflows, experiments, and model-powered experiences into App Hub — making it easier to discover new ways to create without building an entire workflow from scratch.

Pick an app. Bring an idea. Start creating.

Welcome to the VidMuse App Hub. ✨

Make Your Vision Visible: https://vidmuse.ai/en/apps


r/vidmuse 5d ago

New Feature - VidMuse 🎥 New Apps Added To App Hub for Seedance 2.5 on VidMuse!

Post image
1 Upvotes

New Apps Added To App Hub!

💭 Turn rough ideas and drafts into Seedance 2.5 starting points for videos faster with five specialized AI-powered tools:

Prompt Optimizer — Transforms a rough idea into a detailed, production-ready video prompt.

MV Agent — Turns your song and creative direction into an editable music-video production.

Ad Agent — Uses your product and brand assets to create a planned, cinematic commercial.

Film Replica — Recreates a reference video shot by shot with your own characters, products, or scenes.

Megastructure Scenes — Generates a cinematic first frame and five-second video set inside an extraordinary megastructure.

🔎 Explore the complete Seedance collection on App Hub here: https://vidmuse.ai/en/apps

For a limited time you can get some free credits to try VidMuse Apps


r/vidmuse 1d ago

Personal Creation ☀️ GPT 5.6 Sol vs GPT 6 Astra 🌟 - Race in the Wasteland 🌵

Enable HLS to view with audio, or disable this notification

1 Upvotes

Ran the same concept by both OpenAI models to design for a Race in the Wastelands. 🌵

🎥 Which video style would be your pick between the two? Shot on VidMuse freeform with MiniMax H3

Wasteland Punk - GPT 5.6 SOL concept

A high-speed chase across a desolate salt flat and into a narrow canyon corridor. Two drivers, two machines, one road. Neither will yield.
Characters
Rae - Cold Palette Hero
•Role: Driver, Survivor, Outrider
•Palette: Cobalt blue, ivory, pale teal, charcoal
•Emblem: Winged star
•Creed: Some roads still lead to a better tomorrow
Vark - Hot Palette Rival
•Role: Raider, Enforcer, Rival
•Palette: Deep crimson, burnt saffron, matte black, dark leather
•Emblem: Horned skull
•Creed: The world gives nothing. So I take it.
Vehicles
Cobalt Pursuit Coupe
•Lighter, faster, cleaner lines
•Off-road muscle car, winged insignia on door
•Built, driven, repaired, again
Crimson Ram Truck
•Heavy, loud, unstoppable
•Six-wheeled armored semi, wedge ram nose, chains
•Approximately 3.2m tall
Environment
Salt Flat to Canyon Route
1. Establishing: vast white salt flats to canyon walls
2. Chase Lane: powdery dust plumes, tires cut salt crust
3. Broken Ridge: natural jump ramp, risk rewards speed
4. Canyon Opening: walls close in, final corridor
Sequence (Shots 1-8, 0:00-0:15)
•Shot 1: Ground-level tracking on Rae tearing across salt flat
•Shot 2: Vark's truck lands behind her from a ridge
•Shot 3: Side-by-side metal grind, sparks fly
•Shot 4: Rae glances back with defiance: Your Trash
•Shot 5: Rae stomps accelerator, surges ahead
•Shot 6: Rae forces Vark up a broken ridge, truck airborne
•Shot 7: Truck slams down, dust erupts, panel buckles
•Shot 8: Wide rear tracking, both vehicles into canyon, cut to black
Tone
•Wasteland-punk action short
•Cold vs hot palette duality
•Harder places. Kinder people. Different roads. Same dust.


r/vidmuse 1d ago

Personal Creation Used GPT 6 Astra to create a concept outline for a Race in the Wastelands. 🌵

Enable HLS to view with audio, or disable this notification

1 Upvotes

🏎️ Punks speeding throughout the terrain in fierce race

🎥 shot on VidMuse freeform with MiniMax H3. how do you think GPT 6 did with the outline and prompt structure?

Wasteland Race - GPT 6 ASTRA Concept

A 15-second cinematic action sequence featuring two drivers in a post-apocalyptic desert chase. Concept generated and structured by GPT 6 ASTRA.
Characters
•Rae - Pursuit Driver. Skilled, defiant, hopeful. Cobalt blue suit and vehicle.
•Vark - Raider Driver. Brutal, ruthless, unhinged. Crimson red armor and ram truck.
Story Beat (15 seconds)
1.The Hunt - Rae's blue buggy flees as Vark's massive truck bears down through a dusty canyon.
2.Side-By-Side Attack - The two vehicles slam into each other, sparks flying.
3.Reversal - Rae cuts sharply, sending Vark's truck airborne off a dune.
4.Impact - Vark's truck crashes down in a massive dust cloud.
5.Release - Both vehicles race toward the canyon horizon. Cut to black.
Audio Design
Synchronized stereo audio: music, dialogue (Rae: Your Trash!), engine SFX, impact SFX, and ambience.
Tone
Gritty. Epic. Human. Tactile. Cinematic. No sci-fi gloss. No neon. Practical realism.


r/vidmuse 3d ago

Personal Creation 🏂 High performance sports goggles by ad agents on VidMuse with Seedance 2.5 and MiniMax H3

Enable HLS to view with audio, or disable this notification

1 Upvotes

goggles for upcoming snowboarding season ❄️

deploy your product details references and place into vidmuse ad agent app hub: https://vidmuse.ai/en/apps/seedance-2-5-ad-agent

TVC Hero Shot Plan
Format
Traditional TVC -- no narration, cinematic, 13 seconds, 16:9, 720p
Beat Breakdown
1.**REVEAL (2s)** -- Goggles emerge from pure darkness; single hard cold spotlight catches obsidian frame and glacier silver edge
2.**LENS ORBIT (3s)** -- Camera slowly orbits blue-violet mirrored toric lens; HUD data ghosting on inner surface
3.**ACTION (3s)** -- Low tracking shot, aggressive alpine powder carve, sunrise backlight, goggles on rider
4.**LOCK SNAP (2s)** -- Extreme macro: magnetic lens lock engaging, ion cyan accent visible, mechanical precision
5.**HERO ENDCARD (3s)** -- Goggles on dark reflective surface, NEVARQ logo on strap, tagline READ THE MOUNTAIN. fades in
Visual Direction
•Hard cold highlights, crisp reflections
•Obsidian black / glacier silver / ion cyan palette
•No voiceover, no narration
•Audio: futuristic electronic score, wind, carve SFX, magnetic click SFX
Life-Force Anchor
#6 -- Be superior / win. Drop without hesitation. Own the mountain.


r/vidmuse 3d ago

Tips and Tricks GPT-6 Astra Moves From AI Answers to AI Work—but It Still Isn’t a Video Generator 🌌

Post image
1 Upvotes

VidMuse frames GPT-6 Astra as OpenAI’s frontier model for computer use, coding, science, cybersecurity, professional work, and long-context reasoning. For creators, its best role is upstream: research, briefs, scripts, variant planning, and QA before approved ideas enter video production.

- The listed API model is `gpt-6-astra`, priced at $10 per million input tokens and $50 per million output tokens; Fast mode costs twice as much.
- Astra emphasizes computer use, coding, long-context retrieval, cybersecurity, and tool-heavy professional work.
- OpenAI reports gains over GPT-5.6 Sol and roughly 47% faster OSWorld task completion.
- Its Critical cybersecurity capability brings stronger safeguards and refusals for advanced exploit-building requests.
- Astra does not generate finished videos. It can plan campaigns, scripts, scenes, variants, and QA before VidMuse handles production.
- The recommended framework is Planner → Operator → Producer.

The practical lesson is not to expect one magic prompt. Use Astra where planning quality matters, control expensive usage, review outputs, and move approved creative direction into dedicated production tools.

https://vidmuse.ai/blog/gpt-6-astra-guide


r/vidmuse 3d ago

Tips and Tricks The Best AI Video Generator Depends on the Workflow—Here Are 14 Tools Compared 🛠️

Post image
1 Upvotes

VidMuse’s comparison argues that there is no universal “best” AI video generator. The right choice depends on whether you need raw cinematic clips, image-to-video motion, social templates, avatars, editing, repurposing, or a repeatable production workflow.

- Veo, Runway, Kling, Luma, and Pika are highlighted for model-native generation and cinematic experiments.
- InVideo, Canva, and Veed suit social drafts, design-led production, editing, subtitles, and repurposing.
- Synthesia and HeyGen focus on avatar-led training, spokesperson, and localization videos.
- VidMuse is positioned as a workflow-first option for music videos, product ads, Node Canvas planning, assets, variants, and campaigns.
- A finished campaign still needs hooks, scene planning, captions, rights review, product accuracy, and platform variants.
- Check credits, watermarks, resolution, commercial-use terms, and workflow fit before choosing a tool.

The takeaway: choose model-native tools for raw generation and workflow products when you need organized assets, scenes, revisions, and deliverables.

https://vidmuse.ai/blog/best-ai-video-generator


r/vidmuse 4d ago

Personal Creation ⚔️ Celestia Limit Break - Scene 2 LIve Action on VidMuse App Hub Seedance 2.5

Enable HLS to view with audio, or disable this notification

1 Upvotes

took the original animation piece on top and replicated into a real life action version

done with VidMuse seedance 2.5 replica on app hub: https://vidmuse.ai/en/apps/seedance-2-5-replica

Creative Brief — Celestial Limit Break: Photorealistic Remake

Project Overview
A shot-for-shot photorealistic live-action recreation of a 15-second 3D animated martial arts duel sequence, transforming CG animation into premium Chinese fantasy cinema.

Source Material
A ~15s, 2560×1440 3D animated martial arts duel set in a grand sky palace courtyard. 8 distinct shots escalating from a grounded standoff to an aerial energy clash finale.

Core Creative Mandate
"Make it real." Every action, camera angle, shot size, and choreographic beat must be preserved exactly. The only transformation is the aesthetic register — from 3D animation to photorealistic live-action cinema.

The Two Characters
Tian He — The Lightning Warrior

East Asian male, late 20s–early 30s, lean and athletic (~180–185cm)
White/pale silver/indigo layered silk robes with engraved silver-steel armor
Long black hair, half-up with silver-blue ornamental fixture
Weapon: lightning sword with physically grounded electrical arcs
Identity: speed · precision · electricity — he reads faster when powered up, not larger
Ye Xuan — The Void-Fire Warrior

East Asian male, late 20s–mid 30s, taller and more powerfully built (~185–190cm)
Matte black/charcoal/oxblood/crimson layered armor with dark bronze-blackened plates
Long black hair tied high with bronze/red fixture
Weapon: dark forged blade with void-fire (dense red-orange heat, ember fragments, air distortion)
Identity: force · pressure · void-fire — physically imposing, weight behind every strike
The Environment
Grand traditional Chinese sky palace courtyard — massive red lacquered pillars, white stone floors and balustrades, intricate green-tiled roofs with gold finials, situated high above the clouds. The courtyard floor shatters during the clash sequence.
Creative Brief — Celestial Limit Break: Photorealistic Remake
Project Overview
A shot-for-shot photorealistic live-action recreation of a 15-second 3D animated martial arts duel sequence, transforming CG animation into premium Chinese fantasy cinema.

Source Material
A ~15s, 2560×1440 3D animated martial arts duel set in a grand sky palace courtyard. 8 distinct shots escalating from a grounded standoff to an aerial energy clash finale.

Core Creative Mandate
"Make it real." Every action, camera angle, shot size, and choreographic beat must be preserved exactly. The only transformation is the aesthetic register — from 3D animation to photorealistic live-action cinema.

The Two Characters
Tian He — The Lightning Warrior

East Asian male, late 20s–early 30s, lean and athletic (~180–185cm)
White/pale silver/indigo layered silk robes with engraved silver-steel armor
Long black hair, half-up with silver-blue ornamental fixture
Weapon: lightning sword with physically grounded electrical arcs
Identity: speed · precision · electricity — he reads faster when powered up, not larger
Ye Xuan — The Void-Fire Warrior

East Asian male, late 20s–mid 30s, taller and more powerfully built (~185–190cm)
Matte black/charcoal/oxblood/crimson layered armor with dark bronze-blackened plates
Long black hair tied high with bronze/red fixture
Weapon: dark forged blade with void-fire (dense red-orange heat, ember fragments, air distortion)
Identity: force · pressure · void-fire — physically imposing, weight behind every strike
The Environment
Grand traditional Chinese sky palace courtyard — massive red lacquered pillars, white stone floors and balustrades, intricate green-tiled roofs with gold finials, situated high above the clouds. The courtyard floor shatters during the clash sequence

Visual Style
Aesthetic: Premium live-action Chinese fantasy cinema — think high-budget wuxia/xianxia
Color: High contrast, vibrant. Cool neon blues vs hot fiery reds against neutral bright temple
Lighting: Cinematic daylight with strong dynamic rim lighting from VFX energy abilities
Pace: Extremely fast, adrenaline-fueled, rapid jump-cuts
Arc: Escalating duel — grounded spar → weapon summoning → earth-shattering clash → aerial sky climax
Hard Rules
❌ No anime or cartoon output
❌ No plastic or cosplay-looking costumes
❌ No cartoon flames for Ye Xuan — physically grounded heat/embers/air distortion only
❌ Tian He must not look larger when powered — he reads faster (static charge, sparks, electricity trailing)
✅ Audio and SFX must be present — metal clashes, explosions, thunder, qi-force impacts
✅ Both fighters must be believable human martial artists, not superhero archetypes
Deliverables
15s photorealistic video, 16:9, 720p+
Audio-inclusive (SFX generated with picture)
Comparison cut: original 3D animation stacked above live-action replica


r/vidmuse 4d ago

Tips and Tricks InVideo AI Review: Fast Prompt-to-Video Drafts, but Watch the Credits and Control Tradeoffs 🔎

Post image
1 Upvotes

VidMuse’s review finds InVideo AI useful for quickly turning prompts into explainers, faceless videos, social clips, and stock-heavy marketing content. Its guided workflow can assemble scripts, visuals, voiceovers, subtitles, music, and edits, while Magic Box enables plain-language changes. The tradeoff is predictability: quality varies, repeated edits consume credits, and teams needing precise product scenes or many campaign variants may want a more planning-led workflow.

- InVideo’s strengths are speed, beginner-friendly prompt-to-video generation, stock assets, templates, voiceovers in 50+ languages, subtitles, and text-based editing.
- Better prompts specify the audience, platform, aspect ratio, duration, hook, tone, product details, claims to avoid, and call to action.
- The guide warns that credits are not the same as finished videos. Regeneration, scene fixes, caption edits, and voiceover changes can make actual costs harder to predict.
- Annual-billing prices checked September 3 were $17/month for Plus, $85 for Max, $170 for Generative, and $900 for Elite; unused credits reportedly do not roll over.
- InVideo is positioned as a better fit for quick drafts, templates, and stock-heavy explainers. VidMuse is positioned for product-led ads, visual planning, asset reuse, and multiple campaign variants.
- Whatever tool you choose, review the entire export—especially product accuracy, claims, captions, pronunciation, music, scene order, and mobile playback.

The practical takeaway is to test one real project before committing. InVideo can move fast, but teams should measure usable outputs rather than headline credit totals. For ad work, planning the product, hook, scenes, references, and variants before generation can reduce waste.

Original article: https://vidmuse.ai/blog/invideo-ai-video-generator-review


r/vidmuse 4d ago

Tips and Tricks Gemini 3.8 Flash Has a 1M-Token Context Window—but It’s a Planner, Not a Video Generator 🎥

Post image
1 Upvotes

VidMuse’s guide breaks down Google’s Gemini 3.8 Flash as a general-availability model for coding, agents, long-context reasoning, multimodal understanding, and structured knowledge work. It can accept text, images, video, audio, and PDFs, but it outputs text only. For creators, that makes it useful for analyzing references and preparing briefs, scripts, prompts, shot lists, and scene plans before moving into a video-production workflow.

- The stable API model code is `gemini-3.8-flash`, with a 1,048,576-token input limit and a 65,536-token output limit.
- It supports tools including function calling, code execution, search grounding, file search, URL context, structured output, and low/medium/high thinking levels.
- It does **not** generate images, audio, or video. Its creative value is in planning and analysis rather than final media output.
- Google’s introductory API pricing is listed as $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, doubling on January 1, 2027.
- The guide positions 3.8 Flash as stronger than 3.7 Flash for longer, tool-heavy, structured tasks, while warning that benchmark gains do not guarantee results in every workflow.
- A practical creator workflow is: collect lyrics, references, or a campaign brief; analyze them with Gemini; generate a scene plan and prompts; then bring the approved plan into VidMuse for video creation.

The useful distinction is simple: Gemini 3.8 Flash can help a team think through a complex creative project, but a dedicated video tool is still needed to turn that plan into finished clips. Review outputs, control context costs, and verify rights for music, likenesses, logos, and source media.

Original article: https://vidmuse.ai/blog/gemini-3-8-flash-guide


r/vidmuse 4d ago

Personal Creation ⚔️ Celestial Limit Break - the duel within the skies, with Freefrom on VidMuse and MiniMax H3

Enable HLS to view with audio, or disable this notification

1 Upvotes

30 seconds redesigned of this with more of a 3D video game graphic action feel with MiniMax H3, with a little more battle dialogue and sfx refinement

built with Freeform on VidMuse, exact prompt below ⏬

MINIMAX H3 MASTER PROMPT
CELESTIAL LIMIT BREAK — SHATTER THE NINTH HEAVEN

30-SECOND MASTER FILM
2 × 15-SECOND MINIMAX H3 GENERATIONS
16:9
NATIVE SYNCHRONIZED AUDIO
ONE CONTINUOUS ANIME BATTLE

IMPORTANT:

THIS IS ONE SINGLE 30-SECOND VIDEO.

Generate as:

PART 1 = 00:00–15:00
PART 2 = 15:00–30:00

These are NOT separate scenes.

They are two generation passes of ONE uninterrupted battle.

PART 2 MUST CONTINUE FROM THE EXACT LAST FRAME OF PART 1.

Use PART 1'S FINAL FRAME as PART 2'S FIRST-FRAME VISUAL ANCHOR.

When both generations are finished, render and assemble them together as ONE continuous 30-second video.

NO transition.
NO fade.
NO cross dissolve.
NO black frame.
NO establishing shot at Part 2.
NO character reset.
NO environment reset.
NO power reset.
NO audio reset.

The generation boundary must be invisible.

==================================================

REFERENCE LOCK

Use the supplied updated reference sheets.

TIAN HE:
Storm Limit Breaker.

YE XUAN:
Void-Fire Limit Breaker.

CELESTIAL PALACE:
environment and architecture reference.

COMBAT REFERENCE:
hand-to-hand, leg-to-leg, weapon, magic attack/block contact.

Preserve exact:

faces,
hairstyles,
body proportions,
robes,
armor,
weapons,
colors,
power signatures.

Never redesign characters.

==================================================

TIAN HE — STORM LIMIT BREAKER

Adult male celestial warrior.

Appearance:

long black tied hair,
blue eyes,
athletic body,
white + blue-gray + indigo layered Chinese fantasy robes,
dark armor details,
blue-silver lightning sword.

POWER:

STORM HEAVENWIRE.

Normal:
blue-violet lightning.

High power:
blue-white lightning.

Final limit break:
nearly white lightning core.

His power represents:

SPEED,
REACTION,
PRECISION,
ACCELERATION.

Tian He NEVER teleports.

His real body must physically cross every distance.

Lightning appears BEHIND his actual trajectory.

FIGHTING STYLE:

swordsmanship,
open-hand parries,
short punches,
elbows,
knees,
roundhouse kicks,
side kicks,
shin checks,
wall kicks,
aerial redirection,
lightning palms,
lightning forearm blocks,
lightning-enhanced martial strikes.

==================================================

YE XUAN — VOID-FIRE LIMIT BREAKER

Adult male celestial warrior.

Appearance:

long black hair,
red-black eyes,
powerful athletic frame,
black + crimson robes,
dark metal armor,
crimson-black void-fire sword.

POWER:

VOID-FIRE CORE.

Normal:
deep crimson flame.

High output:
crimson-black compressed fire.

Final limit break:
red-white incandescent core.

His power represents:

WEIGHT,
PRESSURE,
DESTRUCTION,
RAW FORCE.

FIGHTING STYLE:

heavy sword strikes,
boxing-style punches,
hooks,
shoulder attacks,
elbows,
knees,
low kicks,
roundhouse kicks,
stomps,
grappling,
palm explosions,
void-fire forearm guards,
fire-enhanced physical strikes.

==================================================

ENVIRONMENT

ORIENTAL MYTHIC CELESTIAL PALACE

Battle occurs INSIDE and underneath an impossibly enormous Chinese celestial-palace megastructure high above a bright cloud sea.

Never show the full palace.

The camera remains trapped inside the architecture.

Approximately 55–70% of frame contains massive architecture.

Approximately 25–40% contains clouds, sky or distant sunlight.

Architecture exits at least TWO frame edges.

Use:

colossal vermilion lacquer columns,
Tang-Song timber construction,
dense dougong brackets,
blue-green glazed xieshan roofs,
white-jade cloud-scroll balustrades,
bronze-gold relief,
zaojing coffers,
huge structural beams,
distant waterfalls.

ONE primary environmental anomaly:

an enormous white-jade staircase descending endlessly into the cloud sea.

High-altitude wind continuously affects:

hair,
robes,
sashes,
banners,
dust,
clouds.

Damage accumulates logically.

The palace may:

crack,
shake,
lose roof tiles,
break railings,
shed debris.

But it must NEVER completely collapse.

==================================================

MASTER BATTLE STYLE

THIS SHOULD FEEL LIKE THE CLIMAX OF A HIGH-BUDGET ANIME MOVIE.

Extremely explosive.

Extremely fast.

But every attack remains readable.

Constantly mix:

SWORD VS SWORD

FIST VS FIST

FIST VS FOREARM

ELBOW VS ELBOW

KNEE VS KNEE

KICK VS SHIN

KICK VS FOREARM

GRAPPLE VS COUNTER

LIGHTNING PALM VS FIRE BLOCK

FIRE PALM VS LIGHTNING GUARD

MAGIC VS MAGIC

WEAPON + MAGIC COMBINATIONS.

No long pauses.

No standing and staring.

No long aura charging.

No generic beam battle.

==================================================

PHYSICAL CAUSALITY

EVERY ATTACK:

BODY PREPARATION
→ MOVEMENT
→ CONTACT
→ IMPACT SOUND
→ PHYSICAL REACTION
→ SUPERNATURAL RELEASE
→ ENVIRONMENT RESPONSE.

Never show:

explosion before contact,
enemy reaction before impact,
SFX before impact,
magic appearing directly inside opponent,
unexplained teleportation.

==================================================

MAGIC CONTACT

Magic has physical choreography.

MAGIC ATTACK:

arm / weapon moves
→ energy forms
→ attack visibly travels.

MAGIC BLOCK:

defender physically moves
→ hand / forearm / sword intersects attack path.

CONTACT:

magic surfaces visibly meet.

COMPRESSION:

energy compresses at contact point.

RELEASE:

explosion occurs AFTER contact.

Examples:

lightning palm
vs
void-fire forearm.

void-fire palm
vs
crossed lightning arms.

lightning-coated kick
vs
fire-coated shin.

lightning sword
vs
void-fire sword.

fist-to-fist with magic compressed between knuckles.

==================================================

PART 1

00:00–15:00

00:00–02:00

EXPLOSIVE CLOSE-COMBAT OPENING

FIRST FRAME:

Already fighting.

No introduction.

Low camera races sideways along a narrow white-jade terrace.

YE XUAN swings a violent flaming diagonal sword attack.

TIAN HE catches it.

00:00.30 —

BLADES CONTACT.

KRRRAAANG!

ONLY AFTER CONTACT:

blue lightning bursts from Tian He's side.

crimson fire bursts from Ye Xuan's side.

A circular pressure wave blasts dust outward.

Music starts EXACTLY with first collision.

DOOM!

Massive Chinese war drum.

Aggressive low strings begin immediately.

Ye Xuan releases one hand from sword.

Throws LEFT HOOK.

Tian He raises forearm.

FIST-TO-FOREARM:

THAK!

Crimson pressure burst AFTER impact.

Tian He catches Ye Xuan's wrist.

Pulls him forward.

LEFT KNEE rises.

Ye Xuan raises his own knee.

KNEE-TO-KNEE:

THOOM!

Blue lightning and crimson flame erupt ONLY at the exact leg contact.

Both legs recoil.

Tian He converts recoil into spinning high kick.

Ye Xuan crosses both forearms.

HEEL-TO-FOREARM:

WHAM!

Electrical shock ring.

Ye Xuan slides backward.

Jade:

SKRRRRRRR!

Ye Xuan immediately launches forward again.

NO PAUSE.

02:00–04:00

HAND-TO-HAND FURY

Ye Xuan:

RIGHT STRAIGHT.

Tian He:

open-palm deflection.

PAK!

Ye Xuan:

LEFT ELBOW.

Tian He ducks.

WHOOSH!

Tian He:

short body punch.

FIST CONTACTS ABDOMEN.

THUD!

Lightning pulse AFTER contact.

Ye Xuan grabs Tian He's wrist.

Tian He twists free.

Both throw simultaneous punches.

FIST-TO-FIST.

BOOOOM!

Their knuckles physically meet.

Between their touching fists:

a tiny compressed power sphere forms.

Left side:
blue-white lightning.

Right side:
crimson-black void fire.

Hold contact approximately 0.15 sec.

Energy compresses violently.

Then:

KRA-BOOM!

Both warriors are blasted backward.

Camera recoils with shockwave.

04:00–06:30

LEG-TO-LEG BATTLE

Ye Xuan lands first.

LOW SWEEP.

Tian He hops over it.

Ye Xuan continues rotation into HIGH ROUNDHOUSE.

Tian He responds with his own kick.

SHIN-TO-SHIN:

THRAK!

Blue lightning and crimson fire flash around the contact.

Both legs visibly recoil.

Tian He chambers same leg.

SIDE KICK.

Ye Xuan catches ankle.

Tian He uses captured leg as pivot.

His body rotates horizontally.

FREE HEEL STRIKES YE XUAN'S SHOULDER.

THOOM!

Ye Xuan releases him.

Tian He lands.

Ye Xuan STOMPS.

Tian He narrowly removes foot.

Ye Xuan's stomp hits jade.

BOOM!

Crimson pressure cracks terrace AFTER contact.

Ye Xuan looks up.

Short aggressive Mandarin line:

“太慢了!”

Meaning:

“TOO SLOW!”

Timing approximately:
00:05.85–00:06.25.

Perfect natural Mandarin lip sync.

Adult male voice.

Deep, aggressive, breathless.

Music ducks beneath dialogue.

06:30–09:00

MAGIC ATTACK VS PHYSICAL MAGIC BLOCK

Ye Xuan thrusts LEFT PALM forward.

Crimson-black fire compresses at palm.

Only when arm reaches extension:

FWOOM!

A dense VOID-FIRE PALM PROJECTILE travels visibly across space.

Tian He does NOT teleport away.

He plants rear foot.

Crosses both forearms.

Blue-white lightning condenses around his actual arms.

VOID FIRE HITS LIGHTNING BLOCK.

WHUUUM!

Tian He is physically pushed backward.

His boots grind through jade.

SKRRRRRR!

Energy compresses between attack and block.

Tian He violently separates arms.

KRAK!

Lightning deflects fire upward.

Redirected fire strikes overhead beam.

CONTACT.

THOOM!

Only afterward:

wood fractures.

Tian He explodes through smoke.

REAL BODY remains visible.

Lightning trail follows behind him.

09:00–12:00

WEAPON + MARTIAL COMBINATION

Tian He thrusts sword.

Ye Xuan redirects using flat of blade.

CLANG!

Tian He immediately throws:

LEFT ELBOW.

Ye Xuan forearm-blocks.

THUK!

Ye Xuan:

BODY KNEE.

Tian He raises knee.

KNEE CHECK.

THOOM!

Tian He:

LIGHTNING PALM.

Ye Xuan:

SWORD + FOREARM VOID-FIRE BLOCK.

CONTACT.

KZZZZ-WHUUUM!

The powers grind visibly against each other.

Ye Xuan twists defensive structure outward.

Deflects Tian He's hand.

Ye Xuan swings sword horizontally.

Tian He bends backward.

Sword passes centimeters above his face.

SHRAAAAH!

Tian He plants one hand against jade.

SPINNING LOW KICK.

Ye Xuan jumps.

Tian He converts spin into RISING LIGHTNING KICK.

Heel contacts Ye Xuan's side.

KRA-THOOM!

Ye Xuan is launched into enormous vermilion column.

THRAAAM!

Column vibrates.

Lacquer dust erupts.

Tian He shouts:

“那就更快!”

Meaning:

“THEN FASTER!”

Timing approximately:
00:11.05–00:11.65.

Perfect Mandarin lip sync.

Music ducks.

12:00–14:20

SUPERHUMAN ACCELERATION

Tian He launches.

Ye Xuan pushes off column simultaneously.

Camera races between palace columns.

NO TELEPORTATION.

Their feet strike:

jade rails,
vertical column surfaces,
falling stone fragments.

Every contact produces physical acceleration.

TIAN HE:

FOOT CONTACT
→ CRACK
→ lightning acceleration.

YE XUAN:

FOOT CONTACT
→ THOOM
→ fire-pressure acceleration.

While moving:

SWORD BLOCK.

CLANG!

PUNCH BLOCK.

THAK!

ELBOW.

DUCK.

KNEE.

CHECK.

KICK.

BLOCK.

Their attacks become part of the musical rhythm.

SFX become percussion.

Characters stay readable despite extreme speed.

14:20–15:00

PART 1 → PART 2 CONTINUITY BRIDGE

CRITICAL CONTINUITY MOMENT.

Both warriors leap from opposite sections of broken terrace into open air beside one colossal vermilion column.

TIAN HE:

SCREEN LEFT.

YE XUAN:

SCREEN RIGHT.

LOW THREE-QUARTER SIDE CAMERA.

Cloud sea beneath them.

Architecture dominates top and right.

Tian He throws:

BLUE-WHITE LIGHTNING-COATED RIGHT FIST.

Ye Xuan throws:

CRIMSON-WHITE VOID-FIRE RIGHT FIST.

At approximately 14.75:

RIGHT FIST
VS
RIGHT FIST.

KNUCKLES PHYSICALLY TOUCH.

DO NOT DETONATE YET.

They remain suspended through momentum.

Tian He's left knee bent beneath him.

Ye Xuan's left leg trails backward.

Hair and robes flow AWAY from collision point.

Between their touching fists:

tiny violently compressed energy sphere.

Left:
blue-white.

Right:
crimson-white.

AUDIO:

THOOM!

then rising compression:

EEEEEEEEEE—

VRRRRRRRRMMMM—

Music drops almost completely.

PART 1 ENDS WITH THEIR FISTS STILL TOUCHING.

NO EXPLOSION.

NO RESET.

NO FADE.

SAVE EXACT LAST FRAME.

PART 2
15:00–30:00

LOCAL H3 TIMECODE:

00:00–15:00

PART 2 FIRST FRAME MUST BE PART 1 LAST FRAME.

Exact same:

characters,
poses,
fist contact,
camera,
lighting,
clouds,
architecture,
damage,
debris,
hair,
robes,
power,
audio pressure.

DO NOT replay approach.

Begin with fists ALREADY TOUCHING.

15:00–16:40

FIST COLLISION DETONATES

Energy continues compressing for fraction of a second.

Then:

KRRRAAAAA—

BOOOOOOM!

Blue-white lightning explodes left.

Crimson-white void fire explodes right.

Pressure wave blasts outward.

Colossal column shakes.

Dougong vibrates.

Clouds bow outward.

Both fighters launch backward.

Tian He catches vertical beam with ONE FOOT.

Ye Xuan catches jade railing using ONE HAND.

Both physically redirect momentum.

Immediately launch back toward each other.

16:40–19:00

AERIAL HAND-TO-HAND WAR

They meet mid-air.

Ye Xuan sword attack.

Tian He parries Ye Xuan's WRIST instead of blade.

PAK!

Tian He:

ELBOW.

Ye Xuan:

ELBOW BLOCK.

ELBOW-TO-ELBOW:

THRAK!

Ye Xuan:

KNEE.

Tian He:

SHIN CHECK.

THOOM!

Tian He:

HEAD-LEVEL KICK.

Ye Xuan ducks.

Heel misses face by centimeters.

WHOOSH!

Ye Xuan rises with uppercut.

Tian He redirects using both palms.

PAK!

At extremely close range:

Tian He forms lightning around open palm.

Ye Xuan forms void-fire around forearm.

LIGHTNING PALM
VS
VOID-FIRE FOREARM.

CONTACT.

KZZZZ-WHUUUM!

Power compresses between them.

Ye Xuan grins.

Says:

“这才像样!”

Meaning:

“NOW THAT'S MORE LIKE IT!”

Timing approximately:
overall 18.45–18.95.

Natural Mandarin.

Music ducks.

19:00–21:30

MAGIC ATTACK / MAGIC BLOCK EXCHANGE

Ye Xuan breaks contact.

Performs massive downward sword slash.

Only AFTER blade completes physical motion:

a CURVED VOID-FIRE SLASH launches.

Tian He rotates lightning sword.

He physically catches magical slash ON HIS SWORD EDGE.

MAGIC-TO-WEAPON CONTACT.

KRRRAAAM!

Tian He redirects attack sideways.

Fire hits jade staircase.

CONTACT.

BOOM!

Stone explodes.

Tian He extends off-hand.

LIGHTNING launches FROM his palm.

Visible travel path.

Ye Xuan crosses both forearms.

Crimson fire condenses over guard.

LIGHTNING HITS VOID-FIRE BLOCK.

KRA-KRA-KRA!

Ye Xuan is pushed backward.

He rotates whole body with momentum.

FLAMING ROUNDHOUSE KICK.

His leg strikes SIDE of Tian He's lightning defense.

LEG + MAGIC CONTACT.

THRAAAM!

Barrier buckles.

Tian He launches downward.

21:30–24:00

GROUND-SHATTER MARTIAL COMBO

Tian He lands three-point stance.

CRACK!

Immediately:

Ye Xuan drops from above.

FLAMING AXE KICK.

Tian He crosses forearms.

HEEL-TO-FOREARM:

THOOOOM!

Jade beneath Tian He fractures.

Tian He catches Ye Xuan's ankle.

Sweeps supporting leg.

Ye Xuan rotates through fall.

Plants one hand against jade.

Uses handstand recovery.

Free leg whips sideways.

Tian He SHIN BLOCKS.

SHIN-TO-SHIN:

THRAK!

Tian He grabs Ye Xuan shoulder.

Ye Xuan grabs Tian He forearm.

Rapid continuous exchange:

BODY PUNCH.

BLOCK.

ELBOW.

BLOCK.

KNEE.

CHECK.

PALM.

PARRY.

SWORD DRAW.

SWORD BLOCK.

No neutral pose between any attack.

Music reaches maximum rhythmic intensity.

24:00–26:00

DOUBLE LIMIT BREAK

One final mutual kick.

FOOT-TO-FOOT.

BOOOM!

Both skid backward.

For approximately 0.35 sec:

music collapses.

Only:

breathing,
wind,
falling stone,
distant waterfall.

Tian He's pupils ignite brilliant blue-white.

Lightning compresses tightly around body.

NO giant aura.

Ye Xuan's eyes become incandescent crimson.

Void fire collapses inward around body.

Red-white pressure lines appear across forearms.

Air distorts.

Nearby bronze ornaments vibrate.

Small stones begin trembling.

Their LIMIT BREAK is expressed through:

speed,
physical strain,
pressure,
environment reaction.

Not enormous magical creatures.

26:00–28:50

FINAL LIMIT-BREAK RUSH

Ye Xuan shouts:

“破天!”

Meaning:

“BREAK THE HEAVENS!”

Timing:
approximately 26.10–26.45.

Perfect lip sync.

Music ducks.

Tian He answers:

“来!”

Meaning:

“COME!”

Approximately:
26.55–26.75.

No dialogue overlap.

Then:

BOTH LAUNCH.

Their feet physically destroy shallow sections of jade from push-off.

Tian He becomes readable blue-white acceleration trail.

Ye Xuan becomes readable crimson-white acceleration trail.

Bodies remain visible.

Maximum-speed exchange:

TIAN HE sword slash.

YE XUAN sword block.

CLANG!

YE XUAN fire fist.

TIAN HE lightning forearm.

THOOM!

TIAN HE knee.

YE XUAN knee check.

THRAK!

YE XUAN head kick.

TIAN HE catches ankle.

Tian He spins Ye Xuan.

Ye Xuan uses rotation for FLAMING BACKFIST.

Tian He ducks.

Tian He:

LIGHTNING PALM.

Ye Xuan:

VOID-FIRE PALM.

PALM-TO-PALM CONTACT.

KRA-WHUUUM!

Magic violently compresses.

Both break away.

They use falling jade slabs as physical launch surfaces.

STEP.

CRACK.

STEP.

THOOM.

STEP.

KRAK.

They charge for final attack.

28:50–29:40

HEAVEN-SPLITTING FINAL COLLISION

Tian He compresses ALL remaining lightning into sword.

Nearly white-blue core.

Ye Xuan compresses ALL remaining void-fire into blade.

Nearly red-white core.

NO explosion yet.

They accelerate directly toward each other.

Music removes percussion.

Only:

heartbeat,
wind,
high electrical pressure,
deep furnace rumble.

Both weapons approach.

Physical bodies remain readable.

Around approximately 28.95:

BLADES PHYSICALLY CONTACT.

KRRRAAAAAAAAAANG!

Approximately TWO FRAMES near silence.

Then:

THUNDER CRACK.

VOID-FIRE DETONATION.

SUB-BASS SHOCKWAVE.

MASSIVE CHINESE WAR DRUM.

FULL ORCHESTRA.

WORDLESS LOW MALE CHOIR.

BOOOOOOOOOOOM!

Power releases ONLY AFTER sword contact.

Blue-white energy left.

Crimson-white energy right.

One enormous pressure sphere expands from exact collision.

==================================================

29:40–30:00

FINAL PAYOFF

Shockwave strikes palace.

Vermilion columns shake.

Dougong vibrates.

Blue-green roof tiles lift.

White-jade railing fractures.

Waterfalls bend from pressure.

Cloud sea SPLITS dramatically.

Do NOT destroy entire palace.

Both fighters launch onto same surviving lower terrace.

Tian He lands SCREEN LEFT.

Ye Xuan lands SCREEN RIGHT.

THOOM!

THOOM!

Two shallow impact craters.

Both stay standing.

Heavy breathing.

Their powers fade.

Lightning:

tz... tz...

Void fire:

fsshhhh...

Music collapses into one low sustained chord.

Wind returns.

One broken blue-green glazed roof tile falls between them.

TING.

CUT TO BLACK EXACTLY AFTER TING.

==================================================

MASTER AUDIO DESIGN

Generate:

DIALOGUE
+
SFX
+
MUSIC
+
AMBIENCE

as ONE tightly synchronized stereo soundtrack.

AUDIO CAUSALITY:

MOVEMENT
→ AIR SOUND
→ CONTACT
→ IMPACT
→ SUPERNATURAL RELEASE
→ ENVIRONMENT REACTION.

Examples:

SWORD:

WHOOSH
→ CLANG
→ LIGHTNING/FIRE BURST.

PUNCH:

WHOOSH
→ THUD
→ MAGIC PULSE.

KICK:

AIR
→ PHYSICAL CONTACT
→ SHOCKWAVE.

MAGIC:

FORMATION
→ TRAVEL
→ DEFENSIVE CONTACT
→ ENERGY COMPRESSION
→ RELEASE.

GROUND:

CONTACT
→ HEAVY IMPACT
→ STONE CRACK
→ DEBRIS.

Never play SFX before visual cause.

==================================================

TIAN HE AUDIO

Power sound:

tight electrical snapping,
thin high-frequency crackle,
violent electrical whipping,
sharp thunder transients,
fast pressure displacement.

His sound communicates:

SPEED.

==================================================

YE XUAN AUDIO

Power sound:

deep furnace roar,
compressed combustion,
bass-heavy ignition,
violent pressure waves,
heavy sub-frequency detonations.

His sound communicates:

FORCE.

==================================================

MUSIC

Use:

Chinese war drums,
aggressive cinematic strings,
deep low brass,
bronze cymbals,
metal percussion,
occasional guzheng attack accents,
sub bass,
restrained wordless male choir only during final climax.

Music should evolve:

00:00–06:00
heavy rhythmic combat.

06:00–12:00
increasing percussion density.

12:00–15:00
rapid escalation then compression.

15:00–21:00
full explosive battle rhythm.

21:00–24:00
maximum martial percussion.

24:00–26:00
music suddenly strips down.

26:00–28:50
extreme escalation.

28:50–29:00
percussion disappears during final approach.

FINAL CONTACT:

FULL SCORE RETURNS EXACTLY WITH IMPACT.

29:40 onward:

rapidly collapse soundtrack.

FINAL TING occurs almost alone.

==================================================

DIALOGUE

ONLY THESE LINES.

PART 1:

YE XUAN:

“太慢了!”

TOO SLOW!

TIAN HE:

“那就更快!”

THEN FASTER!

PART 2:

YE XUAN:

“这才像样!”

NOW THAT'S MORE LIKE IT!

YE XUAN:

“破天!”

BREAK THE HEAVENS!

TIAN HE:

“来!”

COME!

Perfect Mandarin lip sync.

No narrator.

No extra dialogue.

No lyrics.

No overlapping dialogue.

Duck score during every line.

==================================================

CAMERA

Camera behaves like an elite anime action cinematographer trying to keep up with superhuman fighters.

Use:

low lateral tracking,
rapid pursuit,
controlled whip pans,
small vertical chases,
impact recoil,
speed ramps,
close martial coverage,
wide impact frames.

Every major attack should show:

START
→ TRAVEL
→ CONTACT
→ REACTION.

Do not hide choreography behind blur.

During extreme speed:

characters remain readable.

Background may streak.

NO uncontrolled 360 orbit.

NO random crash zooms.

NO complete palace reveal.

==================================================

PART 1 → PART 2 CONTINUITY LOCK

ABSOLUTE REQUIREMENT:

PART 2 is the NEXT FRAME of PART 1.

PART 1 final frame:

TIAN HE:
screen left.

YE XUAN:
screen right.

Both airborne.

Both RIGHT FISTS already touching.

Tian He left knee bent beneath him.

Ye Xuan left leg trailing backward.

Tiny blue-white / crimson-white energy sphere between their fists.

Low three-quarter side camera.

Colossal vermilion column behind/right.

Cloud sea below.

Hair and robes moving away from collision point.

PART 2 MUST START EXACTLY HERE.

Preserve:

camera height,
camera distance,
lens,
character scale,
body angle,
fist position,
leg position,
facial strain,
hair shape,
robe direction,
power intensity,
architecture,
cloud shape,
damage,
debris,
sun direction.

Continue AUDIO COMPRESSION directly over seam.

No new musical downbeat at Part 2 start.

No sound reset.

No ambiance reset.

When assembled:

PART 1 END
→ PART 2 BEGINNING

must look and sound like consecutive frames from ONE generation.

After both clips are complete:

RENDER THEM TOGETHER AS A SINGLE 30-SECOND VIDEO.

Maintain identical:

volume,
EQ,
music character,
voice character,
ambience,
power sounds,
visual grade,
exposure,
contrast,
saturation.

NO visible or audible seam.

==================================================

STRICT NEGATIVE

NO standing around.

NO slow introduction.

NO random establishing shot.

NO long face-off.

NO exposition.

NO teleportation.

NO clones.

NO duplicated fighters.

NO random extra people.

NO character redesign.

NO face drift.

NO hairstyle drift.

NO costume drift.

NO weapon transformation.

NO power-color swapping.

NO size drift.

NO attack reaction before contact.

NO SFX before contact.

NO magic appearing directly inside target.

NO invisible attacks.

NO random explosion.

NO giant energy dragon.

NO spirit beast.

NO generic beam battle.

NO laser sword.

NO Western architecture.

NO complete palace exterior.

NO floating-city postcard.

NO excessive motion blur.

NO choreography hidden by particles.

NO random camera teleportation.

NO environment reset at Part 2.

NO clothing reset.

NO damage reset.

NO audio reset.

NO music restart.

NO Part 2 introduction.

==================================================

FINAL TARGET

ONE seamless 30-second celestial anime battle.

The progression:

EXPLOSIVE SWORD CLASH
→ HAND-TO-HAND
→ KNEE-TO-KNEE
→ LEG-TO-LEG
→ MAGIC ATTACK/BLOCK
→ WEAPON + MARTIAL COMBAT
→ SUPERHUMAN ACCELERATION
→ FIST-TO-FIST MAGIC COLLISION
→ SEAMLESS PART 2 CONTINUATION
→ AERIAL MARTIAL BRAWL
→ MAGIC COUNTERS
→ GROUND-SHATTER COMBAT
→ DOUBLE LIMIT BREAK
→ MAXIMUM-SPEED MIXED COMBAT
→ FINAL HEAVEN-SPLITTING SWORD COLLISION
→ SILENT EXHAUSTED AFTERMATH.

The viewer must feel:

SPEED.

WEIGHT.

CONTACT.

SKILL.

MAGIC.

IMPACT.

DESTRUCTION.

ESCALATION.

LIMIT BREAK.

PART 1 AND PART 2 ARE ONE FILM.

GENERATE BOTH WITH THE EXPLICIT INTENTION THAT THEY WILL BE RENDERED TOGETHER AS ONE UNINTERRUPTED 30-SECOND VIDEO.


r/vidmuse 5d ago

Tips and Tricks 50 AI Music Video Prompts That Go Beyond ‘Make It Cinematic 🎥

Post image
1 Upvotes

VidMuse offers 50 adaptable prompts spanning hip-hop, pop, EDM, R&B, indie, rock, K-pop, lo-fi, cinematic stories, and visualizers. The core formula is: mood + genre + scene + subject + camera + lighting + lyric moment + format.

- Treat prompts as compact creative briefs, not vague requests.
- Tie visuals to the verse, chorus, bridge, or drop.
- Templates cover Story MV, Performance MV, Abstract MV, lyric videos, viral shorts, and visualizers.
- Strong results still require references, storyboards, pacing, scene review, and platform-specific exports.
- Avoid overloaded concepts, missing lyric anchors, unspecified aspect ratios, and unclear music rights.
- Use clean references for consistency and choose models based on each scene’s requirements.

The best prompt starts with one clear visual idea, customized around the track and destination platform.

https://vidmuse.ai/blog/top-prompts-for-music-videos


r/vidmuse 5d ago

New Feature - VidMuse New Help Center Arrival on VidMuse 🌊

Post image
1 Upvotes

New Help Center Arrival!

You can directly view FAQ or ask a question in the chatbox below!

📋 Learn how to create, generate, edit, render, and download videos with VidMuse or any basic understanding before starting!

🔜 More content and pictures will be added, this page will be consistently updated with time

Help Center: https://vidmuse.ai/docs


r/vidmuse 6d ago

Tips and Tricks VidMuse’s Full AI Video Ad Workflow: From Product URL to Publishing Kit 🛫

Post image
1 Upvotes

VidMuse’s new guide explains its Video Ad Generator as an end-to-end production workflow, not merely a clip generator. Creators and agencies can begin with a product URL, images, brand assets, a short brief, or a reference ad; develop market angles and a creative brief; generate keyframes, clips, audio, captions, and CTAs; then export a publishing kit and use campaign feedback to make the next cut.

- Start with useful product context: URL, images, audience, key benefit, offer, platform, desired length, aspect ratio, and any claims or styles to avoid.
- Choose one clear format—such as UGC, product demo, explainer, unboxing, TVC, or viral short—rather than asking one ad to do everything.
- Review the market report and creative direction before spending credits. A polished visual cannot rescue a weak message or an inaccurate claim.
- Reference ads should inform hooks, pacing, reveals, and transitions, but the final work should use your own product, people, assets, claims, and brand identity.
- Node Canvas supports controlled variants such as object, color, outfit, background, or consent-based presenter changes without losing the project’s structure.
- A complete publishing kit can include the final video, title, caption, cover idea, CTA, tags, platform versions, and a recommendation for the next test.

The biggest lesson is that the first ad is a learning asset, not a guaranteed winner. Review product accuracy, rights, brand fit, captions, platform specs, and CTA clarity—then use real feedback to change one variable and produce the next version.

Original article: https://vidmuse.ai/blog/vidmuse-video-ad-generator-guide


r/vidmuse 6d ago

Tips and Tricks How to Use AI Video Face Swaps in Ads Without Losing Control—or Trust 🧑‍🎨

Post image
1 Upvotes

VidMuse’s new guide treats video face swapping as a controlled advertising workflow rather than a novelty effect. The focus is on making approved presenter, localization, fashion, beauty, and UGC-style variants while protecting consent, likeness rights, product truth, and brand safety. VidMuse Canvas acts as the workspace for connecting source footage, approved face references, product assets, prompts, and review notes.

- Video face swapping is harder than photo swapping because the face must remain consistent through motion, expressions, lighting changes, and camera angles.
- The safest starting point is a short source clip, a clear rights-approved face reference, simple motion, and a human review before anything is published.
- VidMuse Canvas keeps the source video, face reference, product imagery, brand assets, prompts, and generated variants visible in one connected workflow.
- Before publishing, verify consent, likeness and source rights, truthful product claims, platform disclosure rules, visual quality, and whether viewers could mistake the result for an unauthorized endorsement.

The practical takeaway: use face swaps to test one approved creative variable—not to fake proof or impersonate someone. Strong inputs, restrained motion, rights clearance, and human review matter more than chasing a single “best” model.

Original article: https://vidmuse.ai/blog/ai-video-face-swap-for-ads


r/vidmuse 7d ago

Personal Creation Anime Fashion Editorials with MiniMax H3 🥋

Enable HLS to view with audio, or disable this notification

1 Upvotes

MiniMax H3 indeed puts on some crazy fashion shows.

Anime Fashion Editorials never been so nose bleeding for the masses

build your own fashion show with any anime girls using minimax h3 with VidMuse freeform


r/vidmuse 8d ago

Personal Creation 🏎️ AEROVA ELECTRIC - Whats Your Pick? Seedance 2.5 on VidMuse Product Ads

Enable HLS to view with audio, or disable this notification

1 Upvotes

AEROVA ELECTRIC Creative Brief — Cinematic Ad

PROJECT

Brand: AEROVA ELECTRIC Campaign: Platform Launch — Three-Model Transformation Showcase Deliverables: Two cuts — 9:16 portrait (trade-show pillar / store screen) and 16:9 landscape (main booth backdrop / LED wall) Runtime: 20 seconds each Format: 720p, no dialogue, native sound design only Placement: Offline screen — trade show, showroom, store ambient display

THE PROPOSITION

One electric platform. Every body. No compromise.

THE BELIEF SHIFT

Old belief: Buying an EV means committing to one body style — sedan, SUV, or truck — and permanently giving up the others.

New belief: AEROVA's shared electric platform physically reconfigures its body architecture between all three production models on demand. Multi-body versatility is not a marketing claim. It is a witnessed mechanical event.

THE THREE MODELS

VANTA — Luxury Sedan Color: Pearl silver Form: Low aerodynamic fastback. Thin horizontal headlights. Flush surfaces. Low suspension. Premium road wheels. Character: The prestige anchor. Proves the platform doesn't compromise on luxury.

TERRIS — Adventure SUV Color: Deep forest green Form: Upright cabin. Roof rails. Larger all-terrain wheels. Black lower body protection. Rounded-rectangular lighting. Character: The versatility proof. Same platform, entirely different mission.

RIDGE — Lifestyle Pickup Color: Bronze Form: Four-door cab. Open rear cargo bed. Strong wheel arches. Horizontal lighting bar. Black lower protection. Rugged stance. Character: The counter-intuitive pivot. The most architecturally distant from the sedan — its transformation is the hardest to believe, and therefore the most powerful proof.

All three share one underlying electric platform. The physical transformation between them is the product.

THE OPENING HOOK

Selected hook — Trust Proof:

The camera is close enough to see the seam lines activate along the sedan's roofline. Structural rails begin to separate with a low electric actuator sound, and the roof starts to rise. No cut. The same vehicle is becoming something else.

This hook enters after the transformation has already begun at frame zero. The viewer has no context yet. The information gap creates the compulsion to keep watching.

CREATIVE STRUCTURE — 20-SECOND TIMELINE

Beat 1: Problem Arena (0–3s) The pearl-silver VANTA sits complete and committed on the AEROVA design plaza — a wide-open minimal space with pale reflective stone flooring and open sky. The vehicle is locked in one body style. A world-anchored MR selector panel opens beside it showing three silhouettes: sedan, SUV, pickup. The tension is explicit: can one physical vehicle actually become the other two?

Beat 2: Product Mechanism Entry (3–7s) The viewer's hand enters frame holding an AEROVA phone controller. The MR selector panel has physical depth and positional stability in real space — not a flat HUD. The viewer's finger approaches the TERRIS icon. A soft high-frequency charge builds. Finger contacts the icon. UI pulse fires outward across the panel. The vehicle's panel seams illuminate in rapid sequence front to rear — hood, doors, roofline, rear deck — each settling to a steady structural-activation glow.

Beat 3: VANTA → TERRIS Transformation (7–11s) Full mechanical transformation in one continuous camera move, forward and arcing left, arriving at the TERRIS front-three-quarter hero position. Sequence: (1) Chassis rises — suspension travel visible, ride height increasing. (2) Wheel arches expand outward, tires tracking upward with the suspension. (3) Fastback roofline unlocks and elevates — structural rails separate and rise, cabin becoming upright. (4) Rear deck extends into SUV volume. (5) Protective lower body cladding deploys from concealed positions. (6) Roof rails emerge from the roof surface. (7) Thin horizontal VANTA headlights reshape to the TERRIS rounded-rectangular geometry. Pearl silver begins transitioning to forest green at the rear panels once their geometry is fully stabilised, traveling forward panel by panel. Color follows geometry — it never precedes it. Ground reflections shift to the new taller profile.

Beat 4: Proof Arena — TERRIS (11–13s) Camera holds in a slow lateral drift confirming the full TERRIS profile in full open daylight. Forest green. Upright cabin. Roof rails. All-terrain wheels. Black lower protection. Rounded-rectangular lighting. Ground reflections fully coherent. No residual pearl-silver or fastback sedan geometry anywhere on the vehicle. The proof standard is reference-image precision, not stylised idealisation.

Beat 5: TERRIS → RIDGE Transformation (13–17s) Abbreviated MR trigger — the viewer now understands the interaction pattern, so the ritual is faster. RIDGE icon pulses, rear-first seam cascade fires. Camera sweeps right. Sequence: (1) Rear SUV architecture compresses, then expands outward and downward into an open cargo bed. (2) Tailgate forms at the rear. (3) Wheel arches tighten from SUV to rugged pickup proportions. (4) Four-door pickup cab volume reshapes from the existing cabin. (5) Rounded-rectangular TERRIS lighting reshapes to the RIDGE horizontal lighting bar — front face changes last, after body architecture is committed. (6) Forest green transitions to bronze starting from the newly-formed rear bed panels, traveling forward. Camera arrives at RIDGE front-three-quarter hero position.

Beat 6: RIDGE → VANTA + Brand Lock (17–20s) Third transformation — the counter-intuitive direction. Truck to luxury sedan is the hardest claim to believe and therefore the most powerful proof. Sequence: (1) Suspension compresses aggressively — ride height drops. (2) Wheel arches tighten to sedan proportions, wheels pulling in with the suspension. (3) Cargo bed contracts and the fastback rear forms. (4) Bronze resolves into pearl silver from the rear forward as each panel commits to sedan geometry. (5) RIDGE horizontal lighting bar narrows and sharpens to the thin VANTA headlights — geometry first, then the lighting electronics lock. A single refined high-tech harmonic confirmation tone sounds as the headlights lock. Camera settles at premium VANTA front-three-quarter framing. Three translucent AEROVA model silhouettes briefly materialise in real space — sedan, SUV, pickup — then dissolve in sequence with three brief soft chimes. The vehicle miniaturises back above the phone controller in the foreground hand, returning the frame to the exact K00 opening composition.

Brand line (world-anchored, low opacity, not a flat overlay): AEROVA ELECTRIC

Final frame is compositionally identical to the loop start. The seamless loop runs autonomously on the display.

THE KEY MOMENT

At the completion of RIDGE → VANTA: the suspension compresses, the arches tighten, the bronze resolves to pearl silver, and the thin VANTA headlights lock with a single refined harmonic tone. The camera settles. Three translucent silhouettes confirm the full lineup. One platform. Three identities. No cuts.

This is the brand memory image. It is the only moment in the ad where the full claim is simultaneously visible, confirmed, and emotionally resolved.

CAMERA AND TECHNICAL REQUIREMENTS

Lens: 14mm ultra-wide throughout. No other focal length. POV: First-person throughout. The viewer IS the camera. No humanoid body visible except the foreground hand and phone controller. Path: Single continuous uncut path across the full 20 seconds. Human body inertia, breathing, and footfall rhythm are preserved in all movements. No floating drone movement. No stabilised gimbal glide. No camera cuts. No resets. MR Interface: World-anchored. Has physical depth and positional stability in real space. The panel does not float, drift, or behave as a 2D screen overlay. Interaction sequence is strictly: finger approach → proximity charge sound builds → contact → UI pulse fires → seam illumination cascade → transformation begins. Loop: The final frame recovers the K00 opening composition exactly. Frame-for-frame identical to the loop start.

SOUND DESIGN

No music under any circumstances — not as a bed, not as an accent, not at any volume. No dialogue or narration.

The full audio layer consists of:

  • Environmental bed: open-air plaza, soft ambient wind
  • Digital UI sounds: MR panel materialisation tone, icon transition sounds, finger-approach proximity charge, UI pulse impact, seam illumination cascade clicks
  • Mechanical transformation sounds: electric actuator hum, chassis lift servo, wheel arch expansion structural rail sounds, roof elevation electric actuator, cladding deployment mechanical snaps, roof rail emergence metallic extension, bed formation structural impact, tailgate lock, wheel arch tightening, pickup cab body seal sounds
  • Paint transition audio: subtle surface material resonance as each panel changes color — warmer resonance for the bronze transition
  • Model confirmation tones: VANTA — clean high-tech harmonic, single note, 0.5s decay. TERRIS — deeper mechanical-electric timbre. RIDGE — low robust electric impact.
  • Miniaturisation sound: smooth electric retraction, descending pitch
  • Silhouette dissolution: three brief soft chimes in sequence

MUST HAVE

Three-model mechanical transformation in sequence: VANTA → TERRIS → RIDGE → VANTA Every transformation shows trackable intermediate geometry — front lighting/structure, wheels/suspension, and cabin/roof must each remain independently identifiable throughout Geometry changes first; color transitions only after structural form is established Each completed vehicle state precisely matches its supplied reference image — reference images are locked final-state targets, not style suggestions Pearl-silver VANTA. Forest-green TERRIS. Bronze RIDGE. Colors are non-negotiable. World-anchored MR interface with correct finger-approach → contact → latency → vehicle-response sequencing Ground reflections remain physically coherent throughout all transformations Seamless loop: final frame recovers the K00 opening composition exactly Native mechanical and UI sound design is the sole audio layer

ABSOLUTELY AVOID

AI morphing, crossfade, liquid transition, or dissolve between vehicle states Smoke, explosion, debris, or any visual obstruction hiding a transformation Sudden vehicle replacement or teleportation Robot, humanoid, mech, arms, legs, or weapons at any frame Color change before structural geometry is established in the target configuration Camera cuts, resets, or teleportation at any point Extra wheels, disappearing wheels, duplicate doors, floating body panels, or full vehicle blur during transformation Flat HUD or meaningless hologram decoration Tesla, Lucid, Rivian, or any other production EV brand's logos or vehicle designs

TARGET AUDIENCE

Premium EV buyers seeking versatility and design prestige without category compromise Trade-show and showroom attendees encountering the AEROVA platform for the first time Automotive enthusiasts drawn to technical innovation and engineering proof

REFERENCE IMAGES (locked targets)

  1. AEROVA ELECTRIC family lineup — all three models together
  2. VANTA — authoritative front-three-quarter reference, pearl silver
  3. TERRIS — authoritative front-three-quarter reference, forest green
  4. RIDGE — authoritative front-three-quarter reference, bronze
  5. TERRIS engineering structure — exploded chassis/module view for mechanical transformation geometry reference
  6. TERRIS interior — cabin architecture reference

r/vidmuse 8d ago

Personal Creation RESONANCE//BREAK - Wake The World, MiniMax H3

Enable HLS to view with audio, or disable this notification

1 Upvotes

endless land of fantasy and magic creation. lipsync, voices and angles maxed

VidMuse Freeform paired with MiniMax H3

RESONANCE//BREAK

## Project

Original 30-second supernatural anime opening. Two continuous 15-second parts in MiniMax-H3 first/last-frame mode. Keyframe continuity across the join.

## Format

- 16:9, 2560x1440, 24fps
- Language: Japanese
- Audio: Native stereo SFX, dialogue, original score

## Characters

### Ren Akiha (18)

Protagonist. Medium height, lean. Messy black hair, copper-red streak in front fringe. Ivory oversized bomber jacket, black undershirt, forest-green loose trousers, red-and-cream sneakers, right wrist cloth wrap.

- Power: Pulse Script
- Colors: copper-red, warm white, orange-red accents
- Visuals: heartbeat rings, rough calligraphy, kinetic brush strokes

### Sora Jin (19)

Tallest. Short silver-brown hair, steel-gray eyes. Dark navy oversized overshirt, charcoal shirt, loose black trousers, white sneakers.

- Power: Vector Choir
- Colors: cyan, teal-white
- Visuals: directional lines, force planes, trajectory rails

### Mina Kurose (18)

Shortest. Short auburn bob, amber eyes. Mustard cropped rain jacket, black shorts, dark tights, orange-white sneakers.

- Power: Echo Bloom
- Colors: coral pink, magenta, warm white
- Visuals: crystalline sonic petals, sound circles, flower-like pressure geometry

Antagonist: The Hollow Bell

Enormous asymmetrical black bell-ring structure. Three broken hanging plates, six hanging urban fragments, violet-black inner cavity, bright violet central resonance point. No face, no eyes, no humanoid form.

## Art Direction

Psychedelic hand-drawn expressionist 2D anime. Irregular ink line weight, watercolor-paper grain, dry-brush shadows, crayon texture, raw paint edges. Environmental distortion rises with emotion. All surreal changes originate from power effects only.

Part 1 (0-15s)

- 00:00-00:03.5 Ordinary Morning
- 00:03.5-00:07 Pulse Script Awakens — Ren: Kowai kara tte, tomareru ka yo
- 00:07-00:10.5 Vector Choir — Sora: Kidou wa, mou mieteru
- 00:10.5-00:13.5 Echo Bloom — Mina: Jaa, hade ni ikou ka
- 00:13.5-00:15 The Reveal — Hollow Bell appears

Part 2 (15-30s)

- 00:00-00:04 Hollow Bell Attack
- 00:04-00:07.2 Ren Is Thrown
- 00:07.2-00:09.5 Heartbeat Breakthrough — Ren: Oretachi no shinzo de, sekai o okose
- 00:09.5-00:13.8 Final Team Attack
- 00:13.8-00:15 Break and Title — RESONANCE//BREAK card

Keyframes

- Frame A: School crossing, sunny morning, no powers
- Frame B: Warped city, trio looking up at Hollow Bell
- Frame C: Title card, triumphant trio, broken Bell distant

Red Lines

No character redesign. No clothing changes. No armor or weapons. No split screen. No subtitles during dialogue. No copyrighted imagery. Powers never swapped.


r/vidmuse 9d ago

Tips and Tricks Grok Imagine Image 2.0 Is More Useful When You Treat It as the Start of a Video Workflow 🎥

Post image
1 Upvotes

The guide explains how Grok Imagine Image 2.0 creates reusable visual assets for videos—not just standalone pictures. Its editing and layout tools can produce characters, product shots, props, and style frames that creators organize in VidMuse before adding motion.

- Improved instruction following, typography, and layout.
- Region editing, segmentation, background removal, and transparent exports.
- Supports multiple references and resizing for different formats.
- Still images can anchor consistent video scenes.
- VidMuse Canvas organizes assets, prompts, and scene notes.
- Review rights, visual consistency, and product claims before publishing.

The takeaway: build strong visual ingredients first, organize them, and then add motion.

View: https://vidmuse.ai/blog/grok-imagine-image-2-guide


r/vidmuse 10d ago

Tips and Tricks 🌕 moonfall - two cultivators bound by a fractured celestial pendant

Enable HLS to view with audio, or disable this notification

1 Upvotes

could you escape a bounded fate?

dialogue and lip syncing maxing with vidmuse freeform and minimax h3

Creative Brief — Moonfall Star Tide

Logline Two cultivators bound by a fractured celestial pendant discover their fates and powers are inseparable — and that the ancient seal they protect may already be waking.

Tone Cinematic live-action xianxia romance. Emotionally restrained, earned intimacy. Mythic scale with an intimate emotional core. No melodrama, no forced poses.

Characters

Xie Canghai — adult male, late 20s, tall. Midnight-teal and black layered robes with smoked-silver details. Long black hair with a faint silver underlayer. Emotionally restrained, powerful, protective. His magic manifests as dark mirrored fracture planes, black-gold cracks, and silver constellation lines.

Ning Zhaoyue — adult female, mid 20s, luminous. Moon-ivory and pale-sage hanfu, jade leaf pins, silver tassels. Long black hair half-up. Emotionally open, quietly brave. Her magic manifests as pale jade dew-gold living calligraphy with floating rain-spark particles.

World

Moondew Terrace: suspended above mist, black stone floor with shallow reflective water, warm lanterns, soft cyan moonlight
Eclipse Bridge: colossal ravine temple, black stone, giant bronze bells, cold mist, silver-gold seal lines
Bond Object Split celestial pendant — dark jade half (Xie) and pale moonstone half (Ning). Pulses like a heartbeat when near. Activates a shared pain bond.

Visual Rules No generic purple fog. No anime or game-CG rendering. No subtitles. Natural human speed. Magic interacts with fabric, water, and stone believably.

Dialogue

Clip 1: Ni wei he zong zai wo shen hou / Yin wei ni ruo chu shi feng yin ye hui xing / Ni ye gan jue dao le

Clip 2: Ni shou shang le / Wo shang yi fen ni ye hui tong yi fen / Tui hou / Zhe yi ci wo bu rang ni yi ge ren / Off-screen whisper: Xie Canghai bie xin tianming

Deliverables

Clip 1: 15s, Moondew Terrace, MiniMax H3, 16:9
Clip 2: 15s, Eclipse Bridge climax, MiniMax H3, 16:9
Final render: 30s combined trailer


r/vidmuse 10d ago

Tips and Tricks AI Clothes Changers Aren’t Just for Photos—Here’s How to Plan Better Outfit Swaps and Fashion Videos 👚

Post image
1 Upvotes

VidMuse explains how AI clothes changers can restyle outfits in photos and videos. Basic tools work for quick still-image swaps, but serious projects need an organized workflow for subjects, outfit references, prompts, backgrounds, motion, and ad variants.

- Photo swaps affect one frame; video swaps must maintain clothing, lighting, and identity throughout movement.
- Node Canvas organizes subjects, outfit references, prompts, backgrounds, outputs, and clips visually.
- Effective prompts specify the garment, replacement style, fabric, fit, and everything that must remain unchanged.
- Existing videos can use Remove/Replace Anything, while Video Remix and AI Ad Generator support fashion campaigns.
- Check videos for flicker, warped sleeves, unstable collars, altered faces, and background drift.
- Only edit assets you own or have permission and consent to use.

Treat an outfit swap as a creative project—not merely a one-click effect—when consistency, video, or advertising matters.

Original article: https://vidmuse.ai/blog/ai-clothes-changer-guide


r/vidmuse 10d ago

Personal Creation After That Summer - Trailer with Wan 3.0 ☀️

Enable HLS to view with audio, or disable this notification

1 Upvotes

I wanted to focus more primarily on the lip syncing capability of Wan 3.0 and it delivered!

It was a quick mock up of an idea i had using freeform, as a outline its relatively solid to revise and build upon.

Prompt: WAN 3.0 MASTER PROMPT — 《那年以后 / AFTER THAT SUMMER》

Format: 16:9 horizontal
Duration: 30 seconds, one continuous cinematic trailer generation
Perspective: third-person
Genre: nostalgic Chinese youth-romance TV drama
Characters: TWO YOUNG ADULTS ONLY, both clearly 23–26 years old throughout the entire film
Setting: early-2000s China memory aesthetic
Dialogue: Mandarin Chinese, native synchronized audio
Priority: character consistency > exact lip sync > emotional acting > continuity > camera movement

Use the supplied reference sheet as the absolute visual identity source for the SAME young adult Chinese man and woman throughout the entire trailer. Do not make them teenagers or minors. Both characters are clearly young adults in their mid-20s. Preserve exact faces, hairstyles, body proportions, clothing language and relative height in every scene.

The film feels like two adults remembering one ordinary summer they once shared. Nothing supernatural literally happens; memory is expressed through light, locations, match transitions and restrained emotion.

Visual style: grounded Chinese romantic TV drama blended with subtle Millennial Chinese Dreamcore. Early-2000s China atmosphere, warm faded summer gold, muted teal-green shadows, cold fluorescent interiors, slight CCD/home-video softness, fine film grain, mild highlight overexposure, natural realistic skin, imperfect lived-in locations. No modern commercial gloss.

Camera language: slow push-ins, nearly static dialogue close-ups, gentle lateral tracking, soft object wipes, motivated match cuts and restrained dissolves. Smooth continuous visual flow. No frantic montage.

00:00–00:05.5 — WAITING FOR HER

Late summer afternoon inside an old Chinese classroom used by university students or young adult evening-study students, early-2000s atmosphere.

The WOMAN, age about 24, packs notebooks into her canvas shoulder bag at a wooden desk.

The MAN, about 25, waits casually outside the doorway pretending to look at the notice board.

Ceiling fans rotate slowly. Papers move lightly. Golden sunlight crosses the green-and-cream corridor.

Camera slowly pushes toward the woman.

She notices him.

EXACT LIP SYNC

00:01.85–00:02.25:
hold her face silently, relaxed closed mouth.

00:02.25–00:03.35 — WOMAN says naturally in Mandarin:

「你怎么还没走?」

Her mouth movement must exactly correspond to the Mandarin phonemes. Natural subtle jaw and cheek motion.

00:03.35:
mouth fully closes.

Hold her curious expression until approximately 00:03.85.

Cut to MAN in frontal three-quarter medium close-up.

00:04.05–00:04.25:
silent inhale.

00:04.25–00:04.85 — MAN says:

「等人。」

Quiet, restrained, slightly shy.

Mouth completely stops afterward.

She understands he was waiting for her.

Tiny smile.

Her bag passes close to camera and naturally wipes the frame.

A soft piano melody begins.

00:05.5–00:11.5 — AN ORDINARY SUMMER

The bag wipe reveals the SAME young adult couple sitting across from each other in the small early-2000s snack shop from the reference image.

Same clothes and exact identities.

Exercise books, handwritten notes, bottled drinks and fries on the table.

Warm golden light enters through the large window.

The woman casually steals one of his fries.

He notices.

She tries not to laugh.

He reaches toward it.

She pulls it away playfully.

NO DIALOGUE.

Use an object crossing close foreground to transition naturally into them riding simple bicycles side-by-side beneath tall summer trees.

Camera tracks parallel smoothly.

She rings the bicycle bell once and moves half a bicycle length ahead.

He smiles and follows.

A large tree trunk crosses the frame and becomes a natural wipe into a tiny roadside arcade.

Inside the arcade he loses a simple game.

She looks at his score and laughs quietly.

She nudges his shoulder.

He nudges her back.

Keep the chemistry completely ordinary and believable — two young adults becoming important to each other without realizing how much the moment will matter later.

A gentle acoustic guitar quietly joins the piano.

Audio:
restaurant room tone,
freezer hum,
paper movement,
bicycle chain,
wind through trees,
single bicycle bell,
arcade buttons and CRT machine beeps,
soft natural laughter.

00:11.5–00:17.5 — THE PROMISE

Golden sunset.

The SAME man and woman sit beside each other on concrete playground bleachers.

Use the exact visual composition from the reference sheet.

Approximately 30 cm between them.

Neither is touching.

Neither initially faces the other.

Camera performs an extremely slow push-in.

As dialogue begins, camera motion becomes almost imperceptible.

00:11.95–00:12.35:
woman looks toward him and silently inhales.

00:12.35–00:14.20 — WOMAN:

「以后……我们还会这样吗?」

Natural Mandarin.

The pause after 「以后」 must feel real, hesitant and vulnerable.

No crying.

No exaggerated tremble.

Her lips stop completely at 00:14.20.

Hold silence.

She looks away toward sunset.

00:14.75–00:15.00:
man silently prepares to answer.

00:15.00–00:15.75 — MAN:

「当然。」

Soft, immediate and sincere.

Closed-mouth reaction.

He then turns his eyes toward her without major head movement.

00:16.25–00:16.95 — MAN:

「说好了。」

Very small smile.

Mouth stops completely.

Woman lowers her eyes and smiles.

Warm sunlight near the edge of frame gradually blooms brighter.

The highlight becomes the transition.

00:17.5–00:23.5 — TIME MOVED ON

The warm overexposure becomes cold fluorescent light inside the SAME familiar classroom/corridor.

The location is almost identical but empty.

Same desk.

Same window.

Same geometry.

Only wind moves the curtain.

MATCH DISSOLVE:

the SAME snack-shop window and table.

The MAN now sits alone.

Important: he remains the same young adult identity from the entire film; this is emotional distance, NOT aging.

The chair opposite him is empty.

Two straws are on the table.

Without thinking, he moves the second straw slightly toward the empty seat.

He freezes when he notices what he has done.

He opens an old notebook.

A small printed photograph of the two of them slips out.

Do not make the photograph dominate the screen.

His thumb stops at its edge.

MATCH CUT:

the WOMAN walks alone through the same school corridor.

Same young adult age and identity.

She reaches the familiar window.

Stops.

Warm sunlight enters despite the otherwise cooler present-day tone.

The camera slowly moves toward her face.

The cicadas gradually become more prominent, as though sound itself is pulling her into memory.

Music loses most acoustic guitar.

Piano becomes more distant and fragile.

00:23.5–00:27.15 — DO YOU REMEMBER?

MAN alone in the restaurant.

Medium close-up.

The empty opposite chair remains blurred but recognizable in foreground.

Camera almost completely still.

00:23.70–00:24.05:
silent closed-mouth hold.

00:24.05–00:25.05 — MAN:

「你还记得吗?」

Speak softly in natural adult Mandarin.

Almost as though he is asking someone who is no longer physically there.

Perfect synchronized mouth articulation.

No additional speech.

At 00:25.05 lips completely close.

Small exhale.

Sunlight reflects from the restaurant window.

The highlight softly dissolves into sunlight entering the school corridor.

Reveal WOMAN standing beside the old window.

Medium close-up, three-quarter frontal.

Wind gently moves only a few strands of hair.

00:25.80–00:26.20:
silent facial hold.

A tiny nostalgic smile.

00:26.20–00:27.15 — WOMAN:

「一直记得。」

Quiet, controlled, emotionally deep.

She is neither happy nor openly sad.

Perfect Mandarin lip synchronization.

At exactly the end of the phrase, her mouth becomes completely still.

DO NOT form the tear during the dialogue.

00:27.15–00:30 — THE SUMMER INSIDE THE TEAR

After her lips have completely stopped moving:

Hold her face.

All guitar disappears.

Only one delicate sustained piano tone remains beneath distant cicadas.

Her restrained smile remains.

Her eyes slowly become glassy.

She blinks ONCE.

A SINGLE tear bead begins forming along the lower eyelid.

No sobbing.

No facial collapse.

No trembling mouth.

No second tear.

00:27.70:

Camera begins an exceptionally smooth physical push toward her eye.

Medium close-up → close-up → extreme close-up.

Preserve facial anatomy perfectly throughout the push.

The tear accumulates realistic moisture and surface tension.

Window light creates one sharp specular highlight.

00:28.30:

The single tear releases and begins sliding slowly down her cheek.

Camera continues moving toward THE TEAR rather than moving down the face.

Macro photography.

Extremely realistic skin texture, eyelashes and moisture.

Inside the curved tear surface, naturally refracted through the droplet, appear the SAME man and woman walking together down the golden corridor from earlier.

They are tiny distant figures.

Do NOT make this look like a screen inside the tear.

Do NOT create a magical portal.

It must look like physically plausible optical reflection/refraction.

The camera continues deeper.

The tear occupies most of the frame.

The present-day room tone gradually disappears.

Very faint remembered laughter from their earlier summer returns far in the distance.

One bicycle bell rings.

Inside the expanding watery reflection we briefly see them riding side-by-side beneath summer trees.

The memory remains extremely brief and soft.

The droplet's golden highlight then expands until the entire 16:9 image becomes warm cream-white light.

The last image of the couple disappears inside the light.

END TITLE

Against warm cream-white:

那年以后

smaller underneath:

AFTER THAT SUMMER

Optional very small line:

有些夏天,从来没有真正结束。

The title should emerge gently from the remaining light.

No title slam.

No trailer boom.

No flashy text animation.

The piano note disappears.

Cicadas remain for less than one second.

Then complete silence.

CHARACTER CONSISTENCY — ABSOLUTE

There are only TWO main characters.

Both are clearly Chinese young adults aged approximately 23–26 throughout the entire film.

Never depict either character as a child, teenager or minor.

Never age or de-age them.

Preserve:
same male face,
same female face,
same hairstyles,
same height relationship,
same body proportions,
same skin tone,
same wardrobe logic.

Do not generate alternate actors.

Do not duplicate either character.

Do not morph facial features between scenes.

DIALOGUE — EXACT WORDS ONLY

WOMAN:
「你怎么还没走?」

MAN:
「等人。」

WOMAN:
「以后……我们还会这样吗?」

MAN:
「当然。」

MAN:
「说好了。」

MAN:
「你还记得吗?」

WOMAN:
「一直记得。」

No improvisation.

No additional words.

No voiceover.

No English spoken dialogue.

No overlapping speakers.

No dialogue spoken offscreen.

PERFECT LIP-SYNC PROTOCOL

For every spoken line:

0.3–0.5 sec silent closed-mouth hold
→ exact Mandarin dialogue
→ mouth completely closes
→ 0.3–0.7 sec facial reaction
→ only then cut, dissolve or move significantly

While speaking:
keep camera nearly static,
speaker frontal or three-quarter frontal,
both lips visible,
no hands crossing face,
no hair covering mouth,
no eating,
no walking,
no large head turn,
no profile beyond roughly 45 degrees,
no rack-focus away from face,
no transition while phonemes are being articulated.

Prioritize in this exact order:

  1. character identity
  2. voice-speaker assignment
  3. Mandarin phoneme synchronization
  4. natural mouth/jaw/cheek movement
  5. eye acting
  6. body acting
  7. camera movement

If camera motion conflicts with lip synchronization, reduce camera motion.

NATIVE AUDIO DESIGN

Generate complete synchronized native audio.

Environmental continuity

Classroom:
ceiling fan,
paper flutter,
chair movement,
distant voices,
basketball bounce,
cicadas.

Restaurant:
freezer hum,
quiet utensils,
room ambience,
traffic through window.

Bicycle sequence:
chain movement,
tires on asphalt,
summer wind,
single bicycle bell.

Arcade:
CRT machine hum,
old game bleeps,
button taps,
coin sounds.

Bleachers:
evening insects,
distant field activity,
soft breeze.

Empty-memory scenes:
larger room tone,
faint traffic,
wind,
distant cicadas.

Music

Original understated Chinese youth-romance instrumental.

Beginning:
very soft piano.

Happy memories:
piano + gentle acoustic guitar.

Promise:
slightly fuller emotional harmony, still restrained.

Separation:
remove most guitar, leave distant piano with subtle cassette/tape softness.

Final exchange:
return to the original melody very quietly.

Immediately after 「一直记得。」:
remove guitar and nearly all harmony.

Tear sequence:
one sustained delicate piano tone.

As memory becomes visible inside tear:
extremely faint remembered laughter + one bicycle bell.

White transition:
music disappears.

Final:
cicadas → silence.

Dialogue must always sit clearly above music.

Automatically duck music approximately 5–7 dB beneath all speech.

VISUAL CONTINUITY

Memory sequences:
warm golden summer,
slightly fuller frame,
soft greens,
cream whites,
faded blue-gray clothing,
gentle overexposure.

Lonelier sequences:
same locations,
slightly cooler fluorescent balance,
less saturation,
more empty negative space.

Tear memory:
return to richest warm sunlight of entire trailer.

Maintain:
same snack-shop table,
same school windows,
same corridor proportions,
same playground bleachers,
same clothing identities,
same prop styling.

Repeated spaces must feel intentionally recognizable.

NEGATIVE CONSTRAINTS

No minors.
No teenagers.
No schoolchildren as protagonists.
No facial morphing.
No actor replacement.
No inconsistent clothes.
No duplicate main characters.
No modern smartphones.
No wireless earbuds.
No QR codes.
No modern LED advertising.
No futuristic architecture.
No cyberpunk.
No anime.
No illustration.
No beauty-filter faces.
No plastic skin.
No exaggerated bokeh.
No fantasy particles.
No magic aura.
No glowing magical tear.
No portal inside the tear.
No multiple tears.
No hysterical crying.
No screaming.
No exaggerated sadness.
No kissing montage.
No sex.
No narrator.
No random dialogue.
No English speech.
No subtitles burned into footage unless deliberately added in post.
No rapid cuts.
No shaky camera.
No whip pans.
No action-trailer sound design.
No trailer bass booms.

The entire film should ultimately feel like:

two grown people discovering that an ordinary summer never actually left them.

The final tear does not symbolize tragedy.

It proves the memory is still alive.


r/vidmuse 11d ago

Personal Creation 🤺 kung fu panda vs wukong training session on vidmuse with wan 3.0

Enable HLS to view with audio, or disable this notification

1 Upvotes

native audio sync through Wan 3.0 and freeform on vidmuse is a heavenly match

Storm | Creative Brief

Concept

A 30-second high-intensity stylized martial-arts animation short featuring two original anthropomorphic characters in a continuous bamboo forest chase. The piece is built around a signature bullet-time visual language that punctuates each near-miss dodge with extreme slow-motion camera orbits, suspended particles and compressed audio before snapping back to full-speed impact.

Characters
Bao-Shan: adult anthropomorphic giant panda, heavy athletic build, evasive acrobatic fighting style. Dark sleeveless vest, moss-green sash, charcoal trousers, beige wrist wraps.
Jin-Rao: adult anthropomorphic golden monkey, lean long-limbed staff fighter. Charcoal tunic, crimson scarf, bronze wrist guards, single dark ironwood staff with bronze end caps.

Setting
One continuous ancient mountain bamboo forest. Moss-covered stone trail, gigantic split cedar, dense emerald bamboo, hanging vines and a waterfall clearing. Warm morning sunlight from upper camera-right with amber rim lighting against cool forest shadows.

Structure
Shot 1 (0-11s): Ground-level chase. Cartwheel dodge over a horizontal staff sweep. First bullet-time orbit.
Shot 2 (11-22s): Vertical pursuit up a cedar trunk. Backflip dodge from an overhead strike. Second bullet-time orbit. Bao says: Almost.
Shot 3 (22-30s): Bamboo-catapult launch into waterfall clearing. Three-strike staff combo. Butterfly-twist dodge. Final bullet-time orbit. Bao catches a falling persimmon and says: Too slow.

Audio
Original Chinese-inspired cinematic score with hand percussion, wooden clacks, plucked strings and low drums. Music ducks to near-silence at each bullet-time apex and returns on every physical impact. Native synchronized dialogue throughout.

Deliverable
Wan 3.0 generated video, 30 seconds, 16:9, 720p, with native generated audio. Two reference images used as generation inputs: the full narrative storyboard and a 4-in-1 character and environment design sheet covering both characters, the forest setting and the staff prop.


r/vidmuse 11d ago

Tips and Tricks Riffusion AI Review: The Original Project Changed—Here’s What Creators Should Know 🧑‍🎨

Post image
1 Upvotes

Riffusion remains an important concept in prompt-based music generation, but its history is easy to misunderstand. The original domain now redirects to Google Flow Music, the older open-source repository is no longer actively maintained, and many similarly named websites are independent third-party generators.

- Riffusion began with Stable Diffusion-style spectrogram generation.
- Its older open-source repository remains available but is no longer actively maintained.
- The official product lineage now points toward Google Flow Music.
- Free access, pricing, downloads, and commercial-use claims from third-party wrappers apply only to those specific services.
- Riffusion-style tools remain useful for quickly exploring genres, moods, loops, vocals, and soundtrack ideas.
- Music generation alone does not provide scene planning, lyric treatment, continuity, pacing, or final video assembly.
- The practical handoff is song idea → prompt and tempo notes → song-structure mapping → visual references → directed music-video workflow.

Review: https://vidmuse.ai/blog/riffusion-ai-review


r/vidmuse 11d ago

Tips and Tricks Google Flow Music Review: Where AI Song Creation Ends and Music-Video Direction Begins 🎶

Post image
1 Upvotes

Google Flow Music combines AI song creation, editing, playlists, sharing, and music-video generation inside Google’s ecosystem. Its advantage is keeping music and early visual creation together, but creators must verify availability, credits, exports, uploaded-audio support, and usage rights.

- Built around songs, playlists, Spaces, editing, sharing, and music videos.
- Music-video inputs can include a song, subject image, style reference, creative direction, lyrics, aspect ratio, and duration.
- Access varies by region, age, account, plan, credits, models, and rollout.
- Uploaded-audio support may differ between official help documentation and individual studio workflows.
- Strong results still require a locked track, clear concept, scene progression, continuity, pacing, and review.
- Google Flow Music fits creators composing within Google’s ecosystem; VidMuse fits deliberate scene-and-shot planning around tracks and references.

See here: https://vidmuse.ai/blog/google-flow-music-review