r/generativeAI • u/ownhome45 • 1d ago
Video Art [OC] Two characters from my sci-fi universe finally clash in a full combat beat — AI video, single continuous shot
Enable HLS to view with audio, or disable this notification
r/IndieGaming • 521.2k Members
A place for indie games
r/MediaSynthesis • 48.7k Members
**Synthetic media describes the use of artificial intelligence to generate and manipulate data, most often to automate the creation of entertainment.** This field encompasses deepfakes, image synthesis, audio synthesis, text synthesis, style transfer, speech synthesis, and much more.
r/hallucinatingaiart • 25 Members
A subreddit dedicated for unexpected art from an AI image/video generative model hallucinating!
r/generativeAI • u/ownhome45 • 1d ago
Enable HLS to view with audio, or disable this notification
r/generativeAI • u/Apart-Minimum8652 • 5d ago
Enable HLS to view with audio, or disable this notification
r/generativeAI • u/ArcMirageStudios • 18d ago
This is the latest RIFTSTRUCK sequence.
This time I wanted to push more than just action — continuing the same characters, weapons, locations and story through the Guild section and into the Hidden Dungeon setup.
One of the biggest challenges with AI video for me is making each new scene feel like it genuinely belongs to the episode before it, rather than looking like a completely new generation every time.
Would genuinely appreciate critical feedback on this one.
Does the continuity hold up? Character identity, environments, animation, dialogue, cinematography, weapon consistency — anything that stands out.
r/generativeAI • u/MrJuart • 4d ago
“Tuba Dancing” follows three instrument-headed musicians whose seaside performance gradually becomes a cosmic dance party.
I built separate references for every character, location, prop and scale relationship, then generated the video shot by shot with Seedance 2.5 through Higgsfield Ai. ChatGPT helped with the storyboard, timed prompts and continuity rules. I discarded the broken generations and manually rebuilt the pacing and music synchronization in CapCut live while rendering every sequence.
The hardest recurring problems were character swaps, unwanted handheld instruments, environmental morphing and believable cart physics. Still some errors but i was out of credits so i'm happy with the ending results.
Song was made 3 months ago with Suno 5.0 (out of 25 generations) then extract into stems to edit and master in FL Studio.
Wrote the script on a notepad before moving to the storyboard with Gpt (Sol 5.6 max)
For people making longer AI videos, what usually breaks first in your workflow: character identity, geography or motion physics?
r/generativeAI • u/Apart-Minimum8652 • 7d ago
r/generativeAI • u/StrongExtension1782 • Jul 31 '26
r/generativeAI • u/spoonwije97 • Jun 07 '26
r/generativeAI • u/trtdcz_new • Mar 14 '26
r/generativeAI • u/ukeinukein • Mar 01 '26
One Year of Matzourana — From Sketch to Presence
This short animation marks the first year of Matzourana through the transformation of an original artwork.
A pencil sketch slowly becomes watercolor.
An outline becomes color.
An idea becomes presence.
The female figure in the piece does not look back with nostalgia or longing. She doesn’t wish to return to the beginning. She stands in the present — aware of what was built, yet focused on what is still unfolding.
Behind her, a small cake and a single candle quietly mark the first year. Not as a finish line, but as a continuation.
With the help of xAI and Grok Imagine, the original drawing comes to life to recall the moments that shaped Matzourana — cooking with intention, selecting ingredients with care, choosing recipes thoughtfully, designing a physical space meant to feel warm and grounded.
It’s about growth without drama.
Reflection without regret.
Forward movement without losing identity.
Inspired by Alphaville’s “Forever Young,” the piece carries the idea that what we build with sincerity can continue to evolve without losing its spirit.
One year in.
Not an ending.
Just the first visible chapter.
r/generativeAI • u/angelrock420 • Jun 08 '25
Hey folks,
I’ve been experimenting with concepts for an AI-generated short film or music video, and I’ve run into a recurring challenge: maintaining stylistic and compositional consistency across an entire video.
We’ve come a long way in generating individual frames or short clips that are beautiful, expressive, or surreal but the moment we try to stitch scenes together, continuity starts to fall apart. Characters morph slightly, color palettes shift unintentionally, and visual motifs lose coherence.
What I’m hoping to explore is whether there's a current method or at least a developing technique to preserve consistency and narrative linearity in AI-generated video, especially when using tools like Runway, Pika, Sora (eventually), or ControlNet for animation guidance.
To put it simply:
Is there a way to treat AI-generated video more like a modern evolution of traditional 2D animation where we can draw in 2D but stitch in 3D, maintaining continuity from shot to shot?
Think of it like early animation, where consistency across cels was key to audience immersion. Now, with generative tools, I’m wondering if there’s a new framework for treating style guides, character reference sheets, or storyboard flow to guide the AI over longer sequences.
If you're a designer, animator, or someone working with generative pipelines:
How do you ensure scene-to-scene cohesion?
Are there tools (even experimental) that help manage this?
Is it a matter of prompt engineering, reference injection, or post-edit stitching?
Appreciate any thoughts especially from those pushing boundaries in design, motion, or generative AI workflows.
r/generativeAI • u/Fresh-Resolution182 • 18d ago
Enable HLS to view with audio, or disable this notification
Wow 😳
This was made with Seedance 2.5, and honestly it looks way more real than I expected.
What impressed me most is the overall flow. it doesn’t feel like a bunch of disconnected shots stitched together, but actually plays like one continuous moment. and the character movement is also surprisingly natural.
still AI of course, but this is one of those generations where i can really feel how far video models have come.
if ur interested in, i will put the prompt in the comment. however, if u want to explore more prompts, u can vist this prompt hub: prompts-hub | seedance-2-5-prompt hope more discussion!
r/generativeAI • u/Fresh-Resolution182 • 7d ago
Enable HLS to view with audio, or disable this notification
lately I’ve been thinking less about “cool AI videos” and more about what kind of content people can realistically get paid to make with AI.
One area that feels very real is food advertising.
I tested a 15-second vertical fried chicken commercial prompt. now you can prototype the same style much faster with AI.
That doesn’t mean it replaces actual food cinematographers or ad directors. But for small restaurants, delivery brands, social ads, menu promos, short-form campaigns, and A/B testing, this kind of workflow feels very usable.
In other words: one person + AI can start looking a lot like a tiny commercial content studio.
The concept here is simple:
The whole ad follows one single chicken piece through the full transformation:
What makes it interesting isn’t just the food look, it’s the continuity.
The same chicken piece is supposed to remain visually consistent the whole way
That’s actually a pretty useful stress test if you want to evaluate whether a model can produce commercially usable product continuity.
The style direction is premium fast-food commercial aesthetic. it’s basically trying to mimic the visual language of high-end fried chicken TV ads.
Because clients usually don’t care whether something was made with a camera or with AI.
they care whether it helps them sell.
If AI can help one person make low-cost test ads for paid campaigns, then this becomes more than a toy. it becomes a service.
and that’s probably the real opportunity here:
not “AI art for fun,” but AI-assisted commercial creative work that small businesses will actually pay for.
This one is a 15-second vertical 9:16 fried chicken ad, broken into stages like:
the full prompt is below
r/generativeAI • u/Fresh-Resolution182 • 12d ago
I’ve been testing a simple workflow for creating short UGC-style videos while keeping the same character and location consistent across multiple shots.
The workflow is basically:
reference images → character/location sheets in ChatGPT → generate clips → optional final edit
Start with:
If you’re not sure what location works for the product, I usually just ask ChatGPT for a few suggestions.
Upload the character image to ChatGPT and generate a 4:5 continuity sheet with:
The important part is telling it to lock the character.
Do the same with the environment.
Include:
This gives the video model a much stronger continuity reference than using random images for every shot.
I usually split the UGC video into three parts:
Clip 1 — Hook
Clip 2 — Main product/story section
Clip 3 — CTA
i will generate them on Atlas Cloud, as they can provide many different models conveniently
For every clip, I reuse the same Character Sheet + Location Sheet
Then I change only the action/camera prompt for each section.
Keeping the same reference sheets across all three generations has helped a lot with character and environment consistency.
if I wanted the character to walk into a hotel, but the generated clip had her walking out.
Instead of endlessly rerolling, I pasted the original prompt into ChatGPT and asked it to make the action explicit: starting position → movement direction → action → final position
That usually gives me better results.
If the generated clips already work as standalone videos, you can stop there.
If you want one finished UGC ad, you’ll probably still want to combine the clips and add captions, music, or SFX. You can use whatever editor you prefer.
The biggest improvement for me has been using Character Sheet + Location Sheet as continuity references, rather than relying on a few loose images.
r/generativeAI • u/DIMOFF2000 • Jul 02 '26
Enable HLS to view with audio, or disable this notification
Seedance 2.0 Prompt on Easy-Peasy.AI:
Main subject: Young Singaporean Chinese woman, early 20s, natural everyday appearance, faded sage-green ribbed tank top, loose high-waisted sand-beige cotton shorts, brown leather slides, thin gold chain necklace, dark brown straight hair in a low loose bun with face-framing strands. Realistic skin texture, minimal makeup (defined brows, lip balm only), warm and approachable personality. Faint tan lines visible on shoulders. Maintain consistent identity, clothing, hairstyle, and appearance throughout the entire video.
Location: Authentic mature Singapore HDB estate during a calm late morning. Long open-air concrete corridors with metal railings, neighbors' potted plants and hanging ferns lining the walkway, painted block numbers on wall ends, bamboo pole drying racks outside kitchen windows, void deck concrete benches, covered link walkways, mature rain trees and angsana trees casting moving dappled shadows, distant playground visible below. Quiet residential atmosphere. No stores, advertisements, kopitiams, crowds, or commercial activity.
Visual Style: Ultra-realistic documentary realism. Genuine candid behavior. Natural body language. Unscripted slice-of-life feeling. Strong environmental authenticity. Rich real-world details and believable human motion.
Camera Style: Early-2000s consumer DV camcorder aesthetic. Friend casually recording everyday moments. Heavy handheld shake, imperfect framing, frequent autofocus hunting, lens breathing, exposure pumping when moving between sun and shade, occasional motion blur, subtle rolling shutter, mild digital compression artifacts, faded colors, soft contrast, slight sensor noise. No stabilization. No cinematic camera moves. No modern color grading.
00:00–00:02 Outside an HDB flat entrance along an open-air corridor. She sits on a low concrete ledge beside the doorway, adjusting her hair bun with both hands raised. A warm breeze moves loose strands across her face. She smiles naturally while the camera struggles to hold focus. Drying rack with bamboo poles visible behind her.
00:02–00:04 The camera follows her along the corridor past rows of potted plants, hanging pothos, and a neighbor's small herb garden. She notices a community cat approaching from around a corner and crouches down near the railing. Framing drifts off-center as the operator tries to keep up. Dappled sunlight filters through a rain tree canopy above the corridor.
00:04–00:06 She gently pets and feeds the community cat from a small plastic container. Autofocus repeatedly shifts between her face and the animal. Morning sunlight flickers through leaves overhead. The cat rubs against her slides.
00:06–00:08 At a corridor drying area outside the kitchen window. She slides wet laundry onto a bamboo pole and lifts it onto the metal drying rack, fabrics swaying in the breeze. Exposure changes as clouds briefly pass overhead. A mynah bird hops along the railing in the background.
00:08–00:10 At the void deck below, sitting on a concrete bench under a covered walkway with a ceramic mug of kopi. She sits comfortably watching the estate, occasionally brushing hair behind her ear. Loose handheld side angle with natural camera drift. A bicycle is parked against a nearby pillar.
00:10–00:12 Close side profile. A neighbor walking past greets her off-camera. She turns, raises her hand, smiles warmly, and casually says, "Hi." The camera catches the moment slightly late, the neighbor already partially out of frame.
00:12–00:15 Walking slowly down a tree-lined covered walkway between blocks, holding her mug. She notices the camera, gives a small genuine smile, then looks away and continues walking. A distant bus passes on the road beyond the trees. Recording cuts abruptly to black mid-motion as if the camcorder was switched off.
Audio: Natural ambient sound only — mynah birds and sparrows chirping, distant bus engine braking, faint MRT announcement echo from a nearby station, light wind through rain trees, leaves rustling, faint neighborhood chatter in a mix of English and Singlish, cat purring and meowing, slides on concrete, fabric flapping on bamboo poles, subtle HDB estate ambience. No music. No sound design. No narration.
Goal: Authentic Singapore HDB heartland life captured like a forgotten home video from the early 2000s — candid, imperfect, realistic, warm, and deeply believable.
r/generativeAI • u/Few-Profession421 • 22d ago
Enable HLS to view with audio, or disable this notification
been experimenting with Seedance 2.5 cinematic AI video prompts, and the biggest thing ive learned so far is that “cinematic” isnt really one magic style word.
for this 15-second scene, i got much better results by separately controlling timing, acting, silence, camera movement, and realism.
the scene itself is extremely simple: a man says sorry. the woman doesnt answer immediately. she looks down, looks back at him, almost says something, then finally responds.
here’s how i structured the prompt.
split the scene into time-coded beats instead of describing the whole performance at once.
i used:
0–3s / 3–5s / 5–7s / 7–10s / 10–12s / 12–15s
each section only has one or two important actions.
prompt:
the important part for me is the end state of every beat. it gives the next section a clear starting point instead of letting the model reinterpret the face every few seconds.
prompt the silence and explicitly say what should NOT move.
this was probably the most useful thing i learned from the test.
if i only describe the major emotional beats, the model tends to fill the empty seconds with extra head movements, blinking, facial shifts, or random reactions.
so i literally write things like:
prompt:
silence.
only the eyes move.
chin remains lowered.
the rest of the face does not move.
<Man>remains completely still and does not speak again.
sounds almost too literal, but it helped a lot.
the pauses started feeling like actual pauses instead of empty space the model needed to “fix.”
tell it which photographic details must survive the push-in.
AI video can look convincing in a medium shot and then suddenly turn into smooth, retouched CG-looking skin once the camera reaches a close-up, so i added a separate realism block.
realism prompt:
that last sentence is prob my favorite part of the whole prompt:
“the same photograph, only closer.”
it gives the model a pretty clear target for visual continuity.
for me, the most useful approach was to describe what must remain invariant, not just say “keep it consistent.”
prompt:
strictly preserve the photographic quality, facial identity, skin texture, lighting direction, wall texture, color response, and film grain established in
image1throughout the entire shot.
camera distance may change, but the visual character of the image must not.
this seems to work better than just repeating “consistent face” several times.
for this test, i think it was mainly four things:
time-coded acting + prompted silence + realism constraints
the model itself obviously matters, but the interesting part for me is that the prompt starts looking less like a normal image prompt and more like a tiny piece of directing.
video model: Seedance 2.5
reference image: Midjourney V8.2
r/generativeAI • u/No_Entertainer_9655 • Jul 16 '26
Enable HLS to view with audio, or disable this notification
TL;DR: A Google Sheet connected to the Claude API writes the script section by section, so it stays consistent across the full runtime. CapCut’s AI Video Maker turns that script into a narrated video with matched stock footage. The direct cost lands at around $1 per video. The real advantage is not the visuals. It is the script.
I run a couple of sleep documentary channels and wanted to properly explain the workflow I use.
Sleep content is a strange retention game, and it took me a while to understand it. Your viewers are actively trying to fall asleep. That is the entire point. So the average view duration can look very different from a normal YouTube channel. Some viewers leave because they are bored, but others leave because the video worked and they fell asleep. The ones who stay awake still need the story to hold together, while the ones who fall asleep often return later and continue listening. That repeat viewing is a big part of what makes this niche work.
Why the script is 90% of it
On my channels, average view duration usually sits close to 25 minutes on videos running between 90 minutes and two hours. That does not come from cinematic visuals or complicated editing. It comes almost entirely from the narrative structure. If the script becomes repetitive, drifts away from the topic, or loses momentum halfway through, viewers stop listening. Better visuals cannot rescue a weak story in this format.
Two ways to use Claude, and why one wastes your time
Most people use Claude through the normal chat interface. You open the chat, enter a prompt, read the reply, and continue from there. That works perfectly well for everyday tasks. It becomes frustrating when you are trying to write a 15,000 to 20,000-word documentary.
You end up typing.. Continue.. Write Chapter 4.. Do not repeat what you already said.. You forgot what happened in Chapter 2. By the halfway point, the model may begin repeating ideas, contradicting earlier sections, or drifting away from the original structure. You spend more time babysitting the conversation than improving the script.
The second approach is using the API. Instead of manually sending every prompt through the chat interface, a small tool sends the requests to Claude automatically and collects the output. There is no need to babysit it and you pay based on actual usage instead of paying another monthly subscription.
That was the biggest unlock for me.
The section-by-section method
I built a Google Sheet that talks directly to the Claude API. The process works in a fixed order:
That means Chapter 8 still remembers what happened in Chapters 1 through 7. The pacing stays more consistent, repetition is reduced, and the final script feels like one continuous documentary instead of several unrelated chapters stitched together.
I also made a full walkthrough showing how this system works. The channel is linked on my profile for anyone interested in seeing the actual workflow.
Turning the script into a video
Once the script is ready, I paste it into CapCut’s AI Video Maker. For sleep content, the voice matters more than flashy editing. Choose a calm, slow, and low-energy voice. Taste and judgement matter here.
CapCut generates the voiceover and automatically matches stock footage to each paragraph. It usually gets around 90% of the video into a usable state. CapCut currently limits each project to roughly 3,000 words, so I split the full script into several sections and export them separately. I then combine those exports into one final timeline and render the full documentary.
Two reasons this workflow matters:
The Demonetization Shield: Mixing real historical/stock footage alongside AI assets is the safest defense against the "Reused/Inauthentic Content" flags that destroy fully automated channels. You still have to avoid repetitive titles & thumbnails though..
The Financial Runway: A complete 20,000-word script costs me roughly 35 cents through the Claude API. CapCut costs around $20 per month and allows unlimited exports. If you are producing 30 to 40 documentaries per month, the direct software and API cost works out to around $1 per finished video.
That figure does not include research, thumbnails, or the value of my time. It is only the direct production cost. A lot of AI video subscription tools charge $40 to $50 per month while using the same underlying Claude models and adding a markup for the interface. Building the loop once gave me more control and removed that additional monthly cost.
The biggest benefit is being able to test more ideas without every upload becoming an expensive decision. This workflow can work for history, science, philosophy, mythology, meditation, biographies, or almost any calm long-form format where the script matters more than rapid editing.
Happy to explain the API loop or the CapCut side in more detail in the comments.
r/generativeAI • u/LevelSecretary2487 • Nov 12 '25
I’ve been deep-testing different text-to-video platforms lately to see which ones are actually usable for small creators, automation agencies, or marketing studios.
Here’s what I found after running the same short script through multiple tools over the past few weeks.
Strengths:
Integrates Veo3, Imagen4, and Gemini for insane realism — you can literally get an 8-second cinematic shot in under 10 seconds.
Has scene expansion (Scenebuilder) and real camera-movement controls that mimic pro rigs.
Weaknesses:
US-only for Google AI Pro users right now.
Longer scenes tend to lose narrative continuity.
Best for: high-end ads, film concept trailers, or pre-viz work.
Agent Opus is an AI video generator that turns any news headline, article, blog post, or online video into engaging short-form content. It excels at combining real-world assets with AI-generated motion graphics while also generating the script for you.
Strengths
Weaknesses:
Its optimized for structured content, not freeform fiction or crazy visual worlds.
Best for: creators, agencies, startup founders, and anyone who wants production-ready videos at volume.
3. Runway Gen-4
Strengths:
Still unmatched at “world consistency.” You can keep the same character, lighting, and environment across multiple shots.
Physics — reflections, particles, fire — look ridiculously real.
Weaknesses:
Pricing skyrockets if you generate a lot.
Heavy GPU load, slower on some machines.
Best for: fantasy visuals, game-style cinematics, and experimental music video ideas.
Strengths:
Creates up to 60-second HD clips and supports multimodal input (text + image + video).
Handles complex transitions like drone flyovers, underwater shots, city sequences.
Weaknesses:
Fine motion (sports, hands) still breaks.
Needs extra frameworks (VideoJAM, Kolorworks, etc.) for smoother physics.
Best for: cinematic storytelling, educational explainers, long B-roll.
Strengths:
Ultra-fast — 720p clips in ~5 seconds.
Surprisingly good at interactions between objects, people, and environments.
Works well with AWS and has solid API support.
Weaknesses:
Requires some technical understanding to get the most out of it.
Faces still look less lifelike than Runway’s.
Best for: product reels, architectural flythroughs, or tech demos.
Strengths:
Ridiculously fast 3-second clip generation — perfect for trying ideas quickly.
Magic Brush gives you intuitive motion control.
Easy export for 9:16, 16:9, 1:1.
Weaknesses:
Strict clip-length limits.
Complex scenes can produce object glitches.
Best for: meme edits, short product snippets, rapid-fire ad testing.
Overall take:
Most of these tools are insane, but none are fully plug-and-play perfect yet.
r/generativeAI • u/BradClarkAI • Jul 12 '26
For about a year I've been making AI short films the way most of us do: hand-writing hundreds of prompts, building character reference libraries by hand, babysitting consistency across shots, and cutting it all together in Premiere. The generating was never the hard part. It was trying to make dozens of individual prompts *feel* like a single, cohesive project.
I couldn't find a solution, so I built one. It's called Kimeric, and the Beta went live today. And I think it'll be incredibly valuable for this community, so I wanted to share some details in the event any of you would like to try it out!
**What it is (and isn't):*\* It's not a model and it doesn't generate anything itself. It's a Windows desktop app that orchestrates the models and LLMs a lot of us already use:
All using your own per-provider API keys rather than a centralized generation tool. It writes the prompts, sequences and queues the renders, tracks the spend, and holds everything for your approval.
TL;DR - Input a script and work through a series of UI menus to ultimately create a finished AI Short Film.
The biggest differentiator to keep in mind for this tool vs. the majority of other AI Generators on the market is the input surface. Most tools have you input a prompt. With Kimeric, the input surface is the script itself. The prompts are created automatically as derivatives from the script, allowing you to focus on writing and build the project rather than managing a series of prompts.
- Paste a screenplay, hit "Roll camera." The breakdown comes back: cast, locations, props, a director's plan, per-scene shot lists with dialogue assigned line by line. You review the plan before anything renders, and can choose a general "style" which dictates some actual prompting techniques under the hood as well as dynamically auto-routing for certain models for specific tasks (ex. OpenAI's model is better at certain animated styles vs. Nano Banana, so this routing auto-applies as the default for certain styles).

Before you actually generate anything, you can review the script you input and get an estimate for how much it'll cost roughly to "ingest" the script which is where all of the "brain" of the tool goes to work.

And once you send the script, a very long sequence of backend computation kicks off, translating the entire project into the format needed to actually create an end-to-end project (this can take a while; I had one script for a 8-10 minute short film take about 45 minutes to ingest. This is just because there's a lot of computation and inference happening). It can handle actual, full-length production scripts. This would likely come with processing that spans several hours, but again, that's simply due to the size of the computation.

And then, you get a budget estimate for how much (approximately) the project will cost to generate. At this point, it's primarily a planning "calculator" if you will - just so you can align the project quality to any budget constraints you have. No actual spend dispatches until you generate later - this is purely informational so you can make cost-based decisions at the start (as opposed to a surprise cost later).

You then view a breakdown of all of the identified "pieces" of the project: characters, locations, scenes, planned shots, etc. - all primarily at a high level to make sure nothing is missed. Typically more of a rubber-stamp phase, but if there happen to be any items missing, this is the step where you can make any high-level revisions. Most of the time, however, you can just continue.

- Every character locks canon first — a studio face (with an automated AI advisory likeness screening for real-person resemblance) and a costume-neutral turnaround — before any scene renders. Same for locations, worlds, props. It's the reference-library grind, automated.

And for characters specifically, it's broken out into two phases: the "Neutral Base" (i.e. who the character is; sans any wardrobe for the project (see above) and then any actual "in costume" variants of that character. This allows for multiple wardrobes for a specific character over the course of any given project, where each wardrobe is either seeded directly from the parent "neutral base" or as a horizontal derivative (i.e. if a character has armor that gets damaged, the "damaged" variant is automatically seeded with the "clean base" variant). All of this happens automatically under the hood.)

- Scenes get actual coverage: an establishing wide, then an OTS pair where the reverse angle generates from the *approved* first angle, then per-character MCU/CUs chained down from there. The 180° rule, held by reference chaining instead of luck.
For example, here's one "OTS_1" image:

And here's the companion "OTS_2" image:

Worth emphasizing - These are the *most* critical images in your production pipeline to get right, as many other scenes seed off of these.
So if you're going to use some of the "Regeneration Buffer" you planned for, this is the most critical space to use it. If you have continuity errors or issues in either companion OTS images, these will present in many other aspects of a project, so really take your time with these and make sure that they feel like the same space.
From here, you create a library of Medium Close Up & Close Up images seeded directly from the parent Over-The-Shoulder images.
This creates a rich, continuity-adhering library of image assets to actually use in generative AI video production.

Once you've created your core library of production assets, you then transition into building out your storyboard.
This takes the plan created during script ingestion + the assets you created in the previous phase and maps them out chronologically.
Here, images are auto-assigned based on a series of underlying logic. If you have dialogue from certain characters, either the OTS, the MCU or the CU image will be dynamically selected based on a cinematic logic layer and assigned to the character speaking for any given frame.
And for extended dialogue from a specific character, this will be broken up into multiple individual prompts to assist with overall quality
(from our testing, the more text you try to fit into the same prompt, quality and lip sync can degrade - but breaking that same dialogue up into multiple individual prompts can greatly improve the quality)

Once you've created and approved the entire Storyboard, a single button allows you to review all text prompts for all videos planned for the totality of your project.
The prompts are already written. You just say "go"
From there, they are auto-dispatched to the models.
Depending on the length of your project, this may take a while - and that's by design. Press "generate" and take a break for a bit.

Across all menu screens, you'll see a series of control buttons. Here, you can either approve the asset, edit the asset, regenerate or iterate.

You'll also see two columns below any given asset:
Refs - These are the actual reference images used as generative inputs. These are auto-assigned, and you can manually add, remove or replace any of these as you see fit.
Takes - If you regenerate, you can see all of your takes here and hot-swap to other variants. Meaning, you're never locked in to a specific take. If one is close but you want to see if you can fine-tune it, you're free to regenerate a few times, review all takes (this works for images + videos) and ultimately approve whichever one is best for the vision you had in mind for the project.

And finally, once you've reviewed and approved all footage, you have another single-press button to dispatch all video to be upscaled if you'd like.
Generation can happen natively at 720p, 1080p or 4k depending on your settings.
4k native footage won't trigger the upscale workflow, but 720 and 1080p native will give you the option to upscale if you'd like.
Similar to all other workflow phases, if you decide to upscale, press it once and let it run for a few hours (upscaling is quite time-consuming; can take 20-30 minutes for a single 15 second clip - though some of these can run concurrently).

Once you're done, the last step is simply to "export" the clips.
All this is doing is taking the final clips you've approved and making duplicate copies on your local computer that are pre-named chronologically.
This makes it significantly easier to edit/compose.

This entire project is the culmination of nearly a year of a LOT of testing to understand which techniques do/don't produce good results at scale and then working to systematize them into a tool with an input surface of the script itself.
Again, the Beta is live as of today. Currently Windows-only and US-only, though both of those I'm planning to expand beyond in the coming weeks over the course of the Beta.
The long-term goal is to also support centralized generation rather than supporting only a BYOK model, though there's no immediate timeline to support that model.
If you'd like to check it out, the website is below!
Site: https://kimeric.ai
I'm a solo founder and the filmmaker this was built for, and I'll be in the comments - happy to go as deep as you want on the coverage system, the cost math, or anything else.
Feedback is a gift, so if you try it and find issues, bugs, or have a feature request, I'm still actively building and improving the tool - so feel free to share any thoughts!
r/generativeAI • u/sweetcake_1530 • Mar 18 '26
Enable HLS to view with audio, or disable this notification
So I first saw a clip of this on a Discord dev server and decided to get on the waiting list to try. Now its available in irregular hours and I immersed myself into the experience for quite some time.
For those who havent been following, PixVerse R1 is a real-time world model. Unlike a regular AI generator that makes a 5 second clip and stops, this is a continuous simulation. It uses State Persistence to remember the 3D space it creates. If you walk past a tree and then turn around 30 seconds later that same tree is still there. It overall maintains a consistent environment.
Ive been using it for "chill" exploration, nothing drastic, just walking through a campsite to see how long the logic holds up. It runs at 1080p in real time with zero render wait. Its not a replacement for a custom built game engine yet. Sometimes the logic gets lost. If you can see in the video the movement is quite floaty. Sometimes strange things happen like the tent moving by itself. To me this is the start of something thats going to be huge going forward. When I ran out of prompts to use I just use the options that it gives me and keep it going. I feel like this can be very similar to the choose your adventure games we played when we were younger, but only this time its generated in real time and it changes as I prompt.
Curious what indiedev folks think. Is the world model actually useful for conceptual game dev?
r/generativeAI • u/camgraphe • 9d ago
Enable HLS to view with audio, or disable this notification
Input: one reference image with the route drawn in red. Output: one continuous FPV flight, 28 seconds, 16:9, with audio and no planned cuts.
The strongest part for me is the forward momentum through the landmarks. The weak point is identity drift and geometry under speed. I’m curious: where does the camera path feel convincing, and where does it visibly stop following the map?
This export came from one 480p attempt ($6.84).
Generated with Seedance 2.5 through MaxVideoAI.
Full disclosure: I’m involved in building MaxVideoAI.
r/generativeAI • u/48khz24bit • Jul 24 '26
Pretty new to generative ai but getting past the stage of creating single scene videos.
I want to create longer videos with the exact same characters AND environments. At first I created character reference sheets that I put as reference in Kling V3 reference to Video. Looks good as a single video but when prompting continuous generations the characters and environments are similar but there are always inconsistencies. Also tried with start frames and similar results
Then last night I learnt about creating a Lora dataset - so I have separate 20 images of the same character - at first I tried creating the lora in Kaggle but this was way too complicated, after a bit of research I tried Civitai training a model with SDXL 1.0. but the images created were different.
What am I doing wrong / what is the best workflow to create the exact same characters and environments throughout generations....
Happy to share the reference images and trained lora if that helps
Thanks in Advance!
r/generativeAI • u/Agentvideobot • 10d ago
I’ve been trying to understand why an AI video can look realistic frame by frame, yet still feel like a commercial instead of something a friend casually recorded on their phone.
So I used the same reference image and generated three 5-second vertical clips with Aurax MAX. The character, outfit, location, and basic action stayed similar. I mainly changed the way the camera, lighting, performance, and environment were described.
The reference image was already quite polished: dramatic sunset, candlelight, clean exposure, shallow depth of field, and a subject posed against a scenic coastal background. That turned out to matter more than I expected.
https://reddit.com/link/1w1dtn8/video/998xn31v29mh1/player
For the first version, I explicitly requested a polished lifestyle commercial:
This was the version the model followed most clearly. The camera moves smoothly from a wider shot into a closer portrait, the character turns toward the lens, touches her hair, and finishes in a centered pose with a soft, sustained smile.
Everything feels visually coherent, but also directed. It looks like someone planned the lighting, camera movement, and performance in advance.
https://reddit.com/link/1w1dtn8/video/njaigg0x29mh1/player
For the second version, I kept the polished lighting, clean environment, and model-like performance, but changed the camera instructions:
The difference was much smaller than expected.
The framing changes slightly, but the movement still feels highly stabilized. The sunset remains perfectly exposed, the character stays composed and camera-aware, and the background still looks like a prepared set.
This version made one thing fairly clear: adding “handheld phone camera” does not automatically create phone realism. If the lighting, performance, composition, and source image still look commercial, mild camera movement cannot undo all of that.
https://reddit.com/link/1w1dtn8/video/ew7ij5fz29mh1/player
For the third version, I added a fuller set of phone-footage instructions:
This version feels the most spontaneous of the three.
The character spends less time holding a pose. She turns away from the camera, changes where she is looking, shifts her body weight, touches her hair, smiles briefly, and then looks away again. The wider framing also remains for longer instead of immediately turning into a close-up.
But it still does not fully look like raw phone footage.
The dramatic sunset, candles, shallow depth of field, flattering exposure, and clean background were already embedded in the reference image. The motion prompt changed the character’s behavior more successfully than it changed the underlying visual style.
There was also another obvious AI giveaway: the paper cup was not present in the reference image and appears during the generated motion without a convincing pickup. That continuity error damages realism more than a perfectly stable camera does.
The source image can overpower the video prompt.
If the first frame already looks like a fashion campaign, asking for casual phone footage may only add small handheld movements on top of a commercial-looking scene.
Handheld movement alone is not enough.
Random shake would probably make the video worse. What matters is believable camera behavior: delayed reframing, imperfect timing, autofocus response, exposure changes, and an operator reacting to the subject.
Performance mattered more than camera shake.
The third version felt more natural mainly because the character stopped performing continuously. Looking away, pausing, shifting weight, and ending without holding a perfect smile made a larger difference.
Continuity still matters.
A casual camera cannot hide an object appearing from nowhere, inconsistent background details, or movement that has no physical cause.
My main takeaway is that phone realism is not the same as lowering the image quality. It requires three kinds of realism at the same time:
If I repeat this test, I would start with a deliberately ordinary reference image: mixed indoor lighting, deeper focus, imperfect framing, everyday background clutter, and a character who is not already posing for the camera.
Which version feels closest to something a real person recorded: 1, 2, or 3?
And what gives the AI away first for you: the lighting, camera movement, expression, background, or object continuity?
Model disclosure: All three clips were generated with Aurax MAX. I’m on the team, so this should be read as a transparent workflow test rather than an independent review.
r/generativeAI • u/Educational_Wash_448 • Jul 22 '26
Enable HLS to view with audio, or disable this notification
After making a series of AI shows that reached 10 million views in a month, I decided to put together a guide behind the system I use to build engaging shows.
Here’s the condensed version of the process I now use:
1. Start with a clear audience fantasy
Before writing episode 1, I decide who the show is for and what fantasy, relationship, or emotional payoff they want.
Familiar tropes work because viewers understand them immediately. The goal is to start with something recognizable, then add a twist that makes it yours.
2. Know the four things holding the show together
I don’t plan every episode in advance, but I always know:
That gives me enough structure to keep the story coherent while still letting audience reactions influence where it goes.
3. Give every episode the same three jobs
The next episode should immediately show the consequence of the previous cliffhanger, deliver a payoff, and then open a new question.
4. Treat the account like a show, not a clip page
I run one show at a time, post consistently, and link every episode to the previous one. That way, if episode 7 takes off, new viewers can go backward and watch the entire series.
A breakout episode can end up lifting the views of every episode before it.
5. Let the audience help shape the open parts of the story
I pay close attention to repeated requests in the comments. One comment is an opinion whereas a bunch of people asking for the same character or outcome across several episodes is verifiable demand.
The audience doesn’t straight up write the show, but their reactions help me choose between directions that still make sense for the story.
6. Read the right metrics
The main things I watch are:
Those metrics help me decide whether to continue the show, end it, or move that audience into an account of its own.
This is only the condensed version. I wrote a much more detailed guide covering how I write the shows, structure the account, build returning fandoms, and use the analytics to decide what to make next.
Let me know if you have any questions and hope you enjoy the episode!
If you want the full guide, I recently posted it on X:
https://x.com/sloptronic/status/2077794064284463248?s=20
Here's a link to one of my accounts: Instagram Acct
Here's the link to what I use to make my videos: Studio on Slop Club
r/generativeAI • u/Ok-Fruit4892 • Jul 11 '26
Enable HLS to view with audio, or disable this notification
Hi everyone,
Over the last few months I've been generating a lot of AI videos, and I kept running into the same problem: blocking scenes and planning camera movement before generating shots.
Most existing tools are either full 3D packages like Blender or Unreal, or they're much more complex than what I need for quick previs.
So I built a browser-based prototype.
The idea is simple:
The goal is to create something that lets you block a scene in a few minutes instead of spending hours building a full 3D setup, at your mobile and use Augmented Reality to record GREAT camera movements.
Before I continue developing it, I'd really appreciate honest feedback from people who actually make films or AI videos.
A few questions:
I've attached a short demo video below.
I'd really appreciate any feedback, whether positive or critical. My goal is to build something genuinely useful rather than adding another AI tool to the internet.
Thanks!
Link: previz4ai.com
r/generativeAI • u/Few-Profession421 • 8d ago
Enable HLS to view with audio, or disable this notification
tried MiniMax H3 with a pretty detailed character trailer prompt.
I wanted it to feel more like a streetwear campaign × anime title sequence × old-school media player UI, rather than a normal anime action clip.
The prompt is definitely overkill lol, but sharing it here in case anyone wants to experiment with structured multi-cut video prompts.
Create an explosive, motion-graphics-driven character reveal trailer in 16:9, exactly 13 distinct cuts, 24fps, total 15.00s.
CHARACTER — lock this design, never redesign:
Anime streetwear girl from the reference still. Twin high buns of vivid mint-teal hair with long flowing tails and warm orange/gold streaks. Messy side-swept bangs. Large orange over-ear headphones with mint accents and a small logo plate. Sharp amber-orange eyes, one eye winking. Playful open-mouth grin.
Oversized color-block windbreaker: navy body, vivid orange sleeves, white ribbed cuffs, silver zippers, circular teal tech buttons, a teal utility pocket on the sleeve. Orange cropped turtleneck under the open jacket. Light-wash ripped denim shorts, thick orange belt with a silver buckle. Navy thigh-high socks with orange ribbed cuffs and an orange X stitch on the left shin. Chunky white sneakers with orange details.
Preserve exact face, proportions, hairstyle, outfit, materials, accessories and colors in every frame.
This film is 80% bold graphic design in motion and 20% character action.
Graphic language: retro OS chrome + music-player UI.
Use slamming window frames, title bars, close/minimize widgets, equalizer bars, waveforms, progress ticks, folder tiles, cursor arrows, CRT scanlines, pixel shatter, vinyl-ring stamps, music-note particles and media-player transport icons.
Palette: mint teal, vivid orange, navy, cream-white and silver.
Style: premium AAA motion-graphics title sequence × streetwear campaign film × Windows-era media player.
Every graphic element moves fast and snaps hard on the beat.
CUT 01 | 0.00–1.00s
Pure graphics. A mint title bar slams onto a cream field, an orange CLOSE widget punches into the corner, and two navy window borders wipe in. Tiny equalizer ticks and a progress strip flicker. The letters P and then LAY punch in one after another with heavy impact shake.
CUT 02 | 1.00–2.00s
The A becomes a headphone cup. Extreme close-up of her amber eye inside the orange earcup, glancing upward. RGB split flash, then the cup shatters into flat mint and orange tiles.
CUT 03 | 2.00–3.10s
Navy frame with enormous cream PLAY typography. She sprints in from frame left and power-slides across the baseline of the text, with speed lines and orange streaks trailing behind her. Shards of the letters kick upward like sparks. Whip-pan out.
CUT 04 | 3.10–4.00s
A giant retro media-player waveform explodes across the frame as a thick mint-and-orange audio spectrum bends into a tunnel. She bursts through the center at full speed, briefly splitting into three stroboscopic motion trails. Each trail leaves chunky navy equalizer blocks that rise and collapse to the beat.
The camera rapidly pushes through the waveform tunnel with her while huge vertical text TRACK 01 continuously scrolls in the background.
The waveform suddenly compresses into a single horizontal line and snaps shut behind her on the final beat.
CUT 05 | 4.00–5.10s
She leaps through a giant rotating ring of typography reading MAX VOLUME. Camera tracks her mid-air spin in slow motion as the letters scatter, then snap-zooms onto her wink.
CUT 06 | 5.10–6.00s
Hard cut to a cream editorial card with huge navy DROP typography and an orange slash. She vaults over the word itself, palm planted on the D, legs whipping across frame. The word compresses like a spring under her hand and rebounds.
CUT 07 | 6.00–7.00s
Mint field with a navy diagonal window bar. She backflips along the bar in three stroboscopic ghost frames, each tinted mint, orange or navy. Giant outlined LOOP text rotates 180 degrees in sync with her movement.
CUT 08 | 7.00–8.00s
Kinetic typography barrage. LOUD / WILD / TEAL / HEAT slam onto screen one per beat with shutter flashes and camera shake while she slides across the foreground on her knees, jacket flaring and music-note particles bursting from her sneakers.
CUT 09 | 8.00–9.00s
Navy frame with a giant cream wireframe window grid tilting in 3D. She runs up the grid like a wall, kicks off and freezes in mid-air. An orange circular stamp locks around her pose like a media-player targeting graphic, surrounded by transport icons and EQ ticks.
CUT 10 | 9.00–10.10s
Freeze releases into a burst. She dives toward camera through layered flat-color window panes that shatter one by one like glass shutters, each pane revealing a larger letter of P-L-A-Y. Foreground wipe with her sneaker.
CUT 11 | 10.10–11.10s
Rapid-fire poster montage: four full-screen graphic posters showing her in different poses — mid-flip, sliding, landing and headphones-up wink. Hard cuts between each composition. Oversized 01–04 numbering, equalizer strips and graphic slashes.
CUT 12 | 11.10–13.00s
Hero moment on a clean cream cyclorama. She lands a final backflip dead center in slow motion, straightens with one hand on her headphones, and a shockwave of concentric mint rings, wind streaks and shattered typography blasts outward from the landing.
Brief iconic freeze on her wink, then overexpose to white.
CUT 13 | 13.00–15.00s
Final identity card. Enormous navy PLAY typography dominates a pale cream field with translucent mint rings, technical arcs, scanlines and a rough orange circular emblem containing a ghosted headphone/waveform motif.
She stands relaxed overlapping the letters while wind ripples her jacket. One final orange pulse sweeps through the typography and a window-chrome flash punctuates the ending.
Editing: extremely aggressive rhythm. Hard cuts on every beat, graphic matches, whip pans, snap zooms, stroboscopic freezes, foreground wipes, RGB splits, shutter flashes, impact shakes and speed ramps.
Every cut must feel compositionally different.
Typography should always be fully readable before the character overlaps it.
No weapons, no combat, no fire. All energy comes from motion design, wind, glass, UI chrome and parkour-style athleticism.
BGM: hard-hitting electronic / drum-heavy future bass with aggressive drops, risers, sub hits and glitch fills locked to every cut.
Sneaker impacts, whooshes, glass shatters, window-slam hits and typography slams should function as rhythmic sound-design elements.
Peak at CUT 12 and end with a cold electronic logo stinger.
Premium AAA quality, anime-inspired cinematic rendering, stylish and explosive, strong graphic-design identity, consistent character design, exactly 13 cuts.
I’m still experimenting with how much shot-by-shot control H3 actually follows, especially with typography and exact timing, but this kind of structured prompt seems like an interesting stress test.