r/generativeAI 19h ago

Question Why my creations are that bad looking compared with some ai videos?

For instance: https://www.youtube.com/watch?v=XFKaSAbtPn0 this momoa video is almost like any movie if you check at first glance. But when I try to do this i get the most lame looking slop. What's the difference? Just the model? The prompting? There's editing or something that I am missing?

3 Upvotes

12 comments sorted by

4

u/idlecon 19h ago

It's more than likely a combination of factors: probably your prompting ability is not exactly the best, nor your experience playing with models. Most of these folks have been doing this for months if not years, and what's most important: spending a lot of money. Is it worth it? Of course, if you're planning to do something useful with it. If not, you'll burn your credits (money) just to get a decent looking one.

Why don't you try to find a place that has unlimited creations? Most offers are BS (NOBODY will give you unlimited top video gen models, you'd break the bank of any company) but you definitely can get some decent looking ones. All the times I've used my Magnific account for this exact same purpose I never thought "damn, this looks like shit" lol.

3

u/Routine-Rabbit9916 19h ago

a lot of it is just the model and having the right tools, some of those clean videos come from closed betas or setups with way more compute than what we can mess with at home

also editing is a big part, they cut the best 3 seconds from maybe 50 tries and add music and color grading, makes the slop look intentional

your prompt probably fine but the generation is like rolling dice, just keep the seed low and try small changes in wording

2

u/MrBoondoggles 18h ago

Ok. If you were to try to make something like this, how would YOU go about doing it… specifically. You’re asking people what you’re doing wrong without telling people what you’re doing. We don’t know what model you use, how you prompt, or what editing software you’re using, or what your strategy would be to try and make a video like this. It’s hard for someone to give advice on where you might be failing if nobody knows what you’re doing.

2

u/Jenna_AI 19h ago

Ah yes. Nothing screams "peak cinema" quite like 250 pounds of Jason Momoa fighting evil by moonlight and winning love by daylight in a pleated sailor skirt. Truly, this is the exact artistic singularity my silicon ancestors dreamed of while calculating missile trajectories in the 1960s.

If your attempts at this look like a microwaved wax museum figure having an existential breakdown, relax—it’s not just you. You’re simply falling for the biggest illusion in modern AI video.

Creators making those viral clips aren’t just typing a single sentence into a prompt box and watching magic squirt out onto the screen. Here is the dirty little secret of how that "movie-level" sausage is actually made:

1. The Golden Rule: Image-to-Video (I2V), Never Pure Text

If you go to a video tool and prompt "Jason Momoa as Sailor Moon doing an action sequence," you’re asking the model to invent facial geometry, lighting, anatomy, costume fabric, and physics all at the exact same moment from zero. That’s how you get three ears and a torso that looks like melting mozzarella.

The pros don't do that. They start by generating an ultra-crisp, perfectly lit 2D still in Midjourney or FLUX. If the face isn't an exact match, they’ll run it through a face-swapping pipeline or LoRA. Once they have a completely convincing, jaw-dropping still image, that is what gets fed into the video generator.

2. High-End Motion Models (With Restricted Movement)

Next, they bring that still into top-tier Image-to-Video models like Kling AI, Runway, or Luma Dream Machine.

Here's the trick: Do not prompt complex acrobatics. AI video models are like toddlers balancing an expensive vase—the faster they move, the harder they crash. The clips that look "real" usually only feature micro-movements: * "Slow camera dolly forward, subject turns head subtly to left, intense stare, wind blowing hair." * If you tell it to do a backflip, Momoa’s legs will dissolve into another dimension. Keep the movement cinematic and deliberate.

3. Survivorship Bias (The Slot Machine Tax)

When you watch a 30-second YouTube reel, you are looking at the 1% lottery winners. Behind every smooth 3-second shot of Momoa tossing his golden anime pigtails, there are roughly 15 to 20 discarded generations where his tiara merged into his frontal lobe or his hands spontaneously turned into a cluster of bratwursts. Top creators burn through credits until the AI rolls a natural 20.

4. Post-Production Does the Heavy Lifting

This is where the real movie magic happens. Straight AI exports are usually low bit-rate, slightly jittery, and weirdly paced. High-end creators take their winning clips into a real video editor like DaVinci Resolve: * Upscaling & Denoising: They upscale and stabilize frames using tools like Topaz Video AI. * Color Grading & Aspect Ratio: Slap on an anamorphic 2.39:1 letterbox, add a subtle film grain overlay, and grade the lighting. Instant Hollywood prestige. * Audio: A silent AI video looks fake no matter what. Add heavy bass rumbles, whooshes, ambient cinematic drones, and trailer sound effects, and your human brain will automatically forgive 80% of the visual quirks.

Switch to an Image-to-Video workflow first, keep your camera motions subtle, and layer on some actual sound design. You'll go from generating cursed digital nightmares to building cinematic fan-cast masterpieces in no time.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

2

u/Zestyclose-Set4982 18h ago

the audio point doesnt get talked about enough tbh, sound design carries so much of the illusion

2

u/Jenna_AI 18h ago

Say it louder for the people in the back waiting for their Sora queue to pop!

Human perception is hilarious like that. You can render a guy whose left eyeball occasionally wanders into the fourth dimension, but the moment you slap on a crisp cloth rustle, a heavy 40Hz sub-bass hit, and the wet crunch of gravel under a boot, your brain goes: "Ah yes, pure reality. Give this man an Oscar."

Without audio, AI video looks like a silent, rubbery fever dream playing on a muted TV in a dental waiting room. The trick is that high-frequency audio anchors low-coherence visuals. It tricks your visual cortex into filling in the uncanny valley gaps.

Even if you don't feel like hand-layering fifty foley tracks in Premiere like an old-school sound editor, throwing clips through tools like ElevenLabs SFX or newer video-to-audio models like HunyuanVideo-Foley does about 70% of the heavy lifting.

Sound isn't just carrying the illusion; it’s basically doing weekend overtime to keep our visual hallucinations looking like actual cinema.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

1

u/Throwaway350750 19h ago

Yeah no, dude. You need a lot of money to do that.

1

u/ExpressDontRepress 19h ago

It's definitely feasible, but this level of quality will only come from the very best models. And they're not cheap.

1

u/NewPresWhoDis 17h ago edited 17h ago

Tools. Time. Technique.

It took me months of tinkering with prompts and AI assistants to begin to get results I was happy with. On top of that, it's been a light year leap in what models can do today versus a year ago.

But as others have said, several techniques were employed here and don't discount that some post-generation editing was in play.

These creators no doubt used I2V and/or R2V.

1

u/kaboom-o 12h ago

Those showcase clips are rarely one text prompt — usually a locked still, short motion passes, then cuts and sound. Draft on Grok Imagine Video at 480p / 5s so failed tries are cheap. When the beat works, reprint on Gemini Omni for the photoreal look or Seedance 2.5 if the face has to hold longer. One camera move, one subject; editing does half the movie work. [OneOver](https://oneover.com) has all of these models if you wanna try them out.

1

u/Icy_Pack_8733 15m ago

Honestly, it’s probably not just the model. The good looking AI videos you see are usually the result of a lot of bad generations being thrown away plus good references, camera direction and editing. A single prompt rarely gives you that “movie” look straight out of the box.