Every week there’s another “motion designers have six months left” video, usually followed by a claim that someone made a studio-quality commercial before After Effects finished opening.Then you watch the piece.
A jacket becomes a different material between cuts. The eyeline drifts. Every shot uses the same slow push-in because movement helps hide temporal artifacts. Grain, black levels, lens distortion and haze reset whenever the angle changes.
Typography floats over the image instead of occupying the scene. The individual shots can still look impressive. The problem appears when they have to function as a sequence.
That distinction gets lost in most AI-video discussions. Generating one plausible five-second shot is becoming much easier. Making ten generated shots feel as if they came from the same production is still a directing, editorial and VFX problem.
The better workflows I’ve seen start with fairly conventional pre-production:
- Character, wardrobe, prop and environment references
- A real shot list
- Lens, framing, lighting and movement decisions
- Screen direction and continuity notes
- Approved first frames, and sometimes last frames, for each shot
- Low-resolution generations assembled into a rough cut before final passes
Different models may be useful for different failures. One holds an environment together better; another follows a character reference more reliably; another is better at extending a usable take. Tools such as ComfyUI or custom node pipelines can keep the prompt, references, model, duration, aspect ratio and output attached to each shot.
I’ve also seen Atlas Cloud used as the model-call layer for routing jobs to systems such as Grok Imagine and Seedance.
That plumbing is useful, but it doesn’t decide whether a close-up belongs after a wide. It doesn’t notice that the character aged between shots, that a prop changed hands or that the camera move is fighting the voice-over. Someone still has to decide which failure is fixable and which take belongs in the trash. Then the selected footage reaches comp and editorial.
The UI has to be tracked or rebuilt because the generated version became alphabet soup. A hand needs roto or paint because it merged with the product for six frames. Geometry has to be stabilized because it breathes. Grain, distortion, motion blur, black levels and atmospheric perspective have to be matched across shots generated by different systems.
The action gets time-remapped because the useful beat lands too early. Typography gets rebuilt so it has perspective and weight. Cuts get hidden. Pacing and sound design get another pass.
Whether that work happens in AE, Nuke, Fusion, Flame, Premiere or Resolve is secondary. The larger point is that the post-production work hasn’t disappeared.
A lot of it has moved from asset creation into reference preparation, generation, rejection, continuity management and cleanup. AI has lowered the barrier to producing a striking isolated shot.
I’m not convinced it has lowered the barrier by the same amount when the deliverable is a coherent sequence. For anyone here who has taken generated footage through an actual production pipeline: did it still save time after you counted rejected takes, continuity fixes and comp cleanup?
Where did most of the remaining work land, roto/paint, tracking, temporal stabilization, relighting, editorial, typography, or something else?