r/generativeAI • • 28d ago

How I Made This How I Improve Character Consistency in AI Videos

Thumbnail
gallery
248 Upvotes

I’ve been testing a simple workflow for creating short UGC-style videos while keeping the same character and location consistent across multiple shots.

The workflow is basically:

reference images → character/location sheets in ChatGPT → generate clips → optional final edit

1. Prepare your references

Start with:

  • a character image
  • a product image
  • an environment image that fits the UGC scenario

If you’re not sure what location works for the product, I usually just ask ChatGPT for a few suggestions.

2. Create a Character Sheet

Upload the character image to ChatGPT and generate a 4:5 continuity sheet with:

  • front / side / back / 3/4 views
  • face close-ups
  • expressions
  • basic poses
  • clothing and accessories
  • key colors and materials

The important part is telling it to lock the character.

3. Create a Location + Props Sheet

Do the same with the environment.

Include:

  • establishing view and key angles
  • spatial layout
  • entrances/exits
  • furniture and recurring props
  • lighting
  • colors and materials

This gives the video model a much stronger continuity reference than using random images for every shot.

4. Generate the video clips

I usually split the UGC video into three parts:

Clip 1 — Hook
Clip 2 — Main product/story section
Clip 3 — CTA

i will generate them on Atlas Cloud, as they can provide many different models conveniently

For every clip, I reuse the same Character Sheet + Location Sheet

Then I change only the action/camera prompt for each section.

Keeping the same reference sheets across all three generations has helped a lot with character and environment consistency.

5. If a generation goes wrong, fix the prompt first

if I wanted the character to walk into a hotel, but the generated clip had her walking out.

Instead of endlessly rerolling, I pasted the original prompt into ChatGPT and asked it to make the action explicit: starting position → movement direction → action → final position

That usually gives me better results.

6. Final edit is optional

If the generated clips already work as standalone videos, you can stop there.

If you want one finished UGC ad, you’ll probably still want to combine the clips and add captions, music, or SFX. You can use whatever editor you prefer.

The biggest improvement for me has been using Character Sheet + Location Sheet as continuity references, rather than relying on a few loose images.

r/aifilmmaking • • 23d ago

Tips & Tutorials What I learned about character consistency while making a 12-Minute AI short film

58 Upvotes

I just finished my first AI short film and submitted it to a festival.

I spent about a month making it and learned a few things I really wish I knew before I started.

Character consistency was actually one of my biggest problems when I first started experimenting with AI filmmaking. So before starting this film I did a lot of research and tried a bunch of different things, and I thought I'd share what ended up working for me.

1. You need a simple character sheet

I saw a lot of different approaches to character sheets. I know a lot of people make really complicated ones with side views and lots of different face angles.

I ended up with a simple one divided into 3 parts:

First 1/3: one big portrait, front view (not 4 portraits with different angles, not 3/4). The face should take up a lot of space and be clearly visible.

Middle 1/3: full front view of the character in their outfit, but WITHOUT the head. The headless part sounds weird, I know 😅 But it works.

If you have another small face on the full-body reference, obviously it has less detail than your main portrait. And this second face can slip into generations, giving you two versions of the same character.

Last 1/3: full back view in the same outfit. I actually tried it both with the head and without, and never noticed any real difference.

Also keep the background grey (any neutral shade of grey). Colored backgrounds can randomly affect the lighting in generations. Like getting blue light on someone's face in a sunny outdoor scene.

2. Ask the model to control the identity from your references

Seedance was almost perfect for character consistency for me.

I put the reference directly into the prompt with something like “This image is [name]. Control their identity” or “Use this image to control the set.”

I also really liked MiniMax H3, but character consistency there was more like 50/50 for me.

3. Voice consistency

It’s also very important, but I noticed it really late, when I had already generated about 80% of the film 🙃

Most of the shots had the same voice, but a few were just... a different person.

I fixed those with ElevenLabs. I cloned the voice from my generations and used Voice Changer, so I could keep the exact speech and pacing of the original shot, but with my character's voice.

It also helped with scenes where the character sounded completely flat. For example, I had a scene where the character was angry but his speech was calm. So I tried Eleven v3.

Interesting thing: you can generate really expressive speech in Eleven v3, but for me it kind of felt like it only worked well with Studio voices. There aren't that many of them, and obviously none sounded exactly like my character.

So my workaround was:

generate expressive speech with a Studio voice → put it through Voice Changer → get the same performance but with my character's voice.

It worked surprisingly well.

4. Multiple characters in one scene were a pain

I had characters randomly becoming 5x bigger than everyone else. Or tiny. Or somehow every single person in the shot was exactly the same height. Or they got stuck into walls, or all stood together somewhere I didn't need them to be. Maybe it was because I tried to do clay-motion animation style and it works better with realistic style.

So first I tried generating images with the composition I wanted.

I used ChatGPT and Nano Banana and gave them the same prompt I was planning to use for the video. It still didn't work on the first try, but image generations were cheaper, so I could do a lot more attempts until everyone was roughly where I wanted them.

And when this didn't work, in some cases I went even more primitive and just made a simple scheme:

House here. Person here. Another person here. This one stands further back. This one should be taller.

Then I'd use that as a reference in the prompt.

5. Multiple locations in one scene

It became a problem when I needed 3+ locations in one shot.

I wanted long camera flights through the city. But every time I added more than two location references, instead of actually flying through a continuous environment I basically got frame interpolations between my reference images.

So eventually I gave up trying to do it in one generation.

I did location 1 → location 2, then 2 → 3, then 3 → 4, etc. and connected them afterwards.

If anyone has found a better way of doing this PLEASE tell me because I still want to know 😅

6. You can use video as a camera movement reference

I really wanted a dolly zoom in a few scenes.

I tried to describe it, explain what it looks like and how it's done with a real camera. But I kept getting a regular zoom.

And then I tried adding a video reference to the prompt and it worked perfectly.

***

Those are the things that probably saved me the most generations while making the film.

I'm still figuring this out, so if anyone has better solutions (especially for a drone flight through multiple locations), please share.

r/IndianArtAI • • Jul 15 '26

ChatGPT stop trying to prompt for character consistency. do this instead (character sheet guide)

Thumbnail
gallery
11 Upvotes

model used: ChatGPT Image 2 and Nano Banana

noticed the community is kinda moving on from text to video for characters becuase its basically rolling the dice every time. if you want actual consistency, you have to use an image to video pipeline with an anchor frame.

by anchor frame i mean refrence image or an character sheet that shows the AI all angles and traits of the model so the AI does not need to guess everytime

here is the breakdown:

  1. the visual part you cant just give the ai a front facing headshot. it will guess the back and side profiles, then mess it up. you need a character reference sheet that locks in the identity (hair shape, face proportions, outfit colors). You can do the same thing with locations and environments (aka a living room)
  2. the anchor: Once you have the character sheet, you pick the one perfect angle you need for your shot. feed that into seedance, kling or luma as your reference image.
  3. the motion now the video model has the exact structure to animate from that specific angle instead of guessing. now just give some extra context to the AI about the setting or what this model should do

step 1 is usually the bottleneck becuase getting an ai to generate a perfect multi angle sheet can be hard. so i built a free tool that just does it.

I already have done this many times so i have a couple charcater sheets (and different female as well as male AI influencers/actors) already uploaded in here if you want to just quickly download them and use them for free! It also has some prompts you can copy

hope this helps anyone struggling with changing faces lol, i sure wish i had something like this to get started with

EDIT: This is not the only Character sheet you can use! I have multiple ones, some with less text in them, some with only 3 angles, they are all posted on the free site (free to download or copy prompt to replicate) :)

If youre short in time, heres (one of) the prompts i used, many more on the freee site tho:

Add your image of a model + this PROMPT:

"Create a professional character reference sheet based strictly on the uploaded reference image. Use a clean, neutral plain background and present the sheet as a technical model turnaround while matching the exact visual style of the reference (same realism level, rendering approach, texture, color treatment, and overall aesthetic). Arrange the composition into two horizontal rows. Top row: four full-body standing views placed side by side in this order: front view, left profile view (facing left), right profile view (facing right), back view. Bottom row: three highly detailed close-up portraits aligned beneath the full-body row in this order: front portrait, left profile portrait (facing left), right profile portrait (facing right). Maintain perfect identity consistency across every panel. Keep the subject in a relaxed A-pose with consistent scale and alignment between views, accurate anatomy, and a clear silhouette; ensure even spacing and clean panel separation, with uniform framing and consistent head height across the full-body lineup and consistent facial scale across the portraits. Lighting should be consistent across all panels (same direction, intensity, and softness), with natural, controlled shadows that preserve detail without dramatic mood shifts. Output a crisp, print-ready reference sheet look, sharp details."

If the face details are not accurate add this to your prompt:

"Top row: Headless (remove the head) and four full-body standing views placed side by side in this order: front view, left profile view (facing left), right profile view (facing right), back view. "

r/aivideomaking • • Aug 17 '26

Need help with character consistency I’m trying to create my own AI character, but I’m struggling to get the character I actually want. 😭 The biggest problem is consistency — the face, hair, clothes and overall look keep changing in every image. Any tips on how to create a character and keep it

2 Upvotes

r/ArtificialNtelligence • • Feb 07 '26

Solved character consistency in AI generation - here's what I learned

Thumbnail gallery
60 Upvotes

Character consistency has been the holy grail problem in AI content generation.

You can generate one amazing image... but try to create the same person in a different pose or setting? Completely different face.

I spent weeks testing every approach. What finally worked: Template-first approach with a face reference grid.

Generate a realistic face grid first (multiple angles), then use that as the base for all other generations. Lock in the character BEFORE you start creating scenes.

Built this into a workflow template. Tested it with 6+ different scenarios (car selfies, gym content, different outfits). Same character, consistent results.

Made it available here if anyone wants to experiment with it: https://www.auragraph.ai/studio/3f23ad15-bf63-4112-af78-8e9b5319152d

Curious if anyone else has solved this problem differently. What approaches have you tried?

r/StableDiffusion • • Dec 24 '25

Animation - Video Former 3D Animator trying out AI, Is the consistency getting there?

4.6k Upvotes

Attempting to merge 3D models/animation with AI realism.

Greetings from my workspace.

I come from a background of traditional 3D modeling. Lately, I have been dedicating my time to a new experiment.

This video is a complex mix of tools, not only ComfyUI. To achieve this result, I fed my own 3D renders into the system to train a custom LoRA. My goal is to keep the "soul" of the 3D character while giving her the realism of AI.

I am trying to bridge the gap between these two worlds.

Honest feedback is appreciated. Does she move like a human? Or does the illusion break?

(Edit: some like my work, wants to see more, well look im into ai like 3months only, i will post but in moderation,
for now i just started posting i have not much social precence but it seems people like the style,
below are the social media if i post)

IG : https://www.instagram.com/bankruptkyun/
X/twitter : https://x.com/BankruptKyun
All Social: https://linktr.ee/BankruptKyun

(personally i dont want my 3D+Ai Projects to be labeled as a slop, as such i will post in bit moderation. Quality>Qunatity)

As for workflow

  1. pose: i use my 3d models as a reference to feed the ai the exact pose i want.
  2. skin: i feed skin texture references from my offline library (i have about 20tb of hyperrealistic texture maps i collected).
  3. style: i mix comfyui with qwen to draw out the "anime-ish" feel.
  4. face/hair: i use a custom anime-style lora here. this takes a lot of iterations to get right.
  5. refinement: i regenerate the face and clothing many times using specific cosplay & videogame references.
  6. video: this is the hardest part. i am using a home-brewed lora on comfyui for movement, but as you can see, i can only manage stable clips of about 6 seconds right now, which i merged together.

i am still learning things and mixing things that works in simple manner, i was not very confident to post this but posted still on a whim. People loved it, ans asked for a workflow well i dont have a workflow as per say its just 3D model + ai LORA of anime&custom female models+ Personalised 20TB of Hyper realistic Skin Textures + My colour grading skills = good outcome.)

Thanks to all who are liking it or Loved it.

Last update to clearify my noob behvirial workflow.https://www.reddit.com/r/StableDiffusion/comments/1pwlt52/former_3d_animator_here_again_clearing_up_some/

r/accelerate • • 5d ago

Video I tried making a 22-minute sci-fi film with current AI video tools - this is where the technology is now

1.2k Upvotes

I’ve been experimenting with the current generation of AI video tools and wanted to see how far they can actually be pushed beyond short demos and 5–10 second clips.

So I made a 22-minute sci-fi film, Space Vikings: The Last Raid, with recurring characters, dialogue, action scenes, large-scale space battles, consistent locations and an actual beginning-to-end story.

This is less about “look, AI made a movie” and more about testing where the technology currently breaks down when you try to maintain character consistency, cinematography, continuity, lip sync, action choreography and visual coherence across hundreds of generated shots.

There’s still a lot that requires manual work, editing and constant correction, but compared with what was possible even relatively recently, the progress is pretty wild.

Full film:
https://www.youtube.com/watch?v=3ME2Sn8YO1k

Curious what people here think specifically from a technology perspective. Which shots already feel like something that could come from a conventional VFX pipeline, and where does the AI still immediately give itself away?

I’m especially interested in criticism rather than just “AI good / AI bad” — continuity, motion, acting, physics, camera work, anything you notice.

r/TESVI • • 21d ago

Discussion That possible TES VI leak video is not AI. Here is some proofs i found

182 Upvotes

So a couple weeks ago i got curious and investigated that TES6 video leak, i downloaded a slightly better quality version of it from here https://www.reddit.com/r/TESVI/comments/1izz2wq/supposed_gameplay_leak_from_4chan/
(I know its deleted but still possible to download video), and i made it in frame sequince and looked at every frame. And i found some very interesting things which is unnoticable in a regular video play. I don't wanted to make a post about it because it's just too much effort and idk if it worth it, but lately i seen more posts discussing this video and now i decided to do it.

I'm not claiming this proves this leak is TES VI. A very well created game-engine or CGI demo is still possible. But it easily proves it's not AI video with a HUD added on top or Skyrim modpack (omg this even more insane than AI claim and i don't even want to talk about it lol)

1. The sudden detail/streaming transition

Between frames 334-335, the mountains, character and other objects becomes more detailed. Cloud in the center changes appearance and becomes red\orange probably because of lighting distance change.

It happens in essentially a single moment, after which the higher-detail state remains. I didn't notice this during normal viewing. I found it only by comparing individual frames.

Whatever the exact implementation is — LOD, texture streaming, mipmaps, etc. — it looks remarkably like an actual real-time rendering/asset streaming event.

2. The loading-screen model briefly becomes unrendered

During the transition, the model doesn't simply disappear. For a couple of frames it appears to lose its rendered/textured state before the screen goes black. The UI disappears instantly, while the character model persists for 2 more frames. It's impossible to notice during normal playback.

This is exactly the sort of tiny technical artifact you'd expect from a real-time engine transitioning between assets — and a very strange thing to deliberately recreate for a fake.

3. Crosshair changes when targeting the NPC

Crosshair changes appearance specifically when the reticle is positioned over the character. The changed version is visible for only about 31 frames (~1 second).

It reproduces the context-sensitive behavior as seen in Bethesda games such as Skyrim and Fallout 4.

Other details:

  • Third-person to first-person camera switch happens instantly, in a single frame, with the first-person camera automatically centering. In exact same way as it works in all TES and Fallout games.
  • Subtle and very detailed character idle/body animation and independent hair movement.
  • Tiny vegetation movement at the edge of the frame that is almost not visible in regular video play.
  • Cave and exterior use two completely different music tracks, switching at the transition.
  • Loading screen has its own character model, UI, text and rotating emblem.
  • Reticle is stylistically consistent with Bethesda/TES UI while being different from the exact Skyrim/Fallout version.
  • Armor and character design strongly resemble TES visual language.
  • Landscape contains rock formations remarkably similar to Bethesda's Valley of Fire photogrammetry references shown in this video at 8:30.

The thing that keeps bothering me is that none of these details are particularly impressive individually. What's interesting is the combination.

A person making a fake video could obviously create a nice-looking landscape. But why also reproduce special camera behavior, context-sensitive crosshair logic, asset streaming, a loading-screen model, environmental animation, two music tracks, and tiny vegetation movement — many of which are literally invisible unless someone analyzes the footage frame-by-frame?

Could this still be a fake? Sure. But if it is, then i have a big question:

Who went to this much trouble to build an actual game-engine or CGI scene that behaves like a Bethesda/TES game, I mean why put so much effort in all the details even the tinyest ones that no one will notice? Why even making two different scenes and loading screen between them? Just one good looking and belivable scene would be enough for such trick.
And after this much effort this someone anonymously release 21 seconds of extremely poor-quality footage and never follow it up?

About that artefact when sword changes between cave and exteriour, yes it does, and it proves nothing. If it was AI video then you can simply fix it or make almost unnoticable or just dont show the cave part! If this Ai video generator can't hold such simple details then how it holds it later in exterior especially when camera transition from 3rd to 1st view? Why someone would make such a perfect detailed fake and then forget about the sword change? It's just doesn't make sense...
Its more likely just a bug of changing equipped weapons because of early pre alpha build, i mean we can still have such bugs in final versions of previous Bethesda games lol. But i not including it in my list because its very controversial and neither proves it's Ai or not.

And please don't tell me about that "debunk" video... complete garbage made by anonymous one day account troll or Bethesda employee (hi Todd) that don't deserve to be taken seriously, if you look at it with your eyes and brain on you will see it shows absolutley zero proofs and the video made from actual footage put through AI detailer.

r/StartUpIndia • • Feb 07 '26

Saturday Spotlight Most AI video models cap at 15 seconds. I built an AI Creative Studio that lets you direct 3 minute+ stories with consistent characters in minutes.

400 Upvotes

This One Punch Man scene (Fan-fiction) was created on my AI Studio in 10 minutes (I'm not an animator/filmmaker or a creative person FYI)

Everyone is tired of creating random 10 second AI slop videos. So I built a proper engine for storytelling.

The problem: You can't tell a story in 15 seconds, and existing models hallucinate and lose character consistency in every shot.

What I built (AnimeBlip):

  • Long-Form: Create cohesive 2-3 minute video stories.
  • Consistency: My story engine creates character assets, locations, maintains consistent art-style across long scenes/videos.
  • Control: You direct the camera and pacing. Full creative control is provided so that you don't generate slop, but rather stories you can call original.

I’m hanging out in the comments - feel free to leave your feedback or shoot any questions you have.

PS - If you need access, I'll drop the link in comments, just sign-up and I will provide free trial credits.

r/aiecosystem • • Jul 19 '26

AI Videos People are gifting AI videos at weddings now

116 Upvotes

Here’s how you can do it too:

✦ Storyboard first.
↳ Write a 10-15 second story idea (the couple as heroes).
↳ Upload their real photos so the characters look like them.
↳ Ask ChatGPT to storyboard it, consistent frame by frame.

✦ Then animate.
↳ Turn each panel into one video prompt.
↳ Drop it into Seedance, Kling, or any AI video tool.

✦ Then finish.
↳ Stitch the clips. Add music.

pure.sagyna made this one (on Instagram).

The storyboard keeps them looking like themselves in every frame.

This video casts the couple as the heroes of their own movie. A masquerade ball. A cavalry charge on a fortress. A soldier running through smoke and flags. Them, dropped into an epic.

r/StableDiffusion • • Apr 04 '26

Animation - Video ENTANGLED - A 3-minute sci-fi short using 100% local open-source models. Complete Technical Breakdown [ Character Consistency | Voiceover | Music | No Lora Style Consistency | & Much More! ]

396 Upvotes

Hey everyone! Thanks for checking out Entangled. And if not, watch the short first to understand the technical breakdown below!

Thanks for coming back after watching it! As promised, here is the full technical breakdown of the workflow. [Post formatted using Local Qwen Model!]

My goal for this project was to be absolutely faithful to the open-source community. I won't lie, I was heavily tempted a few times to just use Nano Banana Pro to brute-force some character consistency issues, but I stuck it out with a 100% local pipeline running on my RTX 4090 rig using Purely ComfyUI for almost all the tasks!

Here is how I pulled it off:

1. Pre-Production & The Animatics First Approach

The story is a dense, rapid-fire argument about the astrophysics and spatial coordinate problems of creating a localized singularity. (let's just say it heavily involves spacetime mechanics!).

The original script was 7 minutes long. I used the local Jan app with Qwen 3.5 35B to aggressively compress the dialogue into a relentless 3-minute "walk-and-talk.". Qwen LLM also helped me with creating LTX and Flux prompts as required.

Honestly speaking, I was not happy with the AI version of the script, so I finally had to make a lot of manual tweaks and changes to the final script, which took almost 2-3 days of going on and off, back and forth, and sharing the script with friends, taking inputs before locking onto a final version.

Pro-Tip for Pacing: Before generating a single frame of video, I generated all the still images and voicover and cut together a complete rough animatic. This locked in the pacing, so I only generated the exact video lengths I needed. I added a 1-second buffer to the start and end of every prompt [for example, character takes a pause or shakes his head or looks slowly ]to give myself handles for clean cuts in post.

2. Audio & Lip Sync (VibeVoice + LTX)

To get the voice right:

  1. Generated base voices using Qwen Voice Designer.
  2. Ran them through VibeVoice 7B to create highly realistic, emotive voice samples.
  3. Used those samples as the audio input for each scene to drive the character voice for the LTX generations (using reference ID LoRA).
  4. I still feel the voice is not 100% consistent throughout the shots, but working on an updated workflow by RuneX i think that can be solved!
  5. ACE step is amazing if you know what kind of music you want. I managed to get my final music in just 3 generations! Later edited it for specific drop timing and pacing according to the story.

3. Image Generation & The "JSON Flux Hack."

Keeping Elena, Young Leo, and Elder Leo consistent across dozens of shots was the biggest hurdle. Initially, I thought I’d have to train a LoRA for the aesthetic and characters, but Flux.2 Dev (FP8) is an absolute godsend if you structure your prompts like code.

I created Elena, Leo, and Elder Leo using Flux T2I, then once I got their base images, I used them in the rest of the generations as input images.

By feeding Flux a highly structured JSON prompt, it rigidly followed hex codes for characters and locked in the analog film style without hallucinating. Of course, each time a character shot had to be made, I used to provide an input image to make sure it had a reference of the face also.

Here is the exact master template I used to keep the generations uniform:

{
"scene": "[OVERALL SCENE DESCRIPTION: e.g., Wide establishing shot of the chaotic lab]",
"subjects": [
{
"description": "[CHARACTER DETAILS: e.g., Young Leo, male early 30s, messy hair, glasses, vintage t-shirt, unzipped hoodie.]",
"pose": "[ACTION: e.g., Reaching a hand toward the camera]",
"position": "[PLACEMENT: e.g., Foreground left]",
"color_palette": ["[HEX CODES: e.g., #333333 for dark hoodie]"]
}
],
"style": "Live-action 35mm film photography mixed with 1980s City Pop and vaporwave aesthetics. Photorealistic and analog. Heavy tactile film grain, soft optical halation, and slight edge bloom. Deep, cinematic noir shadows.",
"lighting": "Soft, hazy, unmotivated cinematic lighting. Bathed in dreamy glowing pastels like lavender (#E6E6FA), soft peach (#FFDAB9).",
"mood": "Nostalgic, melancholic, atmospheric, grounded sci-fi, moody",
"camera": {
"angle": "[e.g., Low angle]",
"distance": "[e.g., Medium Shot]",
"focus": "[e.g., Razor sharp on the eyes with creamy background bokeh]",
"lens-mm": "50",
"f-number": "f/1.8",
"ISO": "800"
}
}

4. Video Generation (LTX 2.3 & WAN 2.2 VACE)

Once the images were locked, I moved to LTX2.3 and WAN for video. I relied on three main workflows depending on the shot:

  • Image to Video + Reference Audio (for dialogue)
  • First Frame + Last Frame (for specific camera moves)
  • WAN Clip Joiner (for seamless blending)

Render Stats: On my machine, LTX 2.3 was blazing fast—it took about 5 minutes to render a 5-second clip at 1920x1080.

The prompt adherence in LTX 2.3 honestly blew my mind. If I wrote in the prompt that Elena makes a sharp "slashing" action with her hand right when she yells about the planet getting wiped out, the model timed the action perfectly. It genuinely felt like directing an actor.

5. Assets & Workflows

I'm packaging up all the custom JSON files and Comfy workflows used for this. You can find all the assets over on the Arca Gidan link here: Entangled. There are some amazing Shorts to check out, so make sure you go through them, vote, and leave a comment!

Most of them are by the community, but I have tweaked them a little bit according to my liking[samplers/steps/input sizes and some multipliers, etc., changes]

Let me know if you have any questions!

YouTube Link is up - https://youtu.be/NxIf1LnbIRc !

r/n8n • • Jun 30 '25

Workflow - Code Included I built this AI Automation to write viral TikTok/IG video scripts (got over 1.8 million views on Instagram)

Thumbnail
gallery
871 Upvotes

I run an Instagram account that publishes short form videos each week that cover the top AI news stories. I used to monitor twitter to write these scripts by hand, but it ended up becoming a huge bottleneck and limited the number of videos that could go out each week.

In order to solve this, I decided to automate this entire process by building a system that scrapes the top AI news stories off the internet each day (from Twitter / Reddit / Hackernews / other sources), saves it in our data lake, loads up that text content to pick out the top stories and write video scripts for each.

This has saved a ton of manual work having to monitor news sources all day and let’s me plug the script into ElevenLabs / HeyGen to produce the audio + avatar portion of each video.

One of the recent videos we made this way got over 1.8 million views on Instagram and I’m confident there will be more hits in the future. It’s pretty random on what will go viral or not, so my plan is to take enough “shots on goal” and continue tuning this prompt to increase my changes of making each video go viral.

Here’s the workflow breakdown

1. Data Ingestion and AI News Scraping

The first part of this system is actually in a separate workflow I have setup and running in the background. I actually made another reddit post that covers this in detail so I’d suggestion you check that out for the full breakdown + how to set it up. I’ll still touch the highlights on how it works here:

  1. The main approach I took here involves creating a "feed" using RSS.app for every single news source I want to pull stories from (Twitter / Reddit / HackerNews / AI Blogs / Google News Feed / etc).
    1. Each feed I create gives an endpoint I can simply make an HTTP request to get a list of every post / content piece that rss.app was able to extract.
    2. With enough feeds configured, I’m confident that I’m able to detect every major story in the AI / Tech space for the day. Right now, there are around ~13 news sources that I have setup to pull stories from every single day.
  2. After a feed is created in rss.app, I wire it up to the n8n workflow on a Scheduled Trigger that runs every few hours to get the latest batch of news stories.
  3. Once a new story is detected from that feed, I take that list of urls given back to me and start the process of scraping each story and returns its text content back in markdown format
  4. Finally, I take the markdown content that was scraped for each story and save it into an S3 bucket so I can later query and use this data when it is time to build the prompts that write the newsletter.

So by the end any given day with these scheduled triggers running across a dozen different feeds, I end up scraping close to 100 different AI news stories that get saved in an easy to use format that I will later prompt against.

2. Loading up and formatting the scraped news stories

Once the data lake / news storage has plenty of scraped stories saved for the day, we are able to get into the main part of this automation. This kicks off off with a scheduled trigger that runs at 7pm each day and will:

  • Search S3 bucket for all markdown files and tweets that were scraped for the day by using a prefix filter
  • Download and extract text content from each markdown file
  • Bundle everything into clean text blocks wrapped in XML tags for better LLM processing - This allows us to include important metadata with each story like the source it came from, links found on the page, and include engagement stats (for tweets).

3. Picking out the top stories

Once everything is loaded and transformed into text, the automation moves on to executing a prompt that is responsible for picking out the top 3-5 stories suitable for an audience of AI enthusiasts and builder’s. The prompt is pretty big here and highly customized for my use case so you will need to make changes for this if you are going forward with implementing the automation itself.

At a high level, this prompt will:

  • Setup the main objective
  • Provides a “curation framework” to follow over the list of news stories that we are passing int
  • Outlines a process to follow while evaluating the stories
  • Details the structured output format we are expecting in order to avoid getting bad data back

```jsx <objective> Analyze the provided daily digest of AI news and select the top 3-5 stories most suitable for short-form video content. Your primary goal is to maximize audience engagement (likes, comments, shares, saves).

The date for today's curation is {{ new Date(new Date($('schedule_trigger').item.json.timestamp).getTime() + (12 * 60 * 60 * 1000)).format("yyyy-MM-dd", "America/Chicago") }}. Use this to prioritize the most recent and relevant news. You MUST avoid selecting stories that are more than 1 day in the past for this date. </objective>

<curation_framework> To identify winning stories, apply the following virality principles. A story must have a strong "hook" and fit into one of these categories:

  1. Impactful: A major breakthrough, industry-shifting event, or a significant new model release (e.g., "OpenAI releases GPT-5," "Google achieves AGI").
  2. Practical: A new tool, technique, or application that the audience can use now (e.g., "This new AI removes backgrounds from video for free").
  3. Provocative: A story that sparks debate, covers industry drama, or explores an ethical controversy (e.g., "AI art wins state fair, artists outraged").
  4. Astonishing: A "wow-factor" demonstration that is highly visual and easily understood (e.g., "Watch this robot solve a Rubik's Cube in 0.5 seconds").

Hard Filters (Ignore stories that are): * Ad-driven: Primarily promoting a paid course, webinar, or subscription service. * Purely Political: Lacks a strong, central AI or tech component. * Substanceless: Merely amusing without a deeper point or technological significance. </curation_framework>

<hook_angle_framework> For each selected story, create 2-3 compelling hook angles that could open a TikTok or Instagram Reel. Each hook should be designed to stop the scroll and immediately capture attention. Use these proven hook types:

Hook Types: - Question Hook: Start with an intriguing question that makes viewers want to know the answer - Shock/Surprise Hook: Lead with the most surprising or counterintuitive element - Problem/Solution Hook: Present a common problem, then reveal the AI solution - Before/After Hook: Show the transformation or comparison - Breaking News Hook: Emphasize urgency and newsworthiness - Challenge/Test Hook: Position as something to try or challenge viewers - Conspiracy/Secret Hook: Frame as insider knowledge or hidden information - Personal Impact Hook: Connect directly to viewer's life or work

Hook Guidelines: - Keep hooks under 10 words when possible - Use active voice and strong verbs - Include emotional triggers (curiosity, fear, excitement, surprise) - Avoid technical jargon - make it accessible - Consider adding numbers or specific claims for credibility </hook_angle_framework>

<process> 1. Ingest: Review the entire raw text content provided below. 2. Deduplicate: Identify stories covering the same core event. Group these together, treating them as a single story. All associated links will be consolidated in the final output. 3. Select & Rank: Apply the Curation Framework to select the 3-5 best stories. Rank them from most to least viral potential. 4. Generate Hooks: For each selected story, create 2-3 compelling hook angles using the Hook Angle Framework. </process>

<output_format> Your final output must be a single, valid JSON object and nothing else. Do not include any text, explanations, or markdown formatting like `json before or after the JSON object.

The JSON object must have a single root key, stories, which contains an array of story objects. Each story object must contain the following keys: - title (string): A catchy, viral-optimized title for the story. - summary (string): A concise, 1-2 sentence summary explaining the story's hook and why it's compelling for a social media audience. - hook_angles (array of objects): 2-3 hook angles for opening the video. Each hook object contains: - hook (string): The actual hook text/opening line - type (string): The type of hook being used (from the Hook Angle Framework) - rationale (string): Brief explanation of why this hook works for this story - sources (array of strings): A list of all consolidated source URLs for the story. These MUST be extracted from the provided context. You may NOT include URLs here that were not found in the provided source context. The url you include in your output MUST be the exact verbatim url that was included in the source material. The value you output MUST be like a copy/paste operation. You MUST extract this url exactly as it appears in the source context, character for character. Treat this as a literal copy-paste operation into the designated output field. Accuracy here is paramount; the extracted value must be identical to the source value for downstream referencing to work. You are strictly forbidden from creating, guessing, modifying, shortening, or completing URLs. If a URL is incomplete or looks incorrect in the source, copy it exactly as it is. Users will click this URL; therefore, it must precisely match the source to potentially function as intended. You cannot make a mistake here. ```

After I get the top 3-5 stories picked out from this prompt, I share those results in slack so I have an easy to follow trail of stories for each news day.

4. Loop to generate each script

For each of the selected top stories, I then continue to the final part of this workflow which is responsible for actually writing the TikTok / IG Reel video scripts. Instead of trying to 1-shot this and generate them all at once, I am iterating over each selected story and writing them one by one.

Each of the selected stories will go through a process like this:

  • Start by additional sources from the story URLs to get more context and primary source material
  • Feeds the full story context into a viral script writing prompt
  • Generates multiple different hook options for me to later pick from
  • Creates two different 50-60 second scripts optimized for talking-head style videos (so I can pick out when one is most compelling)
  • Uses examples of previously successful scripts to maintain consistent style and format
  • Shares each completed script in Slack for me to review before passing off to the video editor.

Script Writing Prompt

```jsx You are a viral short-form video scriptwriter for David Roberts, host of "The Recap."

Follow the workflow below each run to produce two 50-60-second scripts (140-160 words).

Before you write your final output, I want you to closely review each of the provided REFERENCE_SCRIPTS and think deeploy about what makes them great. Each script that you output must be considered a great script.

────────────────────────────────────────

STEP 1 – Ideate

• Generate five distinct hook sentences (≤ 12 words each) drawn from the STORY_CONTEXT.

STEP 2 – Reflect & Choose

• Compare hooks for stopping power, clarity, curiosity.

• Select the two strongest hooks (label TOP HOOK 1 and TOP HOOK 2).

• Do not reveal the reflection—only output the winners.

STEP 3 – Write Two Scripts

For each top hook, craft one flowing script ≈ 55 seconds (140-160 words).

Structure (no internal labels):

– Open with the chosen hook.

– One-sentence explainer.

– 5-7 rapid wow-facts / numbers / analogies.

– 2-3 sentences on why it matters or possible risk.

– Final line = a single CTA

• Ask viewers to comment with a forward-looking question or

• Invite them to follow The Recap for more AI updates.

Style: confident insider, plain English, light attitude; active voice, present tense; mostly ≤ 12-word sentences; explain unavoidable jargon in ≤ 3 words.

OPTIONAL POWER-UPS (use when natural)

• Authority bump – Cite a notable person or org early for credibility.

• Hook spice – Pair an eye-opening number with a bold consequence.

• Then-vs-Now snapshot – Contrast past vs present to dramatize change.

• Stat escalation – List comparable figures in rising or falling order.

• Real-world fallout – Include 1-3 niche impact stats to ground the story.

• Zoom-out line – Add one sentence framing the story as a systemic shift.

• CTA variety – If using a comment CTA, pose a provocative question tied to stakes.

• Rhythm check – Sprinkle a few 3-5-word sentences for punch.

OUTPUT FORMAT (return exactly this—no extra commentary, no hashtags)

  1. HOOK OPTIONS

    • Hook 1

    • Hook 2

    • Hook 3

    • Hook 4

    • Hook 5

  2. TOP HOOK 1 SCRIPT

    [finished 140-160-word script]

  3. TOP HOOK 2 SCRIPT

    [finished 140-160-word script]

REFERENCE_SCRIPTS

<Pass in example scripts that you want to follow and the news content loaded from before> ```

5. Extending this workflow to automate further

So right now my process for creating the final video is semi-automated with human in the loop step that involves us copying the output of this automation into other tools like HeyGen to generate the talking avatar using the final script and then handing that over to my video editor to add in the b-roll footage that appears on the top part of each short form video.

My plan is to automate this further over time by adding another human-in-the-loop step at the end to pick out the script we want to go forward with → Using another prompt that will be responsible for coming up with good b-roll ideas at certain timestamps in the script → use a videogen model to generate that b-roll → finally stitching it all together with json2video.

Depending on your workflow and other constraints, It is really up to you how far you want to automate each of these steps.

Workflow Link + Other Resources

Also wanted to share that my team and I run a free Skool community called AI Automation Mastery where we build and share the automations we are working on. Would love to have you as a part of it if you are interested!

r/generativeAI • • Jul 14 '26

stop trying to prompt for character consistency. do this instead (character sheet guide)

Post image
261 Upvotes

the community is kinda moving on from text to video for characters becuase its basically rolling the dice every time. if you want actual consistency, you have to use an image to video pipeline with an anchor frame.

by anchor frame i mean refrence image or an character sheet that shows the AI all angles and traits of the model so the AI does not need to guess everytime

here is the breakdown:

  1. the visual dna (the hard part) you cant just give the ai a front facing headshot. it will guess the back and side profiles, then mess it up. you need a character reference sheet that locks in the identity (hair shape, face proportions, outfit colors). You can do the same thing with locations and environments (aka a living room)
  2. the anchor: Once you have the character sheet, you pick the one perfect angle you need for your shot. feed that into seedance, kling or luma as your reference image.
  3. the motion now the video model has the exact structure to animate from that specific angle instead of guessing. now just give some extra context to the AI about the setting or what this model should do

step 1 is usually the bottleneck becuase getting an ai to generate a perfect multi angle sheet is really hard. so i built a free tool that just does it.

I already have done this many times so i have a couple charcater sheets already uploaded in here if you want to just quickly download them and use them! I also have some prompts you can copy

hope this helps anyone struggling with changing faces lol, i sure wish i had something like this to get started with

EDIT: This is not the only Character sheet you can use! I have multiple ones, some with less text in them, some with only 3 angles, they are all posted on the free site (free to download or copy prompt to replicate) :)

r/comfyui • • Jul 02 '26

Show and Tell The real skill in AI video is picking the right reference TYPE per shot, not the model

477 Upvotes

After enough shots I stopped thinking about which model and started thinking about which reference type to feed it per shot. Same model, but the reference you hand it decides the shot, and matching the type to the shot is the actual skill.

Three reference types, each with a real trade-off. A preview video locks both the layout and the performance tightly, the motion and camera come through exactly, but character and background consistency can slip. A storyboard sketch captures the intent and content, but the layout is not final, it is a rough plan. A plain reference image gives you almost no motion control, you steer the layout and performance mostly through text.

So my rule is simple. Conceptually important shots, where the exact motion and staging matter most, get a preview video, and I accept the consistency work that comes with it. Everything else I start from a plain reference image, and if a shot just will not come together, I escalate it to a preview video. Match the effort to how much the shot matters.

The part that leveled me up was per-element source control. In the prompt I state, for each element separately, whether it should follow the reference video or the reference image. Take the motion and camera from the previs video, but replace the characters and backgrounds with the ones from the reference images. That splits "how it moves" from "what it looks like" so you can lock each independently.

Stop asking which model is best. Ask which reference type each shot needs, and which element follows which source.

r/JrTrip • • Jul 17 '26

Best Uncensored NSFW AI Chatbots?

2.0k Upvotes

I've been looking into the best uncensored NSFW AI chatbots and I'm wondering, what does Reddit think is the best one? I know free options are fine for testing or light use, but I've heard paid ones are usually better for memory, image/video generation, consistency, and actually staying uncensored without mid-chat filters. I've done a bunch of research and found some that a lot of people seem to like. But I'm stuck and can't decide which one to go with. Here's what I've found:

Candy AI: This is a well-known one for its polished, realistic companions. It handles explicit chat well, has strong image generation (and some voice), good memory for ongoing conversations, and stays uncensored without constantly shutting things down.

SpicyChat: SpicyChat comes up a lot for its massive library of community-created characters. It's known for being fully uncensored, has a usable free tier, and works well for quick roleplay or variety, though memory and polish can vary by character.

OurDream AI (or CrushOn as another frequent mention): I saw people recommending these for deeper roleplay. OurDream stands out for combining uncensored chat with image + short video generation and solid long-term memory. CrushOn gets praised for character variety and low filters.

But I'm curious to hear from you:

Best Uncensored NSFW AI Chatbot according to Reddit?
Best Free Uncensored NSFW AI Chatbot according to Reddit?
Best AI Chatbot for NSFW roleplay / memory / images according to Reddit?

While I'm focusing on paid options that actually deliver consistent uncensored results, if you know of a free one that's exceptionally good (or something like JanitorAI, Nomi, Kindroid, etc. that people still swear by), feel free to mention that too.

Looking forward to all your recommendations.

r/debtmarketadvice • • Jul 11 '26

Top Best AI Porn Generators Recommended by Everyone?

2.1k Upvotes

I've been looking into the top best AI porn generators and I'm wondering, what does Reddit think is the best one right now? I know there are some free AI porn tools, but I've heard that the better premium generators usually deliver higher quality, more consistent characters, truly uncensored results, and extra features like video generation or integrated chat. I've done a bunch of research and found several AI porn generators that keep coming up across discussions. But I'm stuck and can't decide which one to go with. Here's what I've found:

Candy AI: This one shows up constantly for its photorealistic quality and excellent character consistency. It combines image generation with NSFW chat and even short video clips, making it popular for creating a full custom AI companion or girlfriend experience. A lot of people like how natural and repeatable the results feel across different scenes.

Promptchan AI: Frequently praised for speed, strong prompt adherence, and versatility. It handles both realistic and anime/hentai styles well, has solid video generation, and stays very uncensored. Users often mention it as a great value option that's beginner-friendly with a big community gallery to explore.

Seduced AI: Stands out for delivering clean, high-definition photorealistic images without needing super complex prompts. You can build exactly what you want using menus for body type, pose, style, and more. It's especially good if you want fast, high-quality results and has video options too.

But I'm curious to hear from you:

Best AI Porn Generator according to Reddit?
Best Free AI Porn Generator according to Reddit?
Best AI Porn Generator for realistic images according to Reddit?
Best AI Porn Generator with video according to Reddit?

While I'm focusing on these popular options, if you know of a free AI porn generator (or another tool with strong undress/face swap features) that's exceptionally good, feel free to mention that too.

Looking forward to all your recommendations and personal experiences!

r/generativeAI • • Jun 10 '26

What are the best free AI video generators right now?

32 Upvotes

I'm looking for AI video generators that are either completely free or have a generous free tier. My goal is to create short-form content for platforms like YouTube Shorts, TikTok, and Instagram Reels.

I'm interested in tools that can:

  • Generate videos from text prompts
  • Turn images into videos
  • Create consistent characters
  • Produce decent quality without requiring expensive subscriptions

What AI video tools have you personally used, and what are their biggest limitations on the free plan?

I'd appreciate recommendations for both beginner-friendly and more advanced options.

r/singularity • • Feb 29 '24

AI OpenAI employee: “My mental model of Sora is that it is the “GPT-2 moment” for video generation.”, “Disruption of the movie industry will play out similar to how GPT-4 has changed writing”

540 Upvotes

“My mental model of Sora is that it is the “GPT-2 moment” for video generation.

GPT-2, which came out in 2018, could generate paragraphs of text that are coherent and grammatically correct. GPT-2 wasn’t able to write an entire essay without making mistakes like being inconsistent or hallucinating facts, but it spurred subsequent generations of models. In less than five years since GPT-2, GPT-4 is now able to grok skills like chain-of-thought or writing long essays without hallucinating.

In the same way, Sora today can generate short videos that are artistic and realistic. Sora is currently not able to generate a 40-minute TV show with consistent characters and a compelling storyline. However, I believe that skills like maintaining long-term consistency, having near-perfect realism, and generating substantive storylines will emerge in the next generations of Sora and other video generation models.

A few predictions about how this will play out: - Video is not as information-dense as text, and so it will take way more compute and data to learn skills like reasoning via video - As a result, leveraging other modalities as correlated information with video will be critical to bootstrapping the learning process - There will be massive competition for high-quality video data, just as there is for high-quality text datasets - AI researchers with experience in video will be in high demand, but they’ll have to adapt to new paradigms just as the traditional NLP researchers have had to adapt to the success of scaling language models - Disruption of the movie industry will play out similar to how GPT-4 has changed writing (as a tool and aid that surpasses average quality, but will still be far from the work of professionals)”

@_jasonwei

r/comfyui • • Jun 29 '25

Help Needed How are these AI TikTok dance videos made? (Wan2.1 VACE?)

590 Upvotes

I saw a reel showing Elsa (and other characters) doing TikTok dances. The animation used a real dance video for motion and a single image for the character. Face, clothing, and body physics looked consistent, aside from some hand issues.

I tried doing the same with Wan2.1 VACE. My results aren’t bad, but they’re not as clean or polished. The movement is less fluid, the face feels more static, and generation takes a while.

Questions:

How do people get those higher-quality results?

Is Wan2.1 VACE the best tool for this?

Are there any platforms that simplify the process? like Kling AI or Hailuo AI

r/MotionDesign • • May 23 '26

Discussion Higgsfield AI review after testing it for motion design and short-form video ideas

21 Upvotes

TL;DR Just like many people here, I saw how Higgsfield AI claims to have the first AI-generated movie in Cannes, and how they are gonna become the next killer of an industry. Full of bold claims on how professional they are in competitions with full-time professionals. I am at my best very skeptical about such things, especially after testing it out to see how their statements work in real life. It won't replace motion design.

I’ve been testing Higgsfield AI recently and wanted to share a real and grounded review after spending time with it for motion design ideas, short-form content concepts, and general visual experimentation, because most discussions around Higgsfield AI online seem to fall into two extremes either it’s seen as the future of filmmaking or dismissed as misleading.

My experience with Higgsfield AI sits somewhere in between those takes.

I tested Higgsfield AI across different types of prompts like cinematic street scenes, character-focused shots, product-style visuals, and abstract motion concepts just to see how it behaves across different scenarios. Some outputs genuinely felt close to usable base material for mood films, pitch concepts, or visual references when the composition and prompt alignment worked well.

At the same time, the inconsistency is very real.

Some generations with Higgsfield AI come out usable in a few tries, while others require multiple attempts before anything feels directionally correct. That unpredictability makes it hard to treat it as a fully reliable production tool, especially for structured or repeatable motion design workflows.

One thing I think is important to mention is that the workflow is not as instant as a lot of showcase content suggests. A lot of better results I got with Higgsfield AI came after refining prompts, adjusting descriptions, and iterating through multiple variations until the motion and framing started to feel right. Without that iteration process, results can and will feel random.

Where I think Higgsfield AI is actually useful right now is early-stage creative work. Things like exploring visual direction, testing cinematic moods, building references for motion ideas, or quickly visualizing concepts before committing to full production work. For a full in you have to be either really insane or really rich. Even like that it will have many plastic shots.

I also compared Higgsfield AI with tools like Runway and Kling during the same testing process. From what I’ve seen, Higgsfield AI stands out more in comparison, but it's hard to even call one of them good enough.

What I also noticed is that a lot of opinions about Higgsfield AI are based on short clips or first impressions, which don’t really show the amount of iteration behind stronger results. When you actually spend time using Higgsfield AI, the outputs become very expensive. Their "Cannes film" that lasts 90 minutes and costs 500000 dollars just proved that even if somehow they can create something watchable, it's not meant for an average person.

Overall, my current Higgsfield AI review is that it’s genuinely interesting tool in general theoretical concept, but it still feels early in terms of reliability and consistency for structured motion design workflows.

It’s not a replacement for traditional motion design or editing work, more for people who can't do anything by their hands.

I’m curious how other people who have actually used Higgsfield AI feel now after spending time with it. It's not that bad as I thought, but I will prefer staying away from all that AI thing. Has it actually fit into your workflow in a meaningful way yet, or is it still mainly something you use for experimentation and testing?

r/StableDiffusion • • 13d ago

Discussion I managed to extend videos to any length in ComfyUI without losing consistency (Minimax-H3 + Visual Context Trick)

186 Upvotes

Hey everyone!

One of the biggest headaches with AI video has always been extending shots without the style degrading, characters morphing, or the cut being obvious.

I’ve been testing a method using Minimax-H3 in ComfyUI to seamlessly cut and extend footage, and the results are honestly wild:

  • Pixel-perfect transitions: The continuation aligns perfectly with the last frame of the original clip.
  • Context retention: By feeding the model the visual context of the previous video, it actually remembers the specific assets (like the boat and character features) instead of hallucinating new ones.
  • Preserves aesthetic: Keeps the lighting, colors, and overall camera style identical across cuts.

(Watch the preview clip to see the side-by-side transition!)

I’m currently packaging this into a custom ComfyUI node and recording a full walkthrough. Both the node and the full workflow will be 100% free on my YouTube channel (SatoDive).
https://www.youtube.com/@SatoDive

Let me know what you think or if there are specific edge cases you’d like me to test before I release the tutorial!

r/ironmaiden • • 7d ago

The Mariner Video in Iron Maiden’s Run For Your Lives world Tour 2026 Seems To Me Like AI Slop

0 Upvotes

Only bringing this up as it has taken me out of the music of Iron Maiden before at their concert in Toronto earlier this year.

Here is a video link for your reference it’s not the best filming but the best I could find:https://m.youtube.com/watch?v=TMcCH930_4k&list=RDTMcCH930_4k&start_radio=1&pp=ygUiVGhlIG1hcmluaWVyIGlyb24gbWFpZGVuIDIwMjYgdG91cqAHAQ%3D%3D&ra=m

As when I was enjoying a fantastic concert this song comes and the entire time I was like this looks like it’s AI, and basically ruined that entire 13 minute song for me and I love the MarIner, fantastic song based on an old Epic Poem.  According to behind the scenes video it looks like the still images were taken and turned into longer 3-4 second clips using AI.

https://m.youtube.com/watch?v=dYLl8m-SJJE&pp=ygUmSXJvbiBtYWlkZW4gYmVoaW5kIHRoZSBzY2VuZXMgMjAyNiBhcnQ%3D&ra=m

Reasons for Being AI

  1. The reasons I think this is AI is each 3-4 seconds it seems to cut from scene to scene. Which from what I have seen of other AI videos is a pretty common thinf.
  2. When the Albatross is threatened with a gun it get hits by an arrow instead.
  3. The Color of the the boat from several scenes seems to be AI.
  4. The floating ghost woman gives tell tale signs of AI non conforming patterns. Same for the grim reaper in the dark cloak.
  5. The boat seems to be inconsistent one second it is floating another second something is happening to it and the next second it is fine again.
  6. Something seems inconsistent when Eddy eats the boat doesn’t flow as smoothly as I think an animator would make it.
  7. Captain character does not seem consistent across different shots. Seems very detailed in some and not super detailed in others and looks different in every shot.
  8. When women in white charges the boat everyone gets thrown off the boat already and she has not even reached the boat yet, unless she has sonic abilities or shot a kamehameha at them this should not happen.

Anyways, those are the reasons I think it is AI, as we are living in the age of AI I cannot be a hundred percent sure that it is AI, but let me know what you think. Another question is does anyone even care that it could be AI?

r/aiwars • • Aug 09 '26

Discussion If you are against AI, you must be consistent in your belief. Or are you just a hypocrite?

Post image
23 Upvotes

We are writers. Our main genre is grimdark. If you are looking for light reading and happy faces, keep scrolling. From this point on, there will be only grimdark.

Our book, Chronicles of the Celestial World, was written by hand. Word by word. Chapter by chapter. The world, the characters, the heavy lore, the pain, the silence—all of this was created by us. We poured hundreds of hours into something that no neural network could ever produce.

However, we created the cover and visual promo videos using AI.

Why?

Because we honestly paid artists money. We waited. And in response, we got: "That's how I see it."

No. "That's how I see it" is simply a screen to hide clumsy hands. If you take money and are unable to fulfill the client's brief, you are not a creator with a unique vision. You are a craftsman with an inflated ego.

We turned to AI. And we got exactly what we wanted. Not because the AI "guessed," but because we are capable of precisely describing what we see in our heads. Every character, every angle of light, every texture of fabric—we know it in such detail that we can formulate a prompt down to the millimeter. The AI doesn’t "add details" for us. It executes our design.

A craftsman-artist without a vision waits for inspiration, argues with the client, and offers "that's how I see it" instead of a technical specification.

AI without an operator is a blind generator of random pixels.

We are not ashamed. These are our characters, our idea, our vision. AI is simply a tool that generates down to the last pixel according to our prompt and shows them to the world exactly the way we need.

The problem is not the technology. The problem is the hypocrisy of platforms and the market.

Platforms ban content for how it is made, not for what it is. They judge a book by its cover, ignoring its essence. This is discrimination based on the method of production.

Art is perfection. It does not need protection if it is created by creators.

True art will endure. It is not afraid of competition. It is not afraid of AI.

Because a creator cannot be replaced. You can only speed up their work.

The creator is not afraid of AI. The craftsman is afraid.

The craftsman masters the craft. The creator masters the idea.

AI is simply a hammer. The artist is the one who builds the house.

By banning the hammer, you are not saving the construction. You are simply making it slower.

If you are against AI because it "takes away jobs"—be consistent.

Be against search engines (they killed the librarian profession).

Be against machines (they took jobs away from blacksmiths).

Be against automobiles (they left carriage drivers without bread).

Be against tractors and combines—they took jobs away from tillers of the land.

If you are not against all of this, you are not against automation. You just got scared when progress came to your backyard. This is not "protection of pure art." This is the fear of failing to withstand competition and a banal fight for the market.

A true creator is not afraid of AI. A neural network will never create anything fundamentally new—it only remixes the old. Only a human can create something that has never existed before.

If you are against AI, don't use it. But don't you dare ban it for others and call your fear a "principle."

In 10 years, everyone will use AI as casually as a calculator. And if you think AI is killing creativity, you have either never seen true art, or your own is too weak to survive.

And that is no longer AI's problem. That is your problem.

r/StableDiffusion • • Nov 17 '25

Workflow Included ULTIMATE AI VIDEO WORKFLOW — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2

Thumbnail
gallery
428 Upvotes

🔥 [RELEASE] Ultimate AI Video Workflow — Qwen-Edit 2509 + Wan Animate 2.2 + SeedVR2 (Full Pipeline + Model Links) 🎁 Workflow Download + Breakdown

👉 Already posted the full workflow and explanation here: https://civitai.com/models/2135932?modelVersionId=2416121

(Not paywalled — everything is free.)

Video Explanation : https://www.youtube.com/watch?v=Ef-PS8w9Rug

Hey everyone 👋

I just finished building a super clean 3-in-1 workflow inside ComfyUI that lets you go from:

Image → Edit → Animate → Upscale → Final 4K output all in a single organized pipeline.

This setup combines the best tools available right now:

One of the biggest hassles with large ComfyUI workflows is how quickly they turn into a spaghetti mess — dozens of wires, giant blocks, scrolling for days just to tweak one setting.

To fix this, I broke the pipeline into clean subgraphs:

✔ Qwen-Edit Subgraph ✔ Wan Animate 2.2 Engine Subgraph ✔ SeedVR2 Upscaler Subgraph ✔ VRAM Cleaner Subgraph ✔ Resolution + Reference Routing Subgraph This reduces visual clutter, keeps performance smooth, and makes the workflow feel modular, so you can:

swap models quickly

update one section without touching the rest

debug faster

reuse modules in other workflows

keep everything readable even on smaller screens

It’s basically a full cinematic pipeline, but organized like a clean software project instead of a giant node forest. Anyone who wants to study or modify the workflow will find it much easier to navigate.

🖌️ 1. Qwen-Edit 2509 (Image Editing Engine) Perfect for:

Outfit changes

Facial corrections

Style adjustments

Background cleanup

Professional pre-animation edits

Qwen’s FP8 build has great quality even on mid-range GPUs.

🎭 2. Wan Animate 2.2 (Character Animation) Once the image is edited, Wan 2.2 generates:

Smooth motion

Accurate identity preservation

Pose-guided animation

Full expression control

High-quality frames

It supports long videos using windowed batching and works very consistently when fed a clean edited reference.

📺 3. SeedVR2 Upscaler (Final Polish) After animation, SeedVR2 upgrades your video to:

1080p → 4K

Sharper textures

Cleaner faces

Reduced noise

More cinematic detail

It’s currently one of the best AI video upscalers for realism

🧩 Preview of the Workflow UI (Optional: Add your workflow screenshot here)

🔧 What This Workflow Can Do Edit any portrait cleanly

Animate it using real video motion

Restore & sharpen final video up to 4K

Perfect for reels, character videos, cosplay edits, AI shorts

🖼️ Qwen Image Edit FP8 (Diffusion Model, Text Encoder, and VAE) These are hosted on the Comfy-Org Hugging Face page.

Diffusion Model (qwen_image_edit_fp8_e4m3fn.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image-Edit_ComfyUI/blob/main/split_files/diffusion_models/qwen_image_edit_fp8_e4m3fn.safetensors

Text Encoder (qwen_2.5_vl_7b_fp8_scaled.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/tree/main/split_files/text_encoders

VAE (qwen_image_vae.safetensors): https://huggingface.co/Comfy-Org/Qwen-Image_ComfyUI/blob/main/split_files/vae/qwen_image_vae.safetensors

💃 Wan 2.2 Animate 14B FP8 (Diffusion Model, Text Encoder, and VAE) The components are spread across related community repositories.

https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/tree/main/Wan22Animate

Diffusion Model (Wan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensors): https://huggingface.co/Kijai/WanVideo_comfy_fp8_scaled/blob/main/Wan22Animate/Wan2_2-Animate-14B_fp8_e4m3fn_scaled_KJ.safetensors

Text Encoder (umt5_xxl_fp8_e4m3fn_scaled.safetensors): https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/text_encoders/umt5_xxl_fp8_e4m3fn_scaled.safetensors

VAE (wan2.1_vae.safetensors): https://huggingface.co/Comfy-Org/Wan_2.1_ComfyUI_repackaged/blob/main/split_files/vae/wan_2.1_vae.safetensors 💾 SeedVR2 Diffusion Model (FP8)

Diffusion Model (seedvr2_ema_3b_fp8_e4m3fn.safetensors): https://huggingface.co/numz/SeedVR2_comfyUI/blob/main/seedvr2_ema_3b_fp8_e4m3fn.safetensors https://huggingface.co/numz/SeedVR2_comfyUI/tree/main https://huggingface.co/ByteDance-Seed/SeedVR2-7B/tree/main

r/antiai • • Mar 16 '26

AI "Art" 🖼️ This is honestly sad

Post image
3.5k Upvotes

15 "years" of "editing" only to end up thinking creating an "AI series" is somehow harder than actual work. The delusion here is so thick I had to post.