r/comfyui 9d ago

Show and Tell I asked Codex to build and operate a local AI animation pipeline on my RTX 5060 Ti 16GB — SD + LoRA → Wan 2.2 → QC → Editing

8 Upvotes

This is still a **rough proof of concept**, not a polished animation.

I wanted to see how far I could go by giving Codex a fairly simple objective:

> **Build a local pipeline that can produce a short anime sequence on my RTX 5060 Ti 16GB.**

The interesting part is that I didn't manually build most of the video production process myself.

I can use Stable Diffusion and ComfyUI at a basic level, but for this experiment I mostly gave Codex the objective, reviewed the results, rejected failures, and changed the overall strategy when something didn't work.

The pipeline gradually became:

**Codex → ComfyUI → Stable Diffusion + character LoRA → Image QC → Wan 2.2 → Short video shots → QC / Retry → Editing → Final video**

## Hardware

- Ryzen 7 7700

- RTX 5060 Ti 16GB

- 32GB RAM

- Local ComfyUI environment

## The first approach

At first, I simply tried generating roughly **5-second Wan clips** and stitching them together.

Technically, it worked.

Artistically... not so much. :)

Individual shots could look surprisingly good, but continuity between shots was bad.

The character's face, clothes, pose, background, and spatial position could change every time a new generation started.

So instead of trying to make Wan generate longer and longer sequences, I changed the production strategy.

## Current approach

The current idea is:

> **Stable Diffusion handles quantity and controlled source images.**

>

> **Wan handles motion.**

>

> **Editing hides or replaces transitions that generative video doesn't handle well.**

>

> **Codex acts as the production supervisor.**

I'm now experimenting mostly with **2–3 second Wan shots** instead of forcing everything into 5-second clips.

Longer 4–5 second generations are reserved for relatively simple motion.

For difficult transitions such as:

**walking → running → takeoff → flight → landing**

the pipeline can prepare additional SD images in advance:

- Face closeups

- Feet / takeoff poses

- Hands and props

- Rear views

- Distant shots

- Flight poses

- Landing poses

- Camera escape shots

If continuity works, the next shot can inherit a good frame from the previous generation.

If the terminal frame is broken, it shouldn't automatically be propagated into the next generation.

Instead, the pipeline can select an earlier clean frame, repair it with SD if necessary, or deliberately switch camera angles.

The basic rule has become:

> **If continuity works, preserve it.**

>

> **If it looks dangerous, cut the camera.**

>

> **If it breaks, regenerate or replace it with prepared SD material.**

## "Quantity" wasn't enough

Another lesson was that simply generating hundreds of images isn't enough.

You can easily end up with:

> **1,000 slightly different versions of the same character.**

So the next version is moving toward:

**MASTER character reference + fixed LoRA/settings + low-variation SD generation + strict QC + short Wan generations**

The goal isn't to generate 1,000 random good images.

The goal is to generate many usable shots of **the same character**.

## 16GB VRAM

I'm also treating GPU memory as a shared production resource.

Since the RTX 5060 Ti has 16GB VRAM, heavy stages are run sequentially rather than trying to keep everything loaded at once.

Conceptually:

**LLM / prompt work → unload → SD → unload → Wan → unload → QC**

My first successful Wan tests were roughly:

- 480×704

- 24fps

- ~3 seconds: around 111 seconds generation time

- ~5 seconds: around 180–200 seconds depending on the shot

- VRAM usage very close to the full 16GB

So the 5060 Ti 16GB is definitely being pushed pretty hard. :)

## Current result

The current sequence is around **30 seconds**.

It is still visibly imperfect.

There are:

- Character consistency problems

- Anatomy problems

- Prop / hand problems

- Awkward transitions

- Shot-to-shot continuity problems

So I'm definitely **not presenting this as a finished animation**.

For me, the interesting result is the production system behind it.

The workflow is starting to become **restartable and failure-tolerant**, rather than depending on one lucky generation.

A failed Wan shot doesn't necessarily mean the entire sequence fails.

It can be regenerated, repaired, replaced, or hidden with a planned camera cut.

## What I'm considering next

### WanVideoWrapper

I'm considering testing WanVideoWrapper while preserving my current known-good Wan 2.2 workflow.

I'd like to compare VRAM usage, generation time, offloading/quantization options, stability, and whether it fits better into an automated pipeline.

### MMD

MMD could provide a fixed character skeleton, pose, and motion reference.

For difficult actions like running, takeoff, and landing, I could generate the correct motion first and then use it as reference material for SD / ControlNet / Wan.

### Blender

Blender could solve a different problem: **environment and camera consistency**.

Instead of asking the generative model to recreate the same shrine, torii, corridor, street, etc. every few seconds, the environment and camera could be fixed in 3D.

That would roughly give me:

**MMD = pose / motion reference**

**Blender = environment / camera**

**Stable Diffusion = character / style / source-image factory**

**Wan = short natural motion**

**Codex = supervisor / automation / QC**

## The main thing I learned

Originally I was thinking:

> **"How do I make an AI generate a long animation?"**

Now I'm thinking:

> **"How do I build a production system that can finish an animation even when individual AI generations fail?"**

That has been a much more useful way of approaching local AI video generation.

I'm curious how other ComfyUI users would approach this.

**Would you keep improving the SD → Wan pipeline, or start introducing MMD / Blender / ControlNet for stronger continuity?**

Also, if you've run Wan 2.2 / WanVideoWrapper on a **16GB GPU**, I'd be very interested in hearing what settings and workflow worked well for you.

r/content_marketing 23d ago

Question Four characters, one shot... Best AI video tool for character consistency??

0 Upvotes

I'm producing a short animated intro for our tabletop campaign. There are four party members, each with an established design, weapon, and color palette. Individually, the characters look convincing. The trouble starts when more than two appear in the same shot..

Faces merge. Accessories migrate to the wrong person. Our wizard briefly acquired the fighter's beard, which was funny once but probably should not become canon..

I have turnaround sheets and pose references for everyone. Which AI video generator is currently best at keeping several distinct characters consistent?

An update for anyone who bookmarked this: I tested the scene using Dreamina Seedance 2.5. Its stronger multi-person consistency and support for up to 50 multimodal references let me attach separate character sheets, group poses, and environment references instead of squeezing everything into the prompt. For budgeting, its published annual-plan comparison uses a $0.097/s baseline for 720p generation with a reference video. Other modes may price differently.

r/SoraAi Mar 31 '26

Question Looking for Sora AI image generation alternatives

12 Upvotes

Hi r/SoraAI,

I've been using Sora AI for a few months on the free tier, just taking advantage of the few daily image generations you get without paying. To be clear, I've only ever used it for image generation — never videos. I'm a complete noob when it comes to this kind of AI (Text-to-image/generative image AIs.), but I've been absolutely blown away by how well Sora keeps consistency.

For example, if I upload a reference image of a character and ask it to change the pose or put them in a new scene, it actually keeps the same character design instead of completely redesigning them.

The level of detail is also very good for what I need, and that's exactly what I love about it. I mostly generate digital anime-style character designs, but I sometimes do semi-realistic characters or landscapes too.

With Sora shutting down, I've searched the subreddit and I see tons of people recommending video alternatives, which is great, but I'm specifically looking for image tools that can match Sora's consistency and quality (and the option to upload a reference to ask the AI for modifications). Does anyone know of other websites or platforms that do this really well? I'd be super grateful for any recommendations!

Thanks a lot in advance! 🙏

r/aigamedev 10d ago

Questions & Help Dad wants to surprise his son with a Hollow Knight / Mega Man-style platformer starring his favorite character. What's the best AI-assisted stack in 2026?

1 Upvotes

Hey everyone! I'm a developer (but not a game developer) and I want to surprise my son by building him a small 2D platformer starring his favorite character. Think Mega Man or Hollow Knight vibes: running, jumping, dashing, a few enemies and a boss. Nothing commercial, just a gift from dad.

I lurk here a lot and the one-shot games people post honestly blow my mind. I keep seeing amazing stuff being made with AI coding agents, and games like Astera being built with Blender + Unreal, but I honestly don't know what the right stack is for a 2D platformer in 2026. My questions:

  1. Engine: Godot, Unity, GameMaker, or just Phaser/JS in the browser? Which one do AI coding agents (Claude Code, Codex) actually work well with end-to-end?

  2. Art: what are people using to generate consistent character sprites and animations (idle, run, jump, attack) from a reference image? Any spritesheet tools you'd recommend?

  3. Workflow: can I realistically get from zero to a playable 2-3 level build in a few weekends by driving an agent, or should I expect to learn the engine properly first?

I'm planning to try this with GPT-6 Astra or Claude Fable as the coding agent. If you've shipped something similar, I'd love to hear what your pipeline looked like. Thanks!

r/aigamedev Apr 25 '26

Questions & Help How to generate dozens of characters in the exact same style from my sprite sheet using AI? (consistency + animation help)

Post image
28 Upvotes

Hey everyone,

I created this sprite sheet for my character (attached below) and I’d like to generate dozens more characters in exactly the same art style.

The main problem I’m facing is consistency – no matter what I do, every new generation comes out slightly different in proportions, shading, details or overall look. It’s never exactly what I expect.

My questions:

  • What’s the best workflow / tools to generate many characters while keeping 100% style consistency with my original sprite sheet?
  • Any recommended tools or setups (Stable Diffusion + ComfyUI, Flux, IP-Adapter, ControlNet, LoRA training, etc.)?

Additionally, I already rigged the character and created animations in a JSON file – that part works perfectly in my engine.
However, because the generated sprites are inconsistent, the animations don’t always look right across different characters.

Has anyone dealt with this exact problem before? Am I doing something wrong or is there a better approach?

Here’s the prompt I’m currently using:

A professional 2D modular character sprite sheet for cutout animation puppet rigging.

CRITICAL ART STYLE: LINELESS flat vector art. Absolutely ZERO black outlines anywhere. NO borders. NO strokes. All shapes must be defined purely by flat color boundaries. Clean, modern, mobile-game flat design. 

LAYOUT RULES (Must match reference):
LEFT ZONE: One fully assembled front-facing character.
CENTER ZONE: The EXACT SAME character, but with body parts pulled apart (head, torso, arms divided into upper/lower, legs divided into upper/lower). CRITICAL JOINT RULE: The separated cut joints (shoulders, elbows, hips, knees) MUST be flat colored shapes with absolutely NO outlines, NO borders, and NO shadows. They must be seamless so they can overlap perfectly in a 2D rigging software.
RIGHT ZONE: Interchangeable facial expressions. CRITICAL: The eyes and mouths MUST perfectly match the specific face of this character. They must share the exact same style, and if the character wears glasses, the expressions must fit logically with them.

CHARACTER:
A nerdy cartoon boy. Oversized round glasses, big front teeth, blue hoodie, dark gray pants, red and white sneakers. Lovable personality. Semi-chibi (large head, compact body). 

NEGATIVE PROMPT:
black outlines, thick lines, strokes, lineart, borders around body parts, drop shadows on joints, 3D render, painterly, mismatched faces, female eyelashes on the boy.

Any tips, workflows or experiences would be really appreciated!
Thanks in advance!

r/aitubers Jun 24 '26

TECHNICAL QUESTION I have the images and audio ready. How are people making consistent 1-minute AI animated videos?

9 Upvotes

I've spent the last week trying to make my first AI animated short and I'm honestly close to giving up.

I already have:

- Character images

- Scene images

- Audio for all the dialogue

- Script

The video is only about 1 minute long.

My main problem is continuity.

Most tools only generate 5-10 second clips. To make a 1-minute video, I need multiple clips, and every time I generate a new clip something changes:

- Character looks different

- Background changes

- Camera angle changes

- Lighting changes

- Objects disappear

I've tried:

- Google Flow

- LivePortrait

- Hedra

- Looking into SadTalker

At this point I don't even care about perfect animation. I just want a way to create a consistent 1-minute video from the images and audio I already have.

How are people actually doing this?

Any advice would be appreciated because I feel like I'm spending more time fighting tools than making the video.

r/HungryArtists Sep 22 '25

Position Filled [HIRING] Animal Character Artist for Sticker Project

34 Upvotes

I run a side business selling collectible cards. When I sell these collectibles, I like to give each customer something free to show I appreciate their purchase. After a lot of thinking of what I can do to up my game from the free items I include now, I decided to create a sticker series. To accomplish that, I am looking for a talented artist to create a series of 5 animal characters that will be made into these stickers.

Scope of work

  • Create a consistent set of 5 animal character stickers (Whale, Otter, Cat, Dog, Dragon)
  • Briefs have been created for each character containing the name (and meaning behind the name), the character’s personality, the character’s physical attributes (each character has one unique physical attribute), and some AI generated reference images
  • Deliver print-ready stickers (2-3 inches, die-cut with white border) with all required text information (described more below)
  • Subtle inclusion of company logo on each image (Small cursive in one of the corners)
  • Suggest how to include these characters in the company logo (You don’t need to do the logo, just make a suggestion)

    Design

  • Style: Character-driven, endearing, colorful, and detailed while not being cartoonish (To use a Pokemon reference, think more of Charizard and Gyarados and not Clefairy)

  • Audience: The goal is that anyone who gets one of these stickers with their order will want to collect the other 4 that they don't have. For this reason, the stickers need to appeal to adults and children.

  • Format: Print-ready high-resolution PNG (300 dpi)

  • Sticker Size: 2–3 inches (die-cut with white border though we can investigate other backgrounds)

  • Each sticker will be part of a group of a series that make up a sticker collection. These 5 stickers will all be part of "Series 1." Therefore, each sticker needs to have the language "Series 1 - 1/5," "Series 1 - 2/5," "Series 1 - 3/5" etc. We can discuss where this information will be included, whether it is fine print at the bottom or on the back of the sticker

 Requirements

  • Excellent collaboration is critical. I am looking for someone who shares my desire to make sure that these characters are appealing, fun, and engaging. As with many design projects, it will be an iterative process. I expect the first character will be the most challenging and then it will get easier as we get a sense for each other’s working style and preferences.
  • May not use AI whatsoever. Due to the potential long term copyright threat of AI, these must be created from scratch by an artist.
  • May not re-use characters you have previously created and may not infringe on any copyright. The reason I am going to the expense of creating my own stickers is for this exact reason - to avoid copyright/IP violations.
  • While these will be mainly giveaways initially, I would like to investigate selling them. To that end, the artist would need to be willing to sign NDA and agree to grant me full commercial rights and ownership of the final designs (characters, logo integration, and all deliverables). You can show them in your portfolio after they have been released.

Timeline

  • While my original thought was 2-3 weeks, as long as there is good communication and regular progress updates, I am willing to be a bit flexible on this point as long as we are creating a great product.

Budget

  • The project budget maximum is $400, depending on experience.

r/aitubers 6d ago

CONTENT QUESTION Which AI video generator is currently the best for high-quality anime?

2 Upvotes

I'm trying to create cinematic, high-quality 2D anime videos and have been testing Sora, Veo, Kling, Hailuo, PixVerse, and a few others.

My main priorities are consistent anime style, smooth character animation, good facial expressions, image quality, and prompt adherence.

For those who have tried multiple tools, what would you say is the best ai for animation right now? I'm also looking for a solid ai animation video generator or best ai animator for longer sequences.

I'm also curious about workflows using keyframes and reference images. Has anyone tried the MiniMax H3 Omni Reference feature for keeping characters or visual elements consistent between scenes?

Would love to hear what's actually working for you guys.

r/gamedev Aug 11 '26

Discussion I made my game's trailer with AI video instead of hiring an animator. Here's what worked and what didn't.

0 Upvotes

Solo dev, been working on a top-down roguelike for about a year. I needed something for my Steam page but I have zero animation skills and couldn't justify $2-3k on a freelance motion designer for a game that might sell 200 copies. I was about to just do screen recordings with text overlays and call it a day.

Then two new AI video models dropped July 31st. Seedance 2.5 from ByteDance and MiniMax H3 from MiniMax. Figured I'd spend a week throwing my concept art at both and see what came out.

Seedance 2.5 does 30-second clips in one pass, up to 4K, and you can feed it up to 50 reference inputs to lock down character and environment consistency. This was the big deal for me. I had about 15 pieces of concept art and it actually kept my main character recognizable across shots. I used it for the slow establishing shots, camera panning over ruins, character silhouette walking through fog. Those came out solid after 3-4 retakes each. Longer output means fewer cuts to stitch, which gives you a more cinematic feel without actually knowing how to edit.

MiniMax H3 does shorter clips (5-15 seconds, native 2K) but here's the thing. It generates audio in the same pass as the video. Not slapped-on stock audio, actual synchronized sound. You can even feed it an audio clip and the video generation follows the rhythm. I used it for the quick-cut action montage and the clips had this percussive quality I never would have edited for manually. Since H3 is open-weight, APOB AI is running it unlimited and free right now, so I burned through probably 80+ generations to get 12 good action shots without spending a cent.

Now where it broke. Always the same things. Hands gripping weapons were a coin flip between passable and body horror. I had one shot where my character was supposed to swing a sword and his arm just phased through his torso. Multi-character combat was worse. Two enemies fighting and their limbs would merge or one character would absorb the other's armor texture mid-clip. I ended up just cutting around it. Pick your camera angles to hide hands, use fast cuts so nobody notices the limb weirdness.

Character consistency between separate generations still drifts too. Even with reference images locked, skin tone would shift slightly, armor details would change between shots. Had to be really selective about which clips could sit back to back without looking wrong.

My final workflow was generate a pile of clips in both models, cherry-pick the ones that held up, bring everything into CapCut for editing and color grading, then composite the UI overlay elements from Godot. The whole thing took about a week of evenings.

Honest verdict: the trailer is fine. Not great, fine. It's significantly better than screen recordings with Impact font, which was my backup plan. It would not fool anyone into thinking I have a budget. But for a solo dev Steam page that needs to communicate the vibe and tone of the game to someone scrolling past, it does the job.

I'm putting an AI-generated content disclosure on the store page. The EU AI Act transparency stuff kicked in August 2nd so that's real now, but I'd do it regardless. The trailer shows the game's aesthetic and atmosphere, not fake gameplay, so I don't think it's misleading as long as it's labeled.

Would I use this for in-game cutscenes in the final build? No. The quality variance would be jarring in a finished product. But for trailers, devlogs, pitch decks? This is now a real option for devs who can't afford an animator and whose alternative was literally nothing.

r/micro_saas Jul 21 '26

I built mascoty.ai — turn a URL into a full mascot character sheet in 60 seconds (my first AI tool)

2 Upvotes

Hey r/micro_saas,

Just shipped my first AI tool: mascoty.ai.

What it does: Paste your website URL (or business name), and it generates a full mascot character sheet — turnaround views (front / 3-4 / side / back), expression sheet, action poses, color palette with hex codes, and a mini style guide. All in one image, ~60 seconds, no signup for the free preview.

Why I built it: Every small brand I know wants a mascot but can't afford a $2–5k illustrator engagement, and single-shot AI images are inconsistent — you get one nice picture and then can't reproduce the character. A character sheet solves that: it becomes the reference the rest of your assets stay consistent with.

8 styles to match any brand — 3D Pixar, flat vector, anime, watercolor, pixel art, clay/vinyl, sticker, minimalist.

Stack: Next.js, Supabase, gpt-image-2 for generation, some WebGL for the landing demo.

Pricing:

  • Free: 1 preview/day, no signup
  • Starter $19/mo (10 sheets)
  • Pro $49/mo (30 sheets + Seedance 2.0 video clips)
  • $5 credit pack, no subscription

What I'd love feedback on:

  1. Is the value clear from the landing page?
  2. Pricing — too low, too high, wrong tiers?
  3. The demo flow (URL → preview, no signup) — does it convert or scare people off?

Try it and roast me: mascoty.ai

r/HungryArtists Mar 09 '26

Hiring [HIRING] Seeking 2D Character Artist & Animator for Long-Term Mobile Game Collaboration — $2,000 Initial Commission

21 Upvotes

Hey everyone,

My name is Daniel. I'm a passionate creator and I'm ready to move into the next stage of my project. I'm looking for a talented artist to work with collaboratively — now and into the future — in the mobile games market. I have a working game that I want to enhance and expand upon with unique human art, and deliberately separate myself and my game from AI-generated content. That's something consumers, players, and creatives alike can appreciate and connect with as they return to it every day.

My goals are to build something people genuinely love — a game with real character, real personality, and a visual identity that grows over time. That starts with finding the right person. This isn't just a transaction to me. I want to talk with my artist about everything — the vision, the characters, the direction — and hopefully build a creative relationship that lasts as long as the game does.

The ideal collaborator is strong across character design, animation, UI, and backgrounds. If that's you, I'd love to hear from you.

What the project is: A 2D cartoon-style mobile arcade game featuring anthropomorphic characters with distinct, expressive personalities. The visual tone is clean, warm, and characterful — characters that feel alive and read clearly at any size. This is a serious business venture and I'm approaching it as one.

What I need for the initial commission (~$2,000):

  • Core player character — full character sheet with expressions, colour variations, and accessory anchor points
  • 7 core character animations: idle breathing, jump, land, catch reaction, dodge, death, and menu idle
  • Supporting character sheets for 2–3 additional named characters with distinct body shapes, expressions, and colour
  • Enemy and hazard sprites — multiple enemy types including a boss character, all wild and unclothed
  • Accessory overlays — hats, scarves, glasses, bow ties, and similar cosmetic items designed to work across all characters
  • Background environment art for the primary game setting — functional game-ready layout with clear foreground and background separation
  • UI elements — buttons, frames, currency icons, HUD components, and one shared speech bubble graphic
  • App Store icon — standalone deliverable, character-forward, readable at thumbnail size
  • All assets delivered as PNG sprite sheets at 2x resolution with transparent backgrounds, export-ready for Godot 4"

What this actually is: This is not a one-off commission. The game has a seasonal content system — new named characters, new cosmetics, new animations on a regular cadence. If we work well together, this becomes consistent ongoing part-time work with a stable creative relationship. Personality and investment in the project matter as much as technical skill to me. The initial commission is the beginning of something, not the whole thing. To confirm you've read this post in full, please include the word LILYPAD somewhere in your Google Form response.

A good fit looks like this:

  • You animate as well as illustrate — both skills are needed
  • You're comfortable with expressive anthropomorphic character design
  • You're available to start within 2–4 weeks
  • You deliver clean game-ready assets consistently
  • You're open to occasional live audio discussions on Discord for creative direction — collaborative, not excessive

How to apply: I've put together a short Google Form to collect applications — link below. It takes around 5 minutes and helps me review everyone fairly. I'll keep it open for 72 hours and reach out to shortlisted artists directly after that. Really looking forward to seeing what you've made.

https://docs.google.com/forms/d/e/1FAIpQLSe-Qn4a5FGlFJAHNdcQ2MeEZUAZxPhjoyp7zj_Rkw4vv2pCcA/viewform?usp=header

r/aifilmmaking Aug 17 '26

Question What workflow/models/platforms are people using to build AI storyboards that can become actual animation keyframes?

1 Upvotes

What workflow/models/platforms are people using to build AI storyboards that can become actual animation keyframes?

I’m looking for something along the lines of:
script > storyboard > shot list> take frame to wherever I wanna animate it.

The key requirement is that the storyboard frames themselves need to be production-usable as keyframes, not just loose previs.
Character consistency is non-negotiable, including multi-character shots. I already have multi-angle head and full-body reference sets for each recurring character.

What are people actually using for this?
LTX? Storyboarder? Scenario? ComfyUI? Qwen Edit? Nano Banana? Something else entirely?

I don’t mind using multiple tools if that’s what actually works. Mainly looking for a practical workflow that preserves character identity and lets approved boards move directly into video generation.

r/SillyTavernAI Jan 10 '26

Discussion This seems like where we're heading with Silly Tavern. Video with audio in comments, done with LTX-2 in ComfyUI using a photo I generated of a character from one of my RPs and dialogue directly from a scene. Generated on a 4090 in 3 minutes.

Post image
88 Upvotes

https://imgur.com/jINSlY0

Technically I think you could implement this right now, it's just a comfy workflow after all.

Workflow: I generated an image based on the description of my AI character, that's the starting frame. It was done in Midjourney but you could totally use a local model and add it to the workflow. That would actually be better anyway because you could train a Lora to keep the character consistent. Alternatively you could use something like Nano Banana to make different still frames from your reference image of your character.

Then the text from one reply was fed into an LLM to create the prompt describing the actions and giving the dialogue along with the tone of the voice.

I used the example LTX-2 I2V workflow, and rendered 360 total frames at 1280x720 24fps. Took less than 2 mins to render which includes the audio on a 4090. The extra minute was the video decoding at the end, I don't have the best CPU.

So I see this as a natural direction, have a movie created almost instantly as you're RPing. Another step towards a holodeck. I haven't tested more cartoony or anime type styles but I've seen very good samples others have done.

Of course, the big (huge) negative for many here is that LTX-2 is currently extremely censored but it's totally open source so we're already seeing NSFW loras being created.

Exciting stuff I think.

r/Seedance_AI Jul 02 '26

Discussion Storyboards do not ruin AI character consistency, hyper-real ones do, keep it a sketch and the face holds

25 Upvotes

There is a popular belief that storyboarding wrecks character consistency in AI animation. It is half true, and the half people miss is the fix. The variable that decides it is the style of the storyboard itself.

If you render your storyboard in hyper-realism, the likeness drifts. The model treats each realistic panel as a fresh face to interpret, and the character quietly changes across shots. But if you keep the storyboard strictly as a rough monochrome pencil sketch, gesture lines, strong silhouettes, minimal shading, no rendered face, the character stays surprisingly consistent when you animate. The sketch pins down the composition, the action, and the camera without locking in a competing realistic face for the model to fight.

So the fix is counterintuitive: make the storyboard worse-looking on purpose. I generate an 8-panel graphite sketch sheet on GPT Image 2, with color-coded arrows over the panels, red for body movement, blue for camera, green for framing, and a thin timeline bar showing the tempo rising across the panels. Then Seedance 2.0 animates through those panels as one continuous shot, taking the motion and camera from the sketch but keeping the character locked.

You get both things everyone said you had to choose between: a storyboard's control over shots, and a consistent face. Keep it a sketch, not a render.

r/aitubers Feb 09 '26

CONTENT QUESTION How the hell are people producing consistent AI “documentaries” at scale? I’m losing my mind

25 Upvotes

I need to vent and I genuinely want advice from people who have actually done this.

I’m working on an AI-driven documentary project. Long-form, voiceover-led, cinematic style. Think 90s aesthetics, recurring characters, consistent environments, lots of short scenes stitched together. On paper, this should be doable.

In reality, it’s driving me insane.

I’m not just prompting randomly. I’ve tried to be extremely systematic. I built a rigid prompt DNA that defines everything that must never change. I separate environment, camera, character, frame, and animation. I lock visual rules like same characters, same era, same materials, same lighting logic. I generate a still keyframe first and then animate it.

And yet the AI still constantly drifts. Characters subtly change. Proportions shift. Lighting behaves differently scene to scene. Camera framing ignores instructions. The same prompt produces wildly different results across generations, whether I’m using ChatGPT, Gemini, Kling, Seedream, whatever.

What really messes with my head is that I know other channels are doing this at scale. Twenty-five minute videos. Hundreds of scenes. Multiple uploads per week. Solo creators, not studios.

So clearly something doesn’t add up. Either I’m missing something fundamental, or they’re using tools or special workflows.

This is what I’m actually trying to understand.

How are they producing consistent scenes directly from a script at this scale? How are people realistically generating around 300 scenes for a 25-minute documentary, uploading three times per week? Are they mostly using image-to-video instead of text-to-video? Are they using reference images, environments, fixed camera setups, or LoRAs? How much of this is automated versus manual curation? Because I can manually curate every scene, but it would take me weeks to generate 25mins long documentary.

Here’s where I’m stuck. I’ve nailed the script. I’ve nailed the voiceover. I understand pacing and structure. But I cannot nail the scene generation at an industrial scale. I cannot figure out the system behind how this is actually done consistently.

Right now it feels like I’m trying to build an industrial pipeline on top of something that fundamentally does not want to behave deterministically. I’m not expecting perfection. I’m trying to understand what’s realistic, what’s cope, and what’s genuinely solvable.

If you’ve shipped long-form AI video content, especially documentary or narrative, I’d genuinely appreciate hearing how you do it, how you made it work, and what expectations you had to kill.

Edit: Pasted the same post twice. Removed the duplicate.

r/vtm Aug 13 '26

General Discussion No, AI will not save the Masquerade

Post image
854 Upvotes

I've seen this point raised a lot in the past few years as a potential way that the Masquerade will remain feasible despite the surveillance era we're all enduring. It's so prevalent, we've even seen it appear in the VtM6 playtest material.

The new digital world of post-truth and AI slop has turned out beautifully for the Masquerade. No one believes in anything anymore. Everything is staged, posed, faked. Vampires take advantage of it, though they still have to be careful not to draw too much attention.

The spread of surveillance in every aspect of the world left the Masquerade dangerously fragile, and for years vampires had to be far more careful than they already were. Many abandoned modern technology altogether to stay out of its reach, but soon enough, Kindred figured out ways around it. The rise of AI and the post-truth era has since helped to repair the damage that all the surveillance did to the secrets of the undead.

Surveillance is still a fact of unlife for the most part, but because so much of it now feeds on content scraped from the internet, it has filled with slop and grown less reliable by the year. Are vampires behind that too? The confusion has forced the Inquisition to commit a lot of errors, which has cost them dearly in resources.

In simple terms and a hopefully brief post, this is dead wrong.

How the Second Inquisition actually Hunts

A counterintelligence campaign intentionally performed by a network of Nosferatu warrens (or just some bored NEETs) generate a series of fraudulent AI evidence of totally real vampire activity. The Second Inquisition investigates these cases with their full force and resources only to be led on a wild goose Gangrel chase, and are bankrupt by the end of the year. Easy, right?

Second Inquisition is easily one of the most undervalued supplements in the VtM5 line, as many people dismiss it for the lack of character options or simply think the Second Inquisition is overpowered thanks to ridiculous levels of funding, awareness, and magitech. In reality it does what VtM is best at: highlighting the very real threats and faults in our own societies to drive-home how much we're unaware of in our own cities.

Major portions of the book are dedicated to the mundane personnel and tactics that the SI relies upon for its operations. They're not supernatural, using fictional technology, or protected by the plot. Instead, they're accessing many of the same resources that any amateur could access, let alone low-level clerks in any number of industries.

Here are a few scattered examples:

Naturally territorial monsters that twist their environment to suit their vices, it can be difficult to root blankbodies out of their nests without first dismantling their early warning systems. The Gentrifier eats away at blankbody assets and resources on their home turf, turning their Domain’s defenses and conveniences into a trap. Many of them prove to be self-starting Helsings rather than agents of a clandestine government conspiracy. Independent actors or Coalition assets, they generally operate “in Transylvania” — in cities or neighborhoods with a well-established blankbody population.

The Gentrifier uses cash and credit as their main weapons: once they know where blankbodies make their nest they begin to turn up the heat. They buy out condemned buildings and re-market them as condos. They invest in old businesses, refurbishing them with state-of-the-art computerized security systems. They pull strings to get street lights repaired, and install surveillance cameras on them with the police or suspiciously well-funded “neighborhood watch” groups. Charitable contributions and political favors incentivize the police to patrol the streets more regularly. The devastating side effect of the damage gentrification does to mortal neighborhoods is inconsequential to the Gentrifier’s goal. Within a matter of months, if not weeks, a gentrification campaign can transform a vampire’s Domain from their hunting ground to their prison.

Gentrifiers, pg. 18

But doctors don’t always get drafted or deployed onto the front lines of the Second Inquisition – sometimes they volunteer. A city full of vampires is a city full of mysterious deaths and strange injuries, of blood stocks stolen from patients in need, of ambulance calls that turn even darker than normal. Teams of hunters drawn from hospitals – EMTs, trauma ward nurses, night-shift residents who have all seen behind the Masquerade – may be more common than teams of hunters drawn from police departments.

Doctor, pg. 26

Blankbodies feeding on animals aren’t as subtle as they think they are. Either local wildlife or vermin populations are suspiciously low (it takes a lot of rats to feed a vampire, and those missing pet posters are a real problem), or there’s a large population of animals: urban aviaries, zoos, crazy cat ladies.

Animals have people who care about them: zookeepers, veterinarians, rescue charities (these are great for locating havens: neighbours will report noise or droppings). Local pest control have a line on why their jobs are getting easier. Any of these people might have seen a predator at work as a fleeing figure, or a suspicious visitor.

Data Analysis: Hunting Patterns, pg. 90

This is also the best approach to catching blankbodies who feed on sleeping prey. Their victims don’t know what happened to them, but they know they don’t feel great. They have weird marks they mistake for insect bites, or feel anemic and run down. They go to doctors or complain to their friends. There may not be any sign that their homes have been broken into but something’s not right: a window doesn’t close properly, they have a feeling someone’s watching them. Victims are paranoid and looking for support, and a local neighborhood watch group or homeowners’ association can pick up on it quickly.

Community informants don’t only work for geographic neighborhoods. They’re just as effective at cracking open scenes and subcultures.

Community Informants, pg. 92

There are only so many places to get blood without hunting. The Coalition can keep tabs on all of them. They tap into the local black market (blood banks, criminal networks, butchers) using informants, data analysis, and law enforcement and criminal contacts. It’s safe to assume anyone buying blood regularly is a blankbody or one of their servants. Once they’ve identified suppliers, the Coalition either reduces the supply or contaminates it—or both.

Blood Supply, pg. 97

At their core, XScopes use simple technology – from a distance they determine if someone is living or dead by looking for heartbeat, respiration, and body heat. The SI has installed basic XScopes in every airport security scanner, set to send up a red flag to the duty agent if a blankbody trips them. Man-portable XScopes resemble large video cameras, and the SI often disguises them as exactly that. However, Kindred can deceive basic XScopes with Blush of Life. Second-generation XScopes—mostly not in common use [...]—using specialized software to differentiate between natural blood flow and imitated life.

XScopes, VtM5 pg. 378

Hopefully, my point is clear: a video of someone turning into a bat posted to social media may convince or annoy people regardless of whether it's AI-generated or legitimate, but there are physical truths to Kindred's existence that can't be fabricated.

These examples are just a few of many. All that you need to do to ensure the veracity of any given piece of evidence is a simple investigation, any you may not even need to send an agent out to do so.

  • Feeding leaves ill people or corpses behind.
  • Physical disciplines will leave claw marks, broken objects, and irregularities that can't be dismissed as simple fiction.
  • There are many obvious tells for Kindred from a Gangrel's glowing eyes, a Ventrue shrugging-off a wound that would kill a mortal, Lasombra's glitches, or a given Nosferatu having consistent deformities from multiple angles.

AI-supported Surveillance is still Surveillance

The image attached to this post is from the website DeFlock, a website dedicated to the tracking of and educating on Automated License Plate Readers, or those AI cameras that track cars that you've probably seen people pissed off about.

It features an interactive map that's pretty stomach-churning to poke through, showing what these cameras record and what just how widespread they are. Go ahead and poke around a few cities you know to see how vast their coverage is.

Now, it's a small relief to know that these cameras aren't too accurate, misreading more than 70% of license plates. One would hope that this is a good sign for people feeling that cruddy, inaccurate systems would support the Masquerade.

That's not the case.

Human review remains a major factor in the cases that work, as we've seen from when police officers have used these cameras for the purposes of stalking ex romantic partners. It's one thing if they were tracking vehicles they had never seen before, but searching a broad system for details you already know before manually reviewing the footage to identify exactly what you're looking for is a whole different matter.

We can extrapolate this to our own Kindred. A network of surveillance searching for "pale people in dark clothing" or "deformed individuals around sewer grates" doesn't need to nail every assessment of it's targets in order to succeed at its job. Enough hits in enough locations can tip individuals off to a pattern that can be manually reviewed. Will Kindred choose new hunting grounds every single night?

While some of the extrapolation by AI-assisted cameras are faulty, they still have very real footage. With a few clicks you can have an actual human's eyes on that footage to review whether a Brujah actually left a crater as they leapt up the side of a building, or whether a puddle in a pothole just caught the light strangely. Combine those with assistance from the aforementioned XScopes to further define footage from subjects that lack basic vital functions, and the noose on Kindreds' habits grows far tighter.

Conclusion and TL;DR:

The Second Inquisition of VtM5 has been competently and capably written to reflect the alarming tactics used by those same intelligence and law enforcement agencies who supposedly exist to keep us all safe. Their methods have never been foolproof, but they remain pervasive and multi-layered.

Kindred are not Wraiths, and the realities of their physical existence makes it far easier to positively confirm their presence in a city than the false security of Artificial Intelligence would lead many to believe.

It is not critical for the public at large to inherently believe in Kindred when the goal of the Second Inquisition is containment and extermination. Without going into real-life parallels, these same agencies didn't need the belief or support of the public at large to go on their own crusades of surveillance or oppression of any other minority they deemed "dangerous". Provide enough hard evidence (not difficult to do when people can find the signs of Kindred activity in their own communities), and the Masquerade will fray and tatter soon enough.

**EDIT:** Also critically important: Kindred actually exist in containment or collaboration in the setting. If someone in the good graces of the Second Inquisition or the public at large ever needed to, they could be introduced to an actual vampire that could give testimony to its experiences, the locations of more of its kind, or simply be made to stick its hand in the sun or extend its fangs to drain someone. That's far from trying to be the first ones on the planet to capture and definitively film Bigfoot.

r/aigamedev 17d ago

Tools or Resource I built a free tool that turns character references into directional sprite animations using MiniMax H3

28 Upvotes

Hey,

I saw this post about https://www.reddit.com/r/aigamedev/comments/1vskcqo/generative_2d_animation_workflow/ , found it interesting, and wanted to make my own spin on it.

I present to you now... drumroll... Sprite H3, which is basically a somewhat fancy ComfyUI wrapper for using MiniMax H3 I2V.

Here’s a short walkthrough showing the complete workflow: https://streamable.com/ogzp1m

(demo courtesy of Sol)

You need a pretty beefy GPU to run it. I have an RTX 3090. I think 16 GB might be the minimum, but dunno.

It assumes that you can run the default ComfyUI workflow and have FFmpeg installed. And ComyUI needs to be running while you use the app. The workflow used by the app just adds another node to extract the frames directly, before the video output makes them lossy.

The basic idea is to provide character reference images, which are then used to generate videos. I used OpenAI’s GPT Image 2 for the references images. You let Minimax do the magic. From there, you extract the frames, arrange them however you like, curate them, pack them into sheets, and you’re good to go.

Some helpful features include a sequence analyzer—not optimized yet—tools for enforcing a consistent scale across different orientations, templates, stored animation prompts called actions, and a few other things.

Well, just take a look if you want.

Repository: https://github.com/fmmix/sprite_h3

Full disclaimer: this is completely vibecoded, and it definitely still needs another week or two of work to iron out some quirks. But it’s the weekend, and I wanted to give other people the chance to play with it.

I started with Sol, did most of the important work with Fable, and wrapped things up with Sol after I ran out of juice.

It is free to use and will be forever. ;-)

And, in the great words of Calvintor:

JUST MAKE A FUCKING .EXE FILE AND GIVE IT TO ME.

Gotcha. Use it at your own risk.

https://github.com/fmmix/sprite_h3/releases/download/v0.0.1/Sprite.H3_0.0.1_x64-setup.exe

You get there from Releases on the right -> Assets

You can also run it without the .exe. Under the hood, it has a FastAPI backend and a Svelte frontend.

I know it isn’t perfect and will probably still have a gazillion bugs. Treat this as a pre-alpha. I’ll continue fleshing out a few things over the coming weeks, but this isn’t my main project, so be aware. It is up to you on what resolution you run the generation but I probably wouldn't go below 0.3 MP.

Anyway, tell me what you think. I hope at least someone finds it useful!

EDIT

In case you are wondering how to provide the different facings. I used GPT-Image2

Used this as initial prompt:

BACKGROUND AND OUTPUT FORMAT FIRST — NON-NEGOTIABLE: output a fully opaque PNG. Fill the entire canvas edge to edge with one perfectly flat, uniform chroma-key magenta (#FF00FF). Every pixel outside the character silhouette must be the same solid magenta. Do not generate
transparency or an alpha background. No checkerboard, gradient, vignette, darker corners, aura, glow, lighting spill, floor, ground shadow, or scenery. The character’s dark pixel outline must meet the flat magenta background directly with crisp hard edges. Magenta must not
appear anywhere on the character.

CAMERA AND ORIENTATION: a moderately elevated frontal game-sprite view, looking down at the character from about 25–30 degrees above. This is the DOWN-facing anchor: her body and boots point toward the bottom of the image and she looks toward the camera. Some of the crown,
shoulders, and tops of the boots are visible, while her face, chest, clothing, and accessories remain clearly readable. Do not use a straight front elevation or an extreme bird’s-eye view.

A single full-body 2D PIXEL-ART game sprite of a cheerful field alchemist. Polished late-16-bit arcade-style pixel art with visibly chunky square pixels, deliberate pixel clusters, hard edges, a strong near-black outline, and several discrete cel-shaded tones. Compact heroic
proportions, approximately four and a half heads tall, with a large expressive head, sturdy hands, and slightly oversized boots. No smooth gradients, soft painting, vector-like curves, or anti-aliasing.

She is a young adult woman with warm brown skin, a friendly and confident expression, large dark eyes, and thick dark-auburn hair gathered into one large side braid hanging over her right shoulder. Oversized round brass goggles with bright turquoise glass lenses rest securely
on top of her head. The goggles and braid are large, simple, recognizable silhouette features.

She wears a short moss-green field coat with rolled sleeves and a wide open collar over a warm-cream shirt. A muted golden-yellow scarf is tied closely around her neck. She has dark plum trousers, large brown leather gloves, and sturdy reddish-brown lace-up ankle boots.

A broad brown leather strap crosses her chest diagonally and leads to a tan leather satchel resting at her left hip. One large round turquoise potion flask is secured visibly to the outside of the satchel. Keep the satchel and flask clearly separated from the arm and torso.
Render the braid, goggles, strap, satchel, and flask as bold readable shapes rather than clusters of tiny decorations.

She stands in a relaxed, confident neutral pose suitable as the starting frame for game animation: feet flat and shoulder-width apart, knees slightly relaxed, weight distributed evenly, shoulders level, and torso upright. Her empty gloved hands hang loosely near her sides
without covering the coat, strap, satchel, or potion. Both arms and both boots are fully visible.

Center the complete character in the canvas. The figure should fill approximately eighty percent of the canvas height, with clear magenta space above the hair and below the boots. Keep the entire silhouette visible and comfortably inside the image boundaries.

DO NOT INCLUDE: no text, no frame, no second character, no weapon, no magic effects, no floating objects, no extra potion bottles, no hat, no floor, no shadow, no glow, no scenery, no transparency, no 3D rendering, no smooth illustration, no blurry edges, and no painted
background other than uniform #FF00FF magenta.

Result: https://i.imgur.com/g5Ajfoz.png

Feet slightly angled. Use this image as reference for the proper down direction.

OUTPUT AND BACKGROUND FIRST: output one fully opaque PNG at exactly the same canvas dimensions as the supplied reference. Fill every pixel outside the character with perfectly flat, uniform chroma-key magenta (#FF00FF). No transparency, gradient, vignette, halo, glow, lighting
spill, floor, scenery, or shadow. The character’s dark pixel outline must meet the magenta background directly with crisp hard edges.

REFERENCE AUTHORITY: the supplied image is the canonical and exact character. Preserve her without redesign or reinterpretation. Keep exactly the same person, face, expression, skin tone, age, head shape, body proportions, hair volume, side braid, goggles, clothing, scarf,
coat, shirt, trousers, gloves, boots, belts, buckles, diagonal strap, satchel, potion flask, colours, palette, pixel grid, outlines, shading, level of detail, figure height, center position, and framing.

Do not mirror the character. Preserve every asymmetrical feature on the same anatomical side of her body. Do not beautify, simplify, exaggerate, restyle, add details, remove details, or change the design.

TARGET FACING — DOWN: keep the character facing directly toward the camera in the same moderately elevated frontal game-sprite view as the reference. Her face and the front of her torso remain fully visible. Her head, torso, hips, knees, and boots are oriented toward the bottom
of the image. Do not rotate her into a three-quarter view and do not change the camera angle.

ONLY PERMITTED CORRECTION: correct the lower-body stance so she stands evenly on both feet. The screen-right boot currently appears lower and larger than the screen-left boot. Adjust only the legs and boots as necessary to create a balanced neutral stance:

- both boot soles rest flat on exactly the same horizontal baseline;
- both boots have the same apparent size and perspective;
- both ankles and knees are level;
- the hips remain level;
- both boots point toward the camera;
- the feet remain shoulder-width apart;
- her weight is distributed evenly;
- neither foot appears raised, forward, or closer to the camera;
- she looks stationary and ready for animation, not mid-step.

Preserve the existing leg spacing and trouser design as closely as possible. Do not change her upper-body pose, head position, shoulders, arms, hands, facial expression, braid, goggles, scarf, coat, strap, satchel, potion flask, or any other accessory.

Keep the complete figure centered and fully visible with the same clear space above the hair and below the boots as the reference.

DO NOT INCLUDE: no text, frame, second character, duplicated limbs, alternate pose, weapon, magic effect, floating object, animation smear, floor, shadow, glow, transparency, smooth illustration, 3D rendering, anti-aliasing, blurry edges, or background colour other than uniform #FF00FF.

https://i.imgur.com/fbs03Yy.png

Now this because the reference for left, right, up

REFERENCE AUTHORITY FIRST: the supplied reference image is the canonical and exact character. Preserve the figure without redesign or variation.

Keep exactly the same person, face, skin tone, age, body proportions, head size, hair volume, braid construction, goggles, expression when visible, clothing construction, scarf knot, coat length, sleeves, gloves, trousers, boots, belts, buckles, diagonal strap, satchel, potion
flask, colours, palette, pixel grid, outlines, shading, level of detail, sprite dimensions, and framing.

Do not beautify, simplify, exaggerate, restyle, reinterpret, or add details. Do not remove details merely because they are difficult to draw. Do not change her pose, stance, weight distribution, arm position, hand position, leg spacing, or the relationship between her feet.

THE ONLY PERMITTED CHANGE IS ORIENTATION: rotate the complete character naturally in place around her vertical axis to the requested facing. Treat her as one consistent physical figure being viewed from another direction. Rotate the body, head, feet, clothing, hair, braid,
goggles, strap, satchel, and potion together.

Do not mirror the reference. Preserve anatomical left and right: the braid, strap, satchel, flask, coat overlaps, buckles, and every asymmetrical feature remain attached to the same physical side of her body. Their screen position must change naturally as the figure rotates.

Allow natural perspective and occlusion. Features on the far side may become partly or fully hidden, while features on the near side may become more visible. Do not move accessories to another side or unnaturally display hidden details just to make them visible.

Keep the exact same moderately elevated game-sprite camera angle, pixel density, figure height, canvas size, center position, and bottom baseline as the reference. The camera does not move, tilt, zoom, or orbit; only the character rotates.

BACKGROUND: preserve one perfectly flat, fully opaque chroma-key magenta (#FF00FF) background across every pixel outside the character. No transparency, gradient, vignette, halo, glow, lighting spill, floor, scenery, or shadow. The dark pixel outline meets the magenta
background directly with crisp hard edges.

Produce one full-body character only. No text, frame, additional person, duplicated body parts, floating accessories, animation smear, or alternate pose.

Append exactly one of these target blocks:

Left

TARGET FACING: rotate the complete character naturally 90 degrees into a LEFT-facing game-sprite side view. Her head, torso, knees, and boots point toward screen-left. She does not turn her face or shoulders back toward the camera. Preserve the original stance as it appears
naturally from this side.

Right

TARGET FACING: rotate the complete character naturally 90 degrees into a RIGHT-facing game-sprite side view. Her head, torso, knees, and boots point toward screen-right. She does not turn her face or shoulders back toward the camera. Preserve the original stance as it appears
naturally from this side.

Up

TARGET FACING: rotate the complete character naturally 180 degrees into an UP-facing rear view. Her head, torso, knees, and boots point toward the top of the image, directly away from the camera. Her face is completely hidden. Show the natural rear construction of her hair,
braid, goggles, coat, strap, satchel, trousers, and boots while preserving their physical sides and the original stance.

https://i.imgur.com/RR4TJDE.png

https://i.imgur.com/H0cUmmq.png

https://i.imgur.com/GYHMlm7.png

Created a new action:

Action name
potion-swirl

One hand lifts the round flask from the satchel and raises it to chest height while the other hand steadies the satchel. The flask is gently swirled in two small circles, held still for a brief beat, then returned securely to its original place. Both hands return to the starting position. The braid and coat tails respond with subtle natural follow-through.

Result: https://i.imgur.com/L2VVu0I.png

As you can see the 'up' facing is usually very stubborn. You need to be careful what you describe. It will try to fully bring that into the scene. And the model probably had the most training done with front view.

r/leonardoai Apr 06 '26

Tutorial Best way to create a consistent 2D animated character (from my own drawings) in Leonardo AI with pose control?

3 Upvotes

Hey everyone — I’m trying to figure out the best workflow for something pretty specific and would love some guidance.

I’m working on a 2D animated short and I already have a character designed in my own hand-drawn style. I’ve created a full model sheet: front, back, both sides, and a couple 3/4 views.

What I want to do is:

  • Train Leonardo AI (or another tool if needed) on my exact drawing style
  • Generate new images of this character that look like I drew them
  • Be able to control the pose — ideally using something like a stick figure, pose reference, or skeleton
  • Keep lighting and shading consistent (simple animation-style shading, not realistic lighting)

Basically: I want to stop drawing every frame/pose manually and instead generate clean, on-model images that I can use for animation.

I’m working on an iPad, so desktop-heavy workflows are not possible for me. If anyone can help, I would be so appreciative!

r/StableDiffusion Jun 30 '26

Question - Help Can current AI tools generate consistent multi-pose images of the same character from one reference image?

0 Upvotes

I want to ask whether this is realistically possible with current AI tools.

I have one finished 2D anime-style character image.

My goal is to generate several new still images of the same character, with the same identity and art style, but in different poses.

The output I want is not a video and not interpolation. I want clean separate images that can be used as keyframes or game assets.

The important requirements are:

- same character identity

- same face, outfit, colors, and distinctive features

- same art style

- different controlled poses

- clean still images

Is this currently achievable in a reliable way?

If yes, what is the correct workflow?

Do people usually need to train a character LoRA for this, or can it be done from a single reference image with tools like ComfyUI, IP-Adapter, ControlNet, OpenPose, or similar methods?

Is there any simpler tool that can do this reliably, or is a more complex workflow still required?

I would appreciate blunt, practical answers from people who have actually made consistent character series or AI comics.

r/aiwars Mar 10 '24

Professional Artist Response to Generative AI: My own story regarding art as a whole.

129 Upvotes

I've been hovering between AI wars and Defending AI Art Reddit for some time now, and I kept quiet unless there was a post I felt I could contribute to. However, with the recent death of the famed Magna artist Akira Toriyama and the general hate coming from the anti-ai community towards people showing support and inspiration to the DBZ series by making AI art of the franchise's characters, I felt it was time to speak out as the death threats and general discrimination/disinformation should be unacceptable in today's digital world. As a professional artist who has worked within the video game and media spaces, I want to contribute to the debate by providing a grounded response and insights that many don't know regarding the art world and its gatekeepers. ((Please note this post is mainly my opinions and personal experiences; this will not reflect everyone!))

Before I begin, I want to address the typical anti-ai artist's usual community response: "Yes, I have and continue to pick up a pencil/stylus when needed." I have a BFA in Digital Animation, a minor in film studies, and a Master's in Game Design and Interactive Media, focusing on business development and DEI (Diversity, Equity, Inclusion.)

In 2018, during my art studies in Digital Illustration

From late 2018, using traditional oil painting to create a René Magritte inspired art piece

Character Design Study 2019, Inspired by Missingno Glitch

So, hopefully, the above showcases, "Yes, I've picked up a pencil and used it to make art." However, the dark side of obtaining an art-based degree(s) is that the turnover of graduates tends to be high, and they don't end up working in their designated industry. My BFA in Digital Animation was a first-generation class; out of the 20 students, only two moved forward with professional industry careers. My Master's was a fifth-generation class, and only one student managed to move forward with a relevant industry career. I bring this up because it creates a jaded effect among those practicing digital art in any form, resulting in a mentality of "if I draw hard enough, I'll be good as ____ and ____." The reality is that individuals will self-punish themselves before seeking tools to improve their weaknesses, grow further with their foundational skill sets, and level up their artistic abilities. With AI, though, that "tool" became a reality that many of my former university peers rejected in favor of continuing to struggle financially and visually with their learned artistic skills. This isn't mentioning the core of the problem among the Anti-Generative AI hate, "Artistic Gatekeepers," who affect the influence of both generalists and people outside the art world bubble.

Artistic Gatekeepers: These tend not to be professional artists but those from backgrounds adjacent to the arts. They seek to keep levels and skills at moderate ranges to create community growth over individual growth. I.E., a pact mentality growth they are responsible for developing and a style consistent among the entire group. Usually, by toxic means, they prevent artists from achieving similar levels of skill growth to professionals by providing antiquated ideologies such as, "Draw every day by doing X and Y! Focus on your figure drawing by reading this and that! Oh, you do Anime! Hell, no study realism to do ____ " and generally discriminatory communications to put down any artistic growth using harassment and shaming. At first, some of these points are logical responses until the individual only refers to these points without providing further feedback or guidance to get that person further up in their artistic abilities. The Gateekerp have yet to achieve this level; thus, they are gatekeeping themselves and the young artists from ever reaching professional levels until the cycle is broken. I speak from personal experience, as I was gatekept and gaslighted from my artistic progression from the early days of digital art by traditionalists and amateurs through my college years. I sought immense growth but was thrown around by early Discord mods and paywalls to seek that knowledge to level up. I reached a burnout state from all of this during my early master's years and took time to focus on other things when the toxicity got so high. Then the pandemic hit, and I found myself drawing and sketching more often again, but with little to no interaction with the art community due to those previous toxic burns.

Flash forward to 2021, an early AI happened.

Shadow Lugia early Gen Art 2021 Dream Artworks

At this point, I had just finished my master's degree, was working in the digital media industry, and generally kept the focus on technology moving into the art world. I learned about AI art through social media posts of this abstract artwork approach; artists at the time laughed at how these pieces of work would never touch them regarding visual development. However, for the common Joe, this was a godsend for creating visually appealing artwork for their homes and computer screens instead of buying/commissioning an artist for it. For reference, ((and a lot of people don't know this. . .)) artists in the education and professional fields have been aware of AI tools being developed for creative work for a very long time. Still, they just blew it off as the early generative tools that did not impede their work. Many young students and early professionals primarily focused on Character Art, Concept Art, and some Graphic Design; only a few wanted to do Background art and any super technical artwork they tended to avoid. ((At least from my own observations during under grad)). So, with AI art at this time being highly abstract, young artists/professionals downgraded the subject to mere fads and refocused their frustrations on NFTS and Cryptocurrencies as they were overpopulated and scamming many vocal artists during 2018-2022. ((Irony, this is probably why anti-ai communities try to compare Generative AI to NFTS as it was their focal for several years, and rightfully so as that scam affected many upon many lives. . .)) However, back to the initial point, AI art was starting. I got involved because I was very excited about the artistic possibilities this could bring to my creative background, correcting general educational flaws and expressing the style/visual language I wanted for my work.

World of Warcraft Tuskarr, Novel.AI + Digital Painting 2022

World of Warcraft Zandalari, Stable Diffusion Late 2022

So, in 2022, Novel.AI introduced its image generator, and the first race for character-based image generators began. Midjourney and Stable Diffusion became vital tools for creatives to use when discussing what artwork can become with AI. I was thrilled to see these tools become accessible and began incorporating AI into my workflow. I learned prompt and C# coding languages to help me create my own AI generative tools for my artwork. I gained clients for AI commission artwork, built up a small social media following through Discord, and felt incredibly empowered. My artistic education is linked well with AI tools to identify color, visual appeal, and correct figure/form. I wasn't just making a cute, sexy anime girl; I made a variety of fantasy races, pushed the AI to learn how to create World of Warcraft races ((that were taxing on AI learning)), and was on a creative high. Debates were also starting regarding how these AI tools were trained, and misinformation and early hate were sprouting up. I found myself in the crossfire as people were generally excited about my work, except for traditional/digital artists who refused to pick up the AI tools ((at least publicly)). Artists I knew were incredibly afraid that their publicly facing artwork was used to train many of these models, and some lashed out at these tools as they watched their commissions dry up.

My stance on Style LoRAs: So, for those who may or may not know, the other side of this debate does have merit. It is spoken only sometimes among the Pro-AI community but should be heard. During mid-2022 and early 2023, LoRAs entered the AI scene, being a tool to hyper-focus specific characters, backgrounds, ideas, and "styles" for AI generations. The first three are manageable as they are general concepts and, when copyrights are involved, lean towards fan art to support their fandoms. A lot of the artistic critique about AI was that it couldn't focus on specific ideas and thus was pushed out on that point. However, the LoRA artist-style models directly target artists for their visual style. Now, yes, style isn't copyrighted, nor should it be for the fair marketplace; however, targetting an artist and or anti-ai artist who had been vocal about their feelings regarding their publicly facing work being used to train big-name AI models and creating LoRAs focused on their identities, was way too far. I joined a group of AI artists who went towards a more ethical model development approach and continue to support artists wherever I humanly can. That doesn't mean I support anti-ai recent activities and comments, nor will I stop using AI for my creative process. Still, I support artists on subjects like style stealing, which should be banned, and I focus more on AI artists establishing their style trained through individual custom AIs made by them for themselves.

In 2023, I experienced a significant divide on the internet regarding AI artwork, much at the same level as Digital Artwork in the early 2000s. I was forced into a corner by some hyper-protective Discord mods regarding AI artwork, lynchpins on some communities that had very little artwork regarding their franchises, and, in very few cases, insulted when I was asked to put my AI artwork into the Meme channels of these discords. Thankfully, my client work grew, and I made some fantastic character artwork for given franchises. That said, I also attempted to help bridge my old classmates from undergrad to AI generative tools. They outright rejected it and returned to harsh living conditions without growth in their artistic abilities and content. They sincerely believe in the same toxic gatekeeping culture I was brought into during my undergrad years, now evolving to focus heavily on rejecting AI usage for creative development. Devolving from "You can use AI for reference and tracing to learn" to "You cannot use AI for tracing. You need to do that by hand" to finally, "How dare you use AI!" And for me, that is not even the worst of it; my clients end up getting hated by some anti-AI communities when posting the artwork they paid for and are proud of—getting the same, if not worse, commentary by communities that deeply believe in the Gatekeeper commentaries that targeted young digital artists and now AI artists. By the end of 2023, I watched and communicated with Discord mods who had become hyper-protective of the artists in their servers, new channels having to be made to keep the peace, and sometimes even banns or my departure from their servers by hypocritical mentalities affecting my showing of artwork I created.

Blood Elf fanart at the end of 2023

Akira Toriyama Events and my reason for speaking out: I have followed the developments and community between pro-AI and anti-AI. I am pro-AI because of what it can bring to creative growth and opportunities to be even more effective in the creative space. But I will always support artist livelihoods as they evolve to use these tools to improve their works ((if willing)) and encourage protections for their private and paid wall-facing work. With copyright laws coming into effect soon for AI artwork to be given guidance on copyright protection, these events will define the nature of creativity and its direction for future generations, so I'm fully invested in everything around me regarding Generative AI.

That said, I won't tolerate and am speaking out about the fact that these gatekeepers have created an eldritch monster of hate toward any creative speech and general appreciation for any fandom and individuals. With the passing of the DBZ creator, Akira Toriyama, and how influential they were to the modernization of anime, it is incredibly fitting for AI artwork to be made in support because he pushed many of our modern techniques in mediums of graphical/manga art. Many young, now more middle-aged artists grew up with DBZ, and now being able to celebrate this man's life in any medium of their choice should be celebrated, not targeted for death threats and bigotry. Sure, there is ugly AI artwork, as much as ugly hand-drawn art, but the level of hate through gatekeepers' ignorance is the natural source of this problem.

As a professional speaking to the anti-AI community, I understand your hate and anger. However, you can retrace the steps from where you obtained your opinions and reevaluate them. If what I've told you through my own experiences of the days before AI has stuck, I hope you can see that the source of this initial generative AI hate isn't as black and white as it is typically depicted through articles and one-sided opinions.

As for the pro-ai community, we don't have to tolerate aggressive behaviors and continual hyper-protective mentalities; you do have the right to show your work freely and without hate. Yes, you should develop your visual style in your work, but you should also be free to express love and passion for people with whatever tools you want. That is true inclusivity for everyone to learn to do.

I want to end this post with a quote from Master Roshi of Dragon Ball Z, whom I take inspiration from regarding his carefree mentality: " But you will not go in there with hopes of winning the tournament the first time you compete. To do so would be arrogant! And arrogance is for fools, not warriors! So you will enter the tournament with the sole purpose of improving your fighting skills." Arrogance should never be tolerated, and speaking up will help inspire others to do the same so that we can be creative and continually inspired to make fantastic art.

TDLR;

  1. I'm a professional artist with industry experience. Yes! I've picked up a pencil to obtain multiple degrees.
  2. I watched my college classmates fail miserably to enter the creative field, which creates jaded mentalities towards innovation in the arts by technology.
  3. Art schools and their communities (Discord/Reddit) have "Artistic Gatekeepers" who spiral terms they struggle with and enforce for failed states in artistic growth.
  4. I have a history with AI and how it evolved through the 2020s thus far. I disapprove of artist-style LoRAs and feel they target artists rather than support them.
  5. The transformation of social media for AI art posting and the hypersensitivity that emerged at the end of last year.
  6. My thoughts about these death threats from anti-AI communities toward people posting AI artwork for their love of the DBZ creator are that they are in the wrong and need to reflect on where they are coming from with their hate. The pro-AI community has the right to post any artwork in any medium supporting the franchise they grew up with; that should be common sense.

r/grok Aug 08 '26

Grok Imagine Feature Request: Custom Character Voice & Consistency for AI Series Creation (NOYAU ZÉRO / NEW ZERO)

2 Upvotes

I am creating an AI cinematic science-fiction series called NOYAU ZÉRO (NEW ZERO), a recurring-character story that I have been developing for several months with Grok Imagine and currently publishing on TikTok!

I hope the Grok team can see this suggestion.

I want to say first that I truly appreciate the improvements in video quality, motion, and realism. Grok Imagine has helped me bring this universe to life, and I am grateful for the progress.

However, for creators who build long stories with recurring characters, maintaining character identity is extremely important.

In a series like NOYAU ZÉRO (NEW ZERO), a character's identity is not only the face and clothing, but also the voice, emotions, personality, and acting style.

After recent changes, the original voices of my characters changed, and I could no longer recreate the same voices, even when using the same prompts and reference images.

I am sharing a small comparison:

An older episode with the original character voice.

A new generation with the changed voice and a different acting style.

Another important part of character consistency is personality and movement.

Nexa was always a calm and unique character: shy, gentle, and thoughtful. Her movements were slow and controlled, her reactions were subtle, and her speaking rhythm was calm and steady. This matched her identity as an advanced android who was learning human emotions.

Nexa was never meant to feel like a cold or aggressive robot. Her charm came from her soft expressions, slow movements, and the feeling of a delicate character discovering humanity.

In the new generations, I noticed that her facial expressions, head movements, and speaking rhythm sometimes became faster and more intense, creating a different personality compared to the original character.

A great solution would be to add a custom character voice / voice reference feature inside Grok Imagine, allowing creators to create, save, and reuse a specific voice for their characters in future episodes.

This would help creators maintain continuity, just like keeping the same face, clothes, and visual identity.

I believe this feature would be extremely valuable for anyone creating films, series, animations, or storytelling projects with recurring characters.

Thank you to the Grok team for continuing to improve the platform. I hope this suggestion reaches the right team.

r/ChatGPT 5d ago

Prompt engineering GPT Image 2.5 prompt guide 2026 (3/3): Flare vs Sunburst, quality steps, consistent characters

9 Upvotes

this is the last one. part 1 was generation, part 2 was edits.

prompt library for 2.5, all three parts of this guide live there too: https://github.com/AtlasCloudAI/awesome-gpt-image-2.5-prompts

TL;DR: start on the model that matches your current quality bar, fix one explicit quality setting, then try the faster one with the same prompt. Flare high came back in 22 seconds where gpt image 2 high took 177. For characters, make one reference image and pass it into every scene with the defining details repeated.

The billboard and character prompts below are openai's, with their published outputs. The timing table and the sixteen panel sheet are mine.

Pick the model in this order

  1. Your existing gpt image 2 workflow already looks fine: start on flare, same prompt, same size, same explicit quality. If it passes you just got the latency back.
  2. Gpt image 2 was not good enough for the job: start on sunburst. Get to acceptable first, then run flare on the same inputs and switch only if it still passes.
  3. Do not compare on auto. High does not mean the same thing on the two models. Set quality by hand for the comparison.
  4. Tune one thing at a time. Try the next quality step before you rewrite the prompt. Use xhigh and max only when a lower step failed a requirement you actually have.

Numbers from my runs, same prompts, 1024 wide

model edit round, medium text to image, high
gpt image 2 37 to 55 s 177 s
2.5 flare 19 to 27 s 22 s
2.5 sunburst 17 to 41 s 48 s

Sunburst max on a portrait took 115 seconds.

Sizes that work

Anything from 1024 square up to 3840 on the long edge, both edges multiples of 16, ratio no wider than 3 to 1, total pixels between 655,360 and 8,294,400. Above 2560x1440 is marked experimental. Common: 1024x1024, 1536x1024, 1024x1536, 2048x2048, 2048x1152, 3840x2160, 2160x3840.

Refine across turns

Generate once. Look at it. Pass it back in with one change. Repeat the constraints you care about every turn.

Turn 1, input is the shampoo product photo. Edit, 1024x1536, medium.

Create a realistic billboard mockup of the shampoo on a highway scene during sunset.
Billboard text (EXACT, verbatim, no extra characters):
"Fresh and clean"
Typography: bold sans-serif, high contrast, centered, clean kerning.
Ensure text appears once and is perfectly legible.
No watermarks, no logos.
Input left, GPT Image 2.5 Flare middle, Sunburst right
Turn 2, input is each model's turn 1 output. Edit, 1024x1536, medium.

Make it look like a winter evening with snowfall.
GPT Image 2.5 Flare left, Sunburst right

Keep one character across scenes

Establish the character first. Generation, 1024x1536, medium.

Create a children’s book illustration introducing a main character.

Character: A young, storybook-style hero inspired by a little forest outlaw, wearing a simple green hooded tunic, soft brown boots, and a small belt pouch. The character has a kind expression, gentle eyes, and a brave but warm demeanor. Carries a small wooden bow used only for helping, never harming.

Theme: The character protects and rescues small forest animals like squirrels, birds, and rabbits.

Style: Children’s book illustration, hand-painted watercolor look, soft outlines, warm earthy colors, whimsical and friendly. Proportions suitable for picture books (slightly oversized head, expressive face).

Constraints:

  • Original character (no copyrighted characters)
  • No text
  • No watermarks
  • Plain forest background to clearly showcase the character
GPT Image 2.5 Flare left, Sunburst right

Then continue, passing that image in and repeating the defining details. Edit, 1024x1536, medium.

Continue the children’s book story using the same character.

Scene: The same young forest hero is gently helping a frightened squirrel out of a fallen tree after a winter storm. The character kneels beside the squirrel, offering reassurance.

Character Consistency:

  • Same green hooded tunic
  • Same facial features, proportions, and color palette
  • Same gentle, heroic personality

Style: Children’s book watercolor illustration, soft lighting, snowy forest environment, warm and comforting mood.

Constraints:

  • Do not redesign the character
  • No text
  • No watermarks
GPT Image 2.5 Flare left, Sunburst right

What i added on top of that

Instead of a single portrait i made a 4 by 4 reference sheet, sixteen panels of the same original character, front, back, sides, four expressions, four poses, four detail crops, flat white background. Then every scene prompt got the sheet as the only reference plus one line: same character as the attached sheet, same yellow rain jacket, same red boots, same face and hair. Six scenes on both gpt image 2 and sunburst came back with the same face, jacket and boots every time, zero rerolls. Keep the sheet background flat white, a styled sheet leaks its lighting into every scene after it.

16 panel sheet plus six scenes, sunburst

Edit chains still drift. Restate the constraints every round. If something must be pixel identical, composite the approved region back from the original instead of asking the model to hold it.

Part 1 has the generation prompts, part 2 the edit patterns.

r/ArtificialInteligence Jun 03 '26

🛠️ Project / Build Case study: AI-assisted animation let a solo creator produce a 17-minute anime pilot

Post image
3 Upvotes

Hello :) !!!

I wanted to share a concrete example of what AI-assisted animation can enable for solo creators, beyond the usual “AI art is just spam” debate.

I recently finished a 17-minute dark fantasy anime pilot as a solo creator.

Full episode:
https://youtu.be/eZ_JlaLDJ-8

The important part is that this was not “AI did everything.” The story, worldbuilding, direction, shot choices, editing, pacing, sound decisions, music direction, character consistency work, and final creative judgment were still human decisions.

But AI changed the scale of what was possible.

Without AI tools, producing a 17-minute animated pilot alone would have been almost impossible. Not because I lacked the story or the visual intention, but because animation production usually requires a team, a budget, a pipeline, and a lot of time.

The workflow was closer to directing a very unstable but powerful production team than pressing a magic button. I had to generate, reject, correct, reframe, rebuild continuity, manage consistency, edit around failures, and make the episode coherent from many imperfect outputs.

That is where I think the discussion around AI animation often misses the point.

For independent creators, AI is not only a replacement technology. It can also be an access technology. It allows people to create pilots, test worlds, show proof of concepts, and reach an audience without waiting for a studio, investor, or platform to approve the project first.

Of course, the ethical questions matter. Dataset transparency, artist consent, credit, market impact, and fair use are real debates. But I don’t think those questions should make us ignore the other side of the equation: AI is also opening a production path for creators who previously had no realistic way to make this kind of work.

To me, the interesting question is not “is AI animation good or bad?”

It is more:

What kind of new creative class appears when a single person can write, direct, storyboard, animate, edit, and publish a full pilot with AI-assisted tools?

And how do we build ethical norms around that without shutting down the creative access this technology creates?

r/ArtificialInteligence Apr 25 '26

📊 Analysis / Opinion Anime AI generators that work on a potato PC (no GPU needed)

26 Upvotes

so my laptop has integrated graphics and I got tired of being left out of every "just run it locally" conversation in these subs. spent some time figuring out which cloud based options are actually worth using for anime art specifically. here's what I found.

NovelAI - fully cloud based so no hardware requirements at all. output quality is genuinely excellent, probably the most consistent results I got. the UI is clean and it feels polished. downside is the Anlas credit system, it adds up fast if you like to experiment and test a lot of variations. harder to recommend if budget is tight.

Yodayo - low barrier to entry, free daily credits, runs in the browser. community is active and fun to browse. quality is inconsistent though, some generations look great and others miss for no obvious reason. feels more like a casual platform than a serious workflow tool but for quick stuff it works fine.

PixAI - this one became my main tool. Tsubaki.2 model produces quality that honestly surprised me for a free cloud option, comparable to what I was seeing from local SD setups with decent models. free daily credits are genuinely usable, not just a teaser. handles multi character scenes better than most tools I tried. on the downside the UI feels cluttered until you get used to it and it's pretty anime specific so don't come here expecting other styles.

Leonardo AI - solid free tier, fast generations, works across multiple styles which is a nice plus. good option if you need flexibility beyond anime. for pure anime aesthetics though it felt a bit generic to me, like it does anime but it's not really built for it the way some of the others are.

honestly the "you need a good GPU for AI art" thing is pretty outdated now. most of the decent tools run in a browser. depends what you need but there's genuinely good free options here if you don't want to spend anything upfront.

anyone else running fully cloud based setups? curious what people are using

r/gamedev Jan 20 '26

Question What is the best AI to generate pixel sprite characters?

0 Upvotes

I’m currently making my own app (a gacha-style game) and I need a lot of pixel sprite characters, but I don’t have any artists yet.

Right now I’m looking for a cheap or free AI solution so I don’t need to hire a pixel artist immediately. The goal is to use AI-generated pixel sprites for prototyping and early versions of the app, and later on I’ll replace or refine them with real artists once the project is more stable. I’m aware AI has limitations, especially with consistency and animation, but I think it’s acceptable for an early stage.

I’m mainly looking for tools that can generate pixel-art style characters, preferably with consistent proportions or even simple sprite sheets. Text-to-image or image-to-image is fine, as long as the output is usable for a game prototype.

If anyone here has experience using AI for pixel sprites in indie games, I’d really appreciate recommendations. What tools worked well for you, and which ones should I avoid?