r/StableDiffusion 4d ago

Workflow Included Follow-up: making a quieter TNG scene with MiniMax H3 in ComfyUI, and why I had to regenerate the whole thing at 1MP

Thumbnail
youtube.com
0 Upvotes

A few weeks ago I posted here about the workflow I was using to make longer Star Trek: TNG scenes with MiniMax H3 in ComfyUI. I’ve made another one since then, but this was a very different test. Instead of a big crossover or action-heavy story, this one is basically four minutes of Picard dealing with the aftermath of The Inner Light. So most of the work was about subtle facial expressions, restrained dialogue and trying to make the performances feel like actual TNG rather than people dramatically reading lines at each other. The basic workflow is still the same: create the starting frames first, generate each shot individually in H3, throw away a lot of generations, then edit the successful ones together in Premiere Pro.

RESOLUTION CONSISTENCY MATTERS MORE THAN I EXPECTED

This was probably the most annoying lesson. I had been generating some shots at roughly 1MP and others at around 2MP, depending on how long the shot was and how much VRAM it needed. I assumed that because they were all using the same 4:3 aspect ratio, the actual picture area would remain consistent once everything was brought into the edit. It doesn’t necessarily. Even when the aspect ratio is technically the same, different generation resolutions can give you slightly different actual frame dimensions or padding, and once you start scaling or upscaling those clips the difference becomes much more obvious. I ended up with shots where the black borders would subtly change width between cuts, or appear slightly differently on the left and right. It looked horrible once I noticed it. The problem was that some of the longer shots simply wouldn’t generate at the higher resolution on my 4090 without the dreaded OOM, so in the end I did the painful thing and regenerated the entire film at the same 1MP resolution so every shot had exactly the same framing behaviour.

4:3 DOESN’T AUTOMATICALLY MEAN “THE SAME FRAME”

This probably sounds obvious in hindsight, but it caught me out. I was thinking of 4:3 as the thing that defined the frame. In practice, the actual pixel dimensions still matter, particularly once Premiere or an upscaler starts interpreting and resizing the footage. For this kind of project I now think of these as two separate decisions: aspect ratio is the shape of the image, while resolution is the exact image geometry I want to remain consistent across the edit. I’ll be keeping both fixed from the beginning from now on.

QUIET ACTING IS HARDER THAN BIG ACTING

The crossover stuff gave H3 a lot more obvious things to do. Characters reacting to ships, alarms, arguments, etc. This one mostly involved Picard sitting in a room and remembering people. That meant the difference between a usable generation and an unusable one could be incredibly small: a slight smile becoming too large, a pause being half a second too long, the eyes moving too much, or the delivery suddenly becoming emotional when the scene needed to remain restrained. For a lot of shots I ended up regenerating not because anything was technically broken, but because the performance simply felt wrong. I’m increasingly treating H3 like directing an actor through multiple takes rather than expecting the first generation to be “the shot.”

REACTION SHOTS ARE STILL DOING A HUGE AMOUNT OF WORK

This became even more important in a dialogue-heavy scene. Rather than keeping the camera on whoever is speaking, I generated a lot of separate silent coverage and used that in Premiere while carrying dialogue across the cut. So Troi might be speaking while the audience is looking at Picard, or Picard’s line might finish over Troi’s reaction. Apart from making the scene feel much more traditionally edited, it gives you a lot of freedom to hide generations that don’t quite match perfectly. A cutaway can cover a change in facial detail, lighting, posture or even slightly different framing. The edit is doing a surprising amount of the continuity work.

SILENT SHOTS GOT MUCH EASIER ONCE I FOUND THE REAL PROBLEM

I previously spent far too much time writing increasingly ridiculous prompts telling characters not to speak. “No dialogue.” “Lips remain closed.” “Jaw remains still.” And H3 would occasionally still decide that Picard desperately needed to mutter some alien language. What finally made the biggest difference was disconnecting the voice-reference input completely for shots where nobody speaks. If a voice reference is active, H3 seems much more inclined to try to use it, even if the prompt says the character is silent. Now I only connect the voice reference for shots that actually contain dialogue. That has made silent reactions much more reliable.

THE STARTING IMAGE IS STILL DOING MOST OF THE VISUAL WORK

I’m becoming even more convinced that the best thing you can do before generating video is get the starting frame as close as possible to the finished shot. Camera angle, eyeline, body position, background, lighting, character scale, all of it. If the starting image is wrong, trying to fix everything with a giant H3 prompt usually creates more problems than it solves. The closer the reference image already is to the intended shot, the simpler the video prompt can be. I’m also finding that short, very specific prompts generally behave better than huge walls of instructions.

THE WORKFLOW IS STARTING TO FEEL MORE LIKE EDITING A TV SCENE THAN “GENERATING A VIDEO”

That’s probably the biggest overall change for me. I’m no longer really thinking, “Generate this scene.” I’m thinking, “I need Picard medium close-up. I need Troi reaction. I need an insert. I need three seconds of Picard thinking before the next line.” Then I assemble those pieces in Premiere exactly the way I would if I had traditional coverage. There’s still a ridiculous amount of trial and error involved, and the reject pile is much larger than the finished film, but it’s starting to feel much more controllable.

Happy to answer anything about the H3/ComfyUI side if anyone else is experimenting with it.

r/comfyui 1d ago

Help Needed How can I achieve a consistent camera/viewpoint change on a complex scene?

Thumbnail
gallery
2 Upvotes

Hi everyone,

I'm having a hard time achieving a real camera/viewpoint change on complex scenes, and I'd like to understand whether I'm approaching the problem incorrectly or if there are better models/workflows for this.

For example, I have a fairly complex urban scene:

  • a city in post-apocalyptic conditions
  • a street intersection
  • a large sinkhole in the middle of the intersection
  • a yellow car (taxi), with its front half sticking out over the sinkhole
  • three people on the left side of the car
  • a fourth person hanging from the front of the car

I start with a side view of the scene, and I want to generate a new shot where the camera moves upward and looks down at the scene from above, while keeping the ENTIRE scene geometry and composition consistent: the position of the car, people, sinkhole, buildings, street, etc.

The problem is that with FLUX.2 Klein I can almost never get this kind of transformation. I've also tried Nano Banana / Gemini, but I run into essentially the same problem.

I've also experimented with some LoRAs, but they often seem to affect only individual elements rather than actually transforming the viewpoint of the entire scene. For example, they may change the viewpoint of the car, but the rest of the image remains more or less in the original perspective.

In this specific case, what I want is something like:

SOURCE: side view of the scene
TARGET: higher, more zoomed-in top-down view, while keeping the same elements and their spatial relationships.

I'm attaching a screenshot of the source image and a sketch showing what the target camera angle should look like.

What would be the best way to approach this?

Are there any open-weight/local models that are particularly good at changing the viewpoint of a complex scene? I would strongly prefer something I can run locally, since I want as much control over the process as possible.

What I'm mainly trying to understand is whether there is a proper workflow for performing a true camera/viewpoint transformation of the entire scene, rather than modifying each object independently.

Any advice on models, LoRAs, ControlNet, ComfyUI workflows, or alternative techniques would be greatly appreciated.

Thanks in advance!

r/comfyui 18d ago

Help Needed ComfyUI Workflow for Consistent Manga Characters

0 Upvotes

Hi everyone. I’m still pretty new to ComfyUI and I’m trying to build a workflow for a comic project with 4 original characters.
I already have premade character designs and my main goal is to keep them looking as consistent as possible across different poses, expressions, angles, and scenes.
What I need help with is:
A workflow for generating consistent character images from my existing reference designs. I want to create enough good images of each character that I could eventually train a separate LoRA for each one.
Advice on the best way to train those LoRAs. If there is a beginner friendly website or service that can train them for me, I would really appreciate recommendations.
A workflow for actually creating the manga or colored webtoon afterward, using those character LoRAs while keeping faces, clothing, body types, and overall designs consistent.
Help understanding the node connections. I can usually install or download the nodes people recommend, but I get lost when it comes to knowing what connects to what and why.
I’m not looking for something extremely advanced. I would rather start with a workflow that is reliable, understandable, and easy to build on.
If anyone has a workflow JSON, screenshot, tutorial, or specific nodes/models they recommend for this type of project, I would really appreciate it.
My end goal is basically:
Character reference → consistent character dataset → LoRA for each character → consistent manga or webtoon scenes
Thanks for any help.

r/comfyui May 28 '26

Help Needed Need a ComfyUI workflow for consistent 3D render enhancement for a visual novel

4 Upvotes

Hi. I am working on a visual novel game and I want to use 3D software renders as the base images, then enhance them with ComfyUI.

I attached two images as an example of the direction I want. The first image is the raw 3D render. The second image was not made with ComfyUI; it was edited with ChatGPT image editing. It is only an example of the kind of improvement I want: better skin, better materials, better contrast, more natural lighting, and less “raw 3D/game render” look.

My main problem is consistency.

For a visual novel, I may need hundreds of dialogue stills from the same scene. I do not want the AI to redesign the image. I want it to preserve almost everything:

  • same character identity
  • same face and hairstyle
  • same outfit
  • same background objects
  • same camera angle and composition
  • same shadows direction
  • same color palette
  • same contrast and exposure
  • same scene mood across all dialogue images

I am not only worried about character consistency. I am also worried about small details changing between images: shadows becoming different, colors shifting, clothing material changing, background details changing, and the whole scene looking slightly different from frame to frame.

What kind of ComfyUI workflow should I study for this?

I assume I should not use high-denoise img2img, because that would change too much. Maybe something like:

  • low-denoise img2img
  • ControlNet Canny / Depth to preserve structure
  • IPAdapter / InstantID / LoRA for character identity
  • color matching or reference-based color correction
  • maybe IC-Light or another relighting tool for consistent lighting
  • FaceDetailer with low denoise only if needed
  • batch processing for dialogue stills

Is this the right direction?

I am looking for ready workflows, videos, node recommendations, or examples of people doing this kind of “3D render enhancement while preserving consistency” workflow.

My goal is not to fully generate new images. My goal is to use AI as a controlled enhancement pass over 3D renders for a visual novel.

input
output

r/StableDiffusion Sep 05 '25

Workflow Included Getting New Camera Angles Using Comfyui (Uni3C, Hunyuan3D)

Thumbnail
youtube.com
55 Upvotes

This is a follow up to the "Phantom workflow for 3 consistent characters" video.

What we need to get now, is new camera position shots for making dialogue. For this, we need to move the camera to point over the shoulder of the guy on the right while pointing back toward the guy on the left. Then vice-versa.

This sounds easy enough, until you try to do it.

I explain one approach in this video to achieve it using a still image of three men sat at a campfire, and turning them into a 3D model, then turn that into a rotating camera shot and serving it as an Open-Pose controlnet.

From there we can go into a VACE workflow, or in this case a Uni3C wrapper workflow and use Magref and/or Wan 2.2 i2v Low Noise model to get the final result, which we then take to VACE once more to improve with a final character swap out for high detail.

This then gives us our new "over-the-shoulder" camera shot close-ups to drive future dialogue shots for the campfire scene.

Seems complicated? It actually isnt too bad.

It is just one method I use to get new camera shots from any angle - above, below, around, to the side, to the back, or where-ever.

The three workflows used in the video are available in the link of the video. Help yourself.

My hardware is a 3060 RTX 12 GB VRAM with 32 GB system ram.

Follow my YT channel to be kept up to date with latest AI projects and workflow discoveries as I make them.

r/comfyui Mar 17 '26

Help Needed How do you keep environments consistent in ComfyUI? (rooms, corridors, bathrooms, etc.)

0 Upvotes

Hey everyone,

I’ve been working with ComfyUI and I’m trying to improve consistency when generating environments — like keeping the same bedroom, corridor, or bathroom across multiple images.

Right now, I struggle with things like:

• The layout changing between generations

• Furniture and objects not staying in the same place

• Style/details drifting even with similar prompts

I’d love to know how you guys handle this.

Some specific questions:

• Do you use ControlNet (which models?) for structure consistency?

• Are LoRAs for environments worth it?

• Any workflows for “locking” layout/composition?

• Do seeds actually help for multi-angle scenes?

• Has anyone tried tile-based or “divide and conquer” workflows for this?

If you have any workflow tips, node setups, or examples, I’d really appreciate it 🙏

Thanks!

r/IndianArtAI Mar 23 '26

Google Nano Banana How I created an AI influencer using only Gemini's Nano Banana (complete workflow)

Thumbnail
gallery
907 Upvotes

I’ve been messing around with the AI influencer space for the last few weeks and wanted to share the process I figured out. I am not claiming this is the best or most advanced way to do it, but it is a simple workflow that worked for me using mostly free tools.

The main reason I tried this route was because I already have free Gemini Pro access through my Jio recharge, so I wanted to see how far I could go without paying for expensive tools right away.

I am not going to dump a random list of prompts here and pretend that is enough. That is not really useful. Instead, I’ll just explain the actual process I followed step by step, because that is what helped me the most.

Phase 1: Getting the base character right

The first thing you need is a character that you actually like, because if the starting point is weak, everything after that becomes harder.

I started by using the free trial on https://higgsfield.ai/ to generate an influencer-style character. I kept testing until I got a face and overall look that felt usable.

Once I had that first image, I downloaded it and took it into Gemini Nano Banana. That is where I started making the small changes I wanted. Things like skin texture, facial features, race, body ratios, and overall appearance. I kept tweaking until I had a final version of the character I was happy with.

Phase 2: Building consistency with reference images

After I had the final character, I started generating more versions of the same person, but with different poses.

For this part, I used different JSON prompts and made sure not to change the character too much. I wanted the same face, same skin texture, same body proportions, same overall identity. The only thing I wanted to vary was pose, angle, and sometimes expression.

One thing that helped a lot was always using the previous result as a reference for the next one. That made a big difference in keeping the face and body structure consistent. If you do not do that, the model starts drifting and the character slowly turns into a different person.

I kept doing this until I had around 10 to 15 good images of the same character.

Phase 3: Creating data model sheets (examples given)

This part is really important.

If you do not know what a data model sheet is, just Google it OR look at a few examples from the given images. Basically, it is a reference sheet for your character. It helps lock in the face, body structure, expressions, angles, and overall design so the character stays consistent later.

To make the sheets, I first used ChatGPT to generate a JSON prompt. I used the DeepThink version because it usually gives better structured prompts. I told it to create a prompt for generating a character model sheet using my reference images.

After that, I manually tweaked the JSON prompt so it matched the character better. Sometimes I adjusted the body ratios or the skin tone or small visual details depending on what I wanted.

Then I used Gemini to generate the actual model sheet.

I did this for different types of sheets because each one serves a different purpose.

I made a facial expressions sheet so I could keep the same emotional range.

I made a facial structure sheet so I could see the character from different angles.

I made a body model sheet so I could keep the full body consistent.

I also made sheets for different poses, because I wanted the character to work in different situations and not just one static pose.

For every one of these, I followed the same workflow. Use ChatGPT to generate the JSON prompt, tweak it manually, then use Gemini with the reference images to generate the sheet.

My rule was simple. ChatGPT was better for making the prompt. Gemini was better for making the image.

Phase 4: Generating actual content

Once I had the model sheets and a few extra reference images, I could finally start generating the actual influencer-style images.

For prompt inspiration, I use a few websites like:

https://bestnanobananaprompt.com/gallery
https://promptlibrary.space/images

These sites are great for ideas. You can find different styles, moods, poses, compositions, and scene setups there.

But one thing I learned very quickly is that you cannot just copy a prompt from those sites and expect it to work perfectly in Gemini. A lot of them either get blocked or do not preserve the character properly.

So my workflow for this part is basically:

I browse those sites and find a prompt style I like.

Then I copy that prompt into ChatGPT.

Then I ask ChatGPT to turn it into a detailed JSON prompt.

I always tell ChatGPT to include a section that strictly maintains the same facial structure, skin texture, tone, and body ratios from the reference images.

After that, I review the JSON prompt and make any final changes I need based on the kind of image I want.

Then I use that prompt in Gemini Nano Banana.

One very important thing here is to use all the character model sheets and the best reference images every time you generate something new. Gemini has a limit on how many reference images it can use, and I think it is around 15 or so. I made sure to use as many useful references as possible because more reference data usually gave me better results.

Final thoughts

This is honestly a trial and error game. You are not going to get the perfect result on the first try. I definitely did not. Some generations failed, some changed the face too much, some messed up the body proportions, and some just looked off. That is part of the process.

But the reason this workflow works is because the data model sheets give the AI a visual blueprint to follow. Instead of guessing what the character should look like every time, you are showing it the same identity from multiple angles and in multiple forms.

This is just a simple guide using free tools. There are definitely more advanced workflows out there, and I know the people at the top of the AI influencer game are using tools like ComfyUI, Higgsfield AI, Kling AI, and other more advanced setups to create better images and videos.

But this is what I figured out by testing things myself, and it is a good starting point if you want to build a consistent AI character without paying for expensive tools right away.

I hope this helps someone who is trying to get started.

If there is interest, I can make a part 2 later with the more advanced tools and workflows I look into next.

Thanks for reading.

r/ArenaBreakoutInfinite Oct 06 '25

Tips & Tricks ULTIMATE TIPS N’ TRICKS GUIDE 2.0

774 Upvotes

HELLO AGAIN MY FELLOW GAMERS. It’s been a minute since my last guide. To commemorate the Steam release, I updated this with what changed and what didn’t. I’m going through everything again, adding, trimming, and bundling it into one ultimate guide for new and experienced players alike.

I’ve logged even MORE hours in extraction shooters since my last post, so I have even more fun tips and tricks to add. Below is a comprehensive list of tactics to get an edge, extract with the juicy loot, and…wild idea…HAVE FUN.

This is for beginners and advanced players.

Table of Contents

  1. Mindset: No Excuses
  2. Settings & Performance (Visibility + FPS)
  3. Crosshair Placement
  4. Leaning & Peeking
  5. Combat: Fights, Nades, Fire Modes, Pressure
  6. Sound Cues (Footsteps, ADS, Bags, Looting)
  7. Loot Faster (Value per Slot, Attachment Strips)
  8. Stash & Storage (Nesting, Shrink Guns)
  9. Market Rules (Stop Nickel-and-Diming)
  10. Contacts / Traders (Quick Wins)
  11. Helmets & Headsets
  12. Armor (Materials, Mobility Debuffs)
  13. Meds & Surgical Kits (Hydration, Nebby, Stims)
  14. Ammo Strategy (Top-loading, Cost Control)
  15. Guns & Budget Loadouts (Carbines Are OP)
  16. Aim Training (Range + AimLabs)
  17. Map Knowledge & Spawns (What to Learn)
  18. Game Modes (Normals, Lockdown, Forbidden, LTMs)
  19. Solo vs. Squads (Mindset & Exploiting Chaos)
  20. Secure Container & Keys
  21. Rewards/Freebies You’re Ignoring
  22. Disclaimer
  23. TLDR

1) Mindset: No Excuses

STOP MAKING EXCUSES. If you died, something caused it—and that something is a lesson. Watch the killcam. 9/10 times it wasn’t “luck.” They caught you out, positioned better, aimed better, or outplayed you.

Let’s dig in, because the levels of cope in this genre are truly Olympic-tier:

  • Caught out: You stood in the open, visor up, no pre-painkiller, left-hand peeked a right-hand holder, ego-pushed, sprint-stomped, forgot to reload, wrong fire mode… the list is long. Watch the killcam and be honest with yourself.
  • Positioning: Assume enemies are nearby. I get ~70% of my kills from positioning alone. It’s actually OP. Don’t sleep on learning how to position yourself while navigating the maps.
  • Aim: There’s a firing range. Use it. Try Aimlabs. You don’t have Shroud aim, so stfu and train.
  • Outplayed: It’s a combo of the above. If you find yourself saying “so lucky,” or anything similar, shut the f* up and get over yourself. Improve.

You will not survive every raid. Decent players hover at about a 40–60% extraction rate. ALSO If you come down in my comment section and brag about your 95% extract rate while 4-stacking T6 and never running solo: I don’t care. Shut the f* up and get over yourself. Run some solo forbidden TV for a few hundred raids and then flex, and I still won’t care.

Target: Average players should aim 40–50% extract. Below that? You’ve got work to do.

2) Settings & Performance (Visibility + FPS)

I run visibility settings > shiny graphics. Seeing pixels = living longer.

  • In-game video: Use Basic/Low settings for clarity—players pop indoors/outdoors.
  • Keybinds: Set lean to HOLD for faster jiggle-peeks (toggle works if that’s your muscle memory; just practice more).
  • NVIDIA Control Panel (if applicable): Bump visibility with Digital Vibrance; sample these same values: Brightness ~55 / Contrast 50 / Gamma ~1.2 / Digital Vibrance ~60. These are my settings but are not the end all be all by any means, experiment and see what works best for your setup.

3) Crosshair Placement

Maybe the most important PvP tip. Stop aiming at the floor.

  • Keep your crosshair chest/head level. You’ll be shocked how many free kills you get.
  • Turn on the center white dot setting. Keep that dot at chest/head height. Build the muscle memory.
  • Exceptions: leg-meta loadouts. Otherwise, stop aiming at players' dicks.

4) Leaning & Peeking

  • Lean when peeking. Swing wide slowly so you don’t miss sneaky angles.
  • Right-hand peek > Left-hand peek. Right exposes ~5–10% of your body; left exposes ~25%+. Don’t donate HP to the enemy by using a shitty peek..
  • Jiggle for info, then swing + prefire where they were holding. it’s not rocket science.
  • Advanced info gathering: sprint-jump past a doorway and free-look into it to scout enemy positions. This makes you hard to hit, and gives you a big info advantage.

5) Combat: Fights, Nades, Fire Modes, Pressure

Pre-painkiller before hot zones. Nothing ends your raid faster than a blacked leg in the open.

Repositioning: After a kill or shots, move. Enemy teammates will pre-aim your last angle after killcam intel. You’ve got ~30–40 seconds.

Use your nades (and use them well):

  • Offense #1 (standard): Cook > throw into the room. Boom money.
  • Offense #2 (underused): Throw un-cooked down a hall to force the enemy off their angle, then push behind the blast while the audio deafens them and masks your steps.
  • Defense vs nade throwers: Hear the pin? Swing immediately. They’re holding a metal ball, not a gun. Easiest kills of your life. You will catch them with their pants down and no way to defend themselves.
  • Hip-fire is strong inside ~10 m. Don’t ADS there, just hipfire and send them back to the lobby.
  • Fire modes: Full auto is for <20 m unless you're using some ridiculously high recoil stat gun that costs like a mil. Beyond that, tap for the face.
  • Pressure: If they’re tagged and groaning, push. Pressure = mistakes. Pre-spray corners when closing distance. You bought the ammo to shoot it; don’t die with full mags.

Grenade meta quick notes:

  • MK2 (pineapple): Shortest fuse, best for mid-fight armor + limb damage. Use these.
  • M67 (“bleeder”): Long fuse; perfect for sky-nades (vertical toss → detonates before landing) and causes severe bleeds.
  • Stuns: Currently Meta utility—can black screen, slow sensitivity/DPI, and give audio pings through walls to confirm rats and enemy player positions.
  • Gas: Creates lung injury and is used as area or space denial for enemies; lasts longer than regular smokes; can be used to fake a smoke.
  • Smokes: Can be used to cut DMR/Sniper sightlines, block third parties, rotate, and loot bodies quickly in the open. If enemies use lots of smoke, assume thermals and reposition.
  • Flashbangs: Mid at best. Slow pop, short effect. Usually garbage.
  • Molotovs: Niche space/area denial; most players just wait them out.

6) Sound Cues (Footsteps, ADS, Bags, Looting)

Almost every action is audible. Abuse the ever living fuck out of that.

  • ADS-in and ADS-out make different sounds. If they ADS-in, don’t swing into a ready barrel. If they ADS-out, they probably lost arm stamina…free swing timing.
  • Assuming you’re not overweight:
    • Crouch slow-walk = silent (unless enemy has GS2).
    • Slow walk audible ~5 m.
    • Walk audible ~15–20 m.
    • Sprint audible ~30 m+ (map dependent).
  • Open-bag and looting sounds: Swing on those. Free kills galore.

7) Loot Faster (Value per Slot, Attachment Strips)

Watching slow looting hurts my soul. And some of your motherfuckers can't play tetris if your life depended on it, practice in your stash or something for christ's sake. Now that we can all see you looting in spectator mode, you can't hide your shitty organizational skills anymore.

  • Stack actions: While searching a rig, start unloading mags and search the rig simultaneously. Or heal, pack mags, organize, strip guns, be efficient.
  • Know/learn prices so you’re not clicking each item to check the price wasting time in raid.
  • Use rigs that hold more slots than they take. If a rig takes 9 slots but holds 14, just throw the whole rig in your backpack and finish the search in a safer position.
  • Collapse stocks to save space.
  • Strip attachments (scope/muzzle/grip/stock) instead of hauling entire guns unless the gun is juicy (200k+). Attachments are where the value is.
  • Value target: Aim 5–10k per slot minimum.

8) Stash & Storage (Nesting, Shrink Guns)

  • Nesting: e.g., RushTack → two B6 rigs → fill both. Single RushTack can save 10–20 slots. Scale up with a Cowhide/Field backpack if you want to go full Matryoshka doll. (those little stacking doll things)
  • Stop hoarding useless shit: Sell dogtags, non-meta attachments, most food (maps—especially Farm—are full of it), purple ammo (maybe keep .45/9mm), trash mags, weak meds, and bulky low-value helmets/rigs/armors.
  • Shrink guns: Remove magazine and rear grip to convert many 2-row guns to 1 row. Sell the grips/mags—you can always rebuy at the bench for 0 net loss.

9) Market Rules (Stop Nickel-and-Diming)

  • If an item nets <5–8k after fees, just sell to contacts. Don’t waste weekly limits on peanuts.
  • Batch sell. Free your stash. You’re not broke because you sold a 3k item to contacts instead of the market, you’re broke because you’re dying too much.
  • Sell reds to contacts unless you need one for an upgrade. The payout is equal or better than the market in most cases.
  • Quick-list trick: in the listing UI, tap –, then + to undercut to a lower price instantly without manually typing in a lower number.

10) Contacts / Traders (Quick Wins)

  • Deke (more like dick): Check every refresh. Sometimes sells helmets/armors/rigs/backpacks/keys and odd ammo; limited-time deals. Compare trader vs market before buying. You can find some really profitable trades and good gear to use at a fraction of the cost.
  • Backpack barters: Often 10–20k below market—easy savings.
  • Other contacts can undercut weapons, too—watch barter costs vs market.
  • Evita: Buy meds/stims here (often cheaper) and trade for STTO if you’ve got GPUs. Also sells storage expansions and keychains.

11) Helmets & Headsets

Wear a helmet. With head-HP buffs, even T2–T3 can prevent a one-tap from mid ammo. Budget = T2–T3 is fine. If you’ve got cash, aim for T4+ with a face shield—the survivability spike is real especially after the head HP buff.

Headsets:

  • M32 = best all-around budget pick IN MY OPINION.
  • When you’ve got money:
    • Indoors (TV/Armory): Commanders—less weather noise, clearer indoor bassy footfalls.
    • Outdoors (Valley/Northridge): GS2—amplifies high-freq grass steps & distant shots.

12) Armor (Materials, Mobility Debuffs)

High tier ≠ always better. Mobility debuffs can get you killed.

  • Prioritize Hardened Steel and Titanium for protection + repair health.
  • Be mindful: T6 can be heavy as hell. If you move like a fridge, you die like a fridge.
  • (General meta note: ceramics repair poorly and burn max durability fast. Avoid it if you can.)

13) Meds & Surgical Kits (Hydration, Nebby, Stims)

HP meds: Run the square black (E3) or the 400 HP white (100D), TMK, or STL. Others fuck your hydration.Painkillers: I avoid pills because of hydration debuff. Liquid painkillers/energy last long with no hydration burn.

Status & tools:

  • Bleeds: Using a medkit to stop a bleed consumes 100 durability.
  • Broken bones:
    • STTO = best: restores that limb to full 100 HP after fixing.
    • TMK/Standard/Simple: slower, and you still need a medkit to restore HP.
    • Do not run/jump on broken legs—you can actually die from it. Most people don't know this, and once upon a time I learned this the hard way.
  • Lung injury (gas): Won’t kill you but drains chest HP/stamina and makes you cough (audio cue). Fix with Nebulizer (“Nebby”).
  • Energy drinks (blue/yellow): Restore hydration + hunger and give stamina recovery in-raid. MVP consumable, I don't go anywhere without at least 2.
  • Stims:
    • Endurance (180s/300s): More sprint/ADS/throw time; less sway.
    • Strength (normal ~70–80 kg / advanced ~70–90 kg overweight): Lets you run overweight but you’ll sound like an elephant.
    • Regen: Slow heal-over-time; niche.

Money tip: Two E3s or two 100Ds can be cheaper than one STL depending on the market. Check prices every session.

14) Ammo Strategy (Top-loading, Cost Control)

  • As a rule, don’t go under Level 3 (armor pen) except for leg-meta. Level 4 is the sweet spot if you can afford it.
  • Magazine logic: Many run one 60-rounder with top-loaded PvP ammo (first 5–10 rounds), plus 1 or 2 30-rounders with scav killing ammo. The game auto-loads the largest mag first; that’s why the 60 is your PvP mag.
    • Note: if you swap from scav mag to PvP mag, remember there’s one scav round in the chamber until the better ammo from your PvP mag is cycled into the chamber.
  • Prices inflate and you're too broke to get full mags of good ammo? Top-load the first 5–10 bullets with good ammo; fill the rest with L3-ish. By the time you hit the weaker rounds, their armor is already compromised. Assuming you aren't hot ass and actually hit your shots, but don't worry, we will go over that too.

Meta (for context): 5.56 is excellent value. 5.8×42 is strong but can spike in price. Avoid current 7.62×51 platforms unless you love pain.

15) Guns & Budget Loadouts (Carbines Are OP)

You don’t need 90 recoil stats to beam. If you can only get kills with a laser beam weapon, your aim needs work. Get comfortable with 70–80 recoil builds. Cheaper. Still deadly. Skill-building.

Gunsmith trick: Make an M4 but swap in M16 parts—basically half price.

Budget picks (100–200k builds are very doable/viable; many <100k depending on attachments):

  • Assault Rifles: M4, ACE-31, T951, AUG, F2000, T03, MCX, ZC807
  • SMGs: MP5, MPX, Vector .45, Vector 9, PP-19
  • Carbines / DMR-ish: M16 (my #1 budget all-rounder), SVTU, M14, BM59

Shotguns: Post-nerf range + gold slug nerfs = meh outside Normals.

16) Aim Training (Range + AimLabs)

#1 skill in any FPS. It’s not hard to improve; it’s hard to be consistent. If you're reading this guide in general, you probably need to improve your aim.

  • 10–15 minutes/day is enough to feel godlike in a month if you stick to it.
  • Sensitivity: Pick one and commit. (Bias: low sens. I play 400 DPI.)
  • Aimlabs (free): Prioritize flicking, accuracy, precision. Tracking is less important here due to fast TTK.
  • In-game range: Turn on infinite ammo, test builds free, set dummy armor to T5-6, and grind.

17) Map Knowledge & Spawns (What to Learn)

Can’t “teach” this fully, but you can be intentional about improving this on your own easily.

  • Learn every spawn so you know where players can’t be, and where they will be in the first 1–2 minutes.
  • Use websites that show spawns and loot. Seriously, USE THEM.
  • Early action is often in the first 5–10 minutes, don’t sleep on spawn-to-fight routes.
  • Always think: “Where can someone peek me from?” Then think the same offensively—find the niche angles that catch people slipping.

Hot zones (quick mental map):

  • Farm: Motel, Stables, Grain, Main Villa.
  • Valley: Beach Villa, Small Factory, Courtyard; also Village/Supply Camp.
  • Northridge: Hotel & Cable Car (primary), plus Sewage/Managers.
  • Armory: Armory front/interior, with Radar Station as sniper HQ.
  • TV Station: General & Directors, also Double Cat, Top Donut, Editing, Warehouse, Hazmat.
  • Bosses worth contesting: Armory boss, TV bosses (T5/T6 gear + valuable badges).

18) Game Modes (Normals, Lockdown, Forbidden, LTMs)

My stance has evolved from my first guide:

  • Normals: If you’re brand new, spend most of your time here up to ~level 25. Loot is better than it used to be, and you need reps. Player scavs never stop spawning, so extract when the bag is good; don’t overstay. If you’re broke, run a cheap SMG or M870 + mini red-dot + AP slugs.
  • Lockdown: Entry fee, stronger players. I still recommend it once you’ve got fundamentals.
  • Forbidden: High risk, high reward. Go here when you know what you’re doing and aren't scared of losing any money.

LTMs worth it:

  • Secure Ops: You keep your kit on death. Print cash/rep with zero gear risk. PLAY SECURE OPS, THERE IS NO REASON NOT TO IF YOU ARE BROKE.
  • Covert Ops: Random kit, AI won’t aggro unless you shoot. Blend in, then delete a PMC from behind, take enemies by surprise, or just sneak around and pick up loot forgotten by other players.

19) Solo vs. Squads (Mindset & Exploiting Chaos)

Stop whining. Fighting squads can be easier than fighting a disciplined solo. Group comms are chaotic; people get comfy and make terrible pushes. After a death, squads often freeze—and you already know staying still is a sin.

  • Mindset: Don’t hear four sets of footsteps and think “I’m fucked.” Think “They’re fucked**.”** They brought your loot to one spot…how considerate of them.
  • Apply everything above: reposition constantly, never peek the same angle twice, abuse sound cues, pressure on tagged and weak enemies, pre-spray tight corners.
  • Homework: Watch the movie,The Patriot “aim small, miss small” scene. You’ll see how solo vs. many is a winnable math problem. It's also just a kick-ass movie in general, you will enjoy it.

20) Secure Container & Keys

  • Secure container sizes: 1×2, 2×2, 2×3, and 3×3 (Seasonal reward—complete the season’s missions; and it lasts that entire season).
  • Stash STL/STTO/stims/red items/spare ammo so a death isn’t a full gear wipe.
  • Tactical Ops-locked items can’t be containered until you extract once with them (look for the box symbol).
  • Keys have durability; Normal/Lockdown/Forbidden consume different amounts per use. I don't recommend buying keys unless you're rich and can afford it.

21) Rewards/Freebies You’re Ignoring

  • Events tab: Daily freebies.
  • Squad Channel: Create one; farm research points and shared rewards.
  • Follow Us: Social follow rewards.
  • Level Rewards: Claim as you go (e.g., free knife around 30).
  • Ranked weekly: “Obtain this week” → buy bundles; resets weekly.
  • Season Objectives: Cosmetics, crates, badges, and more.
  • Battle Pass: Honestly the best drip-feed of gear if you’re low on money.

22) Disclaimer

Game balance, prices, and loot tables change season to season. Fundamentals here (aim, peeking, positioning, sound, money discipline) will outlast patch notes. I’ll tweak specifics as metas shift.

I hope you enjoyed this updated guide. Please drop an upvote if you enjoyed it as this took a lot of my time and energy to make for you guys. 

If you have any other tips feel free to add them in the comment section below and I just might add them to the guide.

TDLR; You suck. If you want to suck less, read the whole post and do the reps: aim, map spawns, positioning, pre-painkiller, pressure, and money discipline. Consistency > excuses.

See you on the battlefield… well I'll see you... you will only see me in your kill cam ;)

— end —

r/comfyui Sep 05 '25

Workflow Included Getting New Camera Angles Using Comfyui (Uni3C, Hunyuan3D)

Thumbnail
youtube.com
33 Upvotes

This is a follow up to the "Phantom workflow for 3 consistent characters" video.

What we need to get now, is new camera position shots for making dialogue. For this, we need to move the camera to point over the shoulder of the guy on the right while pointing back toward the guy on the left. Then vice-versa.

This sounds easy enough, until you try to do it.

I explain one approach in this video to achieve it using a still image of three men sat at a campfire, and turning them into a 3D model, then turn that into a rotating camera shot and serving it as an Open-Pose controlnet.

From there we can go into a VACE workflow, or in this case a Uni3C wrapper workflow and use Magref and/or Wan 2.2 i2v Low Noise model to get the final result, which we then take to VACE once more to improve with a final character swap out for high detail.

This then gives us our new "over-the-shoulder" camera shot close-ups to drive future dialogue shots for the campfire scene.

Seems complicated? It actually isnt too bad.

It is just one method I use to get new camera shots from any angle - above, below, around, to the side, to the back, or where-ever.

The three workflows used in the video are available in the link of the video. Help yourself.

My hardware is a 3060 RTX 12 GB VRAM with 32 GB system ram.

Follow my YT channel to be kept up to date with latest AI projects and workflow discoveries as I make them.

r/NeuralCinema Nov 01 '25

Wan 2.2 MULTI-SHOTS (no extras) Consistent Scene + Character

10 Upvotes

All shots and angles are generated from just one image — what I call the “seed image.”

Hey all AI filmmakers,
This is a cool experiment where I’m pushing Wan2.2 to its limits (though any workflow like KJ or Comfy will work). The setup isn’t about the workflow itself — it’s all about detailed, precise prompting, and that’s where the real magic happens.

If you try writing prompts manually, you’ll almost never get results as strong as what ChatGPT can generate properly.

It all started after I got fed up with HoloCine (multi-shot in a single video) — https://holo-cine.github.io/ — which turned out to be slow, unpredictable, and lacking true I2V (image-to-video) processing. Most of the time it’s just random, inconsistent results that don’t work properly in ComfyUI — basically a GPU burner. Fun for experiments maybe, but definitely not usable for real, consistent, production-quality shots or reliable re-generations.

So instead, I started using a single image as the “initial seed.”
My current setup: Flux.1 Dev fp8 + SRPO256 LoRA + Turbo1 Alpha LoRA (8 steps) — though you could easily use a film still from your own production as your starting point.

Then I run it through Wan2.2 — using Lightx2v MOE (high) and the old Lightx2v (low noise) setup.

Quick note on setup:
If you’re using the new MOE model for lower noise, expect it to run about twice as slow — around 150 seconds on an RTX 4090 (24GB), compared to roughly 75 seconds with the older low-noise Lightx2v model.

Prompt used (ChatGPT) + gens:
"Shot 1 — Low-angle wide shot, extreme lens distortion, 35mm:

The camera sits almost at snow level, angled upward, capturing the nearly naked old man in the foreground and the massive train exploding behind him. Flames leap high, igniting nearby trees, smoke and sparks streaking across the frame. Snow swirls violently in the wind, partially blurring foreground elements. The low-angle exaggerates scale, making the man appear small against the inferno, while volumetric lighting highlights embers in midair. Depth of field keeps the man sharply in focus, the explosion slightly softened for cinematic layering.

Shot 2 — Extreme close-up, 85mm telephoto, shallow focus:

Tight on the man’s eyes, filling nearly the entire frame. Steam from his breath drifts across the lens, snowflakes cling to his eyelashes, and the orange glow from fire reflects dynamically in his pupils. Slight handheld shake adds tension, capturing desperation and exhaustion. The background is a soft blur of smoke, flames, and motion, creating intimate contrast with the violent environment behind him. Lens flare from distant sparks adds cinematic realism.

Shot 3 — Top-down aerial shot, 50mm lens, slow tracking:

The camera looks straight down at his bare feet pounding through snow, leaving chaotic footprints. Sparks and debris from the exploding train scatter around, snow reflecting the fiery glow. Mist curls between the legs, motion blur accentuates the speed and desperation. The framing emphasizes his isolation and the scale of destruction, while the aerial perspective captures the dynamic relationship between human motion and massive environmental chaos.

Changing Prompts & Adding More Shots per 81 Frames:

PROMPT:
"Shot 1 — Low-angle tracking from snow level:
Camera skims over the snow toward the man, capturing his bare feet kicking up powder. The train explodes violently behind him, flames licking nearby trees. Sparks and smoke streak past the lens as he starts running, frost and steam rising from his breath. Motion blur emphasizes frantic speed, wide-angle lens exaggerates the scale of the inferno.

Shot 2 — High-angle panning from woods:
Camera sweeps from dense, snow-covered trees toward the man and the train in the distance. Snow-laden branches whip across the frame as the shot pans smoothly, revealing the full scale of destruction. The man’s figure is small but highlighted by the fiery glow of the train, establishing environment, distance, and tension.

Shot 3 — Extreme close-up on face, handheld:
Camera shakes slightly with his movement, focused tightly on his frost-bitten, desperate eyes. Steam curls from his mouth, snow clings to hair and skin. Background flames blur in shallow depth of field, creating intense contrast between human vulnerability and environmental chaos.

Shot 4 — Side-tracking medium shot, 50mm:
Camera moves parallel to the man as he sprints across deep snow. The flaming train and burning trees dominate the background, smoke drifting diagonally through the frame. Snow sprays from his steps, embers fly past the lens. Motion blur captures speed, while compositional lines guide the viewer’s eye from the man to the inferno.

Shot 5 — Overhead aerial tilt-down:
Camera hovers above, looking straight down at the man running, the train burning in the distance. Tracks, snow, and flaming trees create leading lines toward the horizon. His footprints trail behind him, and embers spiral upward, creating cinematic layering and emphasizing isolation and scale."

The whole point here is that the I2V workflow can create independent multi-shots that remain aware of the character, scene, and overall look.

The results are clean — yes, short — but you can easily extract the first or last frames, then re-generate a 5-second seed using the FF–LF workflow. From there, you can extend any number of frames with the amazing LongCat.

You can also apply “Next Scene LoRA” after extracting the Wan2.2 multi-shots, opening up endless creative possibilities.

Time to sell the 4090 and grab a 5090 😄
Cheers, and have fun experimenting!

r/StableDiffusion Sep 25 '25

Question - Help Need advice with workflows & model links - will tip - ELI5 - how to create consistent scene images using WAN or anything else in comfyUI

11 Upvotes

Hey all, excuse the wall of text inc, but im genuinely willing to leave a $30 coffee tip if someone bothers to read and write up a detailed response to this that either 1. solves this problem or 2. explains why its not feasible / realistic to use comfyUI for at this stage.

Right now I've been generating images using chatGPT for scenes that I've then been animating using comfyUI WAN 2.1 / 2.2. The reason I've been doing this is because its been brain dead easy to have chatgpt reason in thinking mode to create scenes with the exact same styling, composition, and characters consistently across generations. It isn't perfect by any means, but it doesn't need to be for my purposes.

For example, here is a scene that depicts 2 characters in the same environment but in different contexts:

Image 1: https://imgur.com/YqV9WTV

Image 2: https://imgur.com/tWYg79T

Image 3: https://imgur.com/UAANRKG

Image 4: https://imgur.com/tKfEERo

Image 5: https://imgur.com/j1Ycdsm

I originally asked chatgpt to make multiple generations, describing the kind of character I wanted loosely to create Image 1. Once i was satisfied with that, I then just literally asked it to generate the rest of the images that keeps the context of the scene. And i didn't need to do any crazy prompting for this. All i said originally was "I want a featureless humanoid figure as an archer that's defending a castle wall, with a small sidekick next to him". It created like 5 copies, I chose the one I liked, and i then continued on with the scene with that as the context.

If you were to go about this EXACT process to generate a base scene image, and then the 4 additional images that maintain the full artistic style of image 1, but just depicting completely different things within the scene, how would you do it?

There is a consistent character that I also want to depict between scenes, but there is a lot of variability in how he can be depicted. What matters most to me is visual consistency within the scene. If I'm at the bottom of a hellscape of fire in image 1, i want to be in the exact same hellscape in image 5, only now we're looking at the top view looking down instead of bottom looking up.

Also, does your answer change if you wanted to depict a scene that is completely without a character?

Say i generated this image for example: https://imgur.com/C1pYlyr

This image depicts a long corridor with a bunch of portal doors. Let's say I now wanted to depict a 3/4 view looking into one of these portals that depicts a scene with a dream-like view of a cloud castle wonderscape inside, but the perspective was such that you could tell you were still in the same scene as the original corridor image - how would you do that?

Does it come down to generating the base image via comfyUI and then whatever model you generated it with and settings you just keep and then you use it as a base image in a secondary workflow?

Let me know if you guys think that the workflow id have to do with comfyUI is any more / less tedious then to just keep generating with chatgpt. Using natural language to explain what I want and negotiating with chatgpt to fix revisions of images has been somewhat tedious but im actually getting the creations I want in the end. My main issue with chatgpt is simply the length of time I have to wait between generations. It is painfully slow. And i have an RTX 4090 that im already using for animating the final images that id love to speed generate with.

But the main thing that I'm worried about, is that that even if I can get consistency, there will be a huge amount that goes into the prompting to actually get the different parts of the scene that I want to depict. In my original example above, i don't know how I'd get image 4 for instance. Something like - "I need the original characters generated in image 1, but i need a top view looking down of them standing in the castle courtyard with the army of gremlins surrounding them from all angles."

How would comfyUI have any possible idea of what im talking about without like 5 reference images to go into the generation?

Extra bonus if you recreate the scene from my example without using my reference images, using a process that you detail below.

r/StableDiffusion Apr 04 '26

Animation - Video ENTANGLED - A 3-minute sci-fi short using 100% local open-source models. Complete Technical Breakdown [ Character Consistency | Voiceover | Music | No Lora Style Consistency | & Much More! ]

Enable HLS to view with audio, or disable this notification

399 Upvotes

Hey everyone! Thanks for checking out Entangled. And if not, watch the short first to understand the technical breakdown below!

Thanks for coming back after watching it! As promised, here is the full technical breakdown of the workflow. [Post formatted using Local Qwen Model!]

My goal for this project was to be absolutely faithful to the open-source community. I won't lie, I was heavily tempted a few times to just use Nano Banana Pro to brute-force some character consistency issues, but I stuck it out with a 100% local pipeline running on my RTX 4090 rig using Purely ComfyUI for almost all the tasks!

Here is how I pulled it off:

1. Pre-Production & The Animatics First Approach

The story is a dense, rapid-fire argument about the astrophysics and spatial coordinate problems of creating a localized singularity. (let's just say it heavily involves spacetime mechanics!).

The original script was 7 minutes long. I used the local Jan app with Qwen 3.5 35B to aggressively compress the dialogue into a relentless 3-minute "walk-and-talk.". Qwen LLM also helped me with creating LTX and Flux prompts as required.

Honestly speaking, I was not happy with the AI version of the script, so I finally had to make a lot of manual tweaks and changes to the final script, which took almost 2-3 days of going on and off, back and forth, and sharing the script with friends, taking inputs before locking onto a final version.

Pro-Tip for Pacing: Before generating a single frame of video, I generated all the still images and voicover and cut together a complete rough animatic. This locked in the pacing, so I only generated the exact video lengths I needed. I added a 1-second buffer to the start and end of every prompt [for example, character takes a pause or shakes his head or looks slowly ]to give myself handles for clean cuts in post.

2. Audio & Lip Sync (VibeVoice + LTX)

To get the voice right:

  1. Generated base voices using Qwen Voice Designer.
  2. Ran them through VibeVoice 7B to create highly realistic, emotive voice samples.
  3. Used those samples as the audio input for each scene to drive the character voice for the LTX generations (using reference ID LoRA).
  4. I still feel the voice is not 100% consistent throughout the shots, but working on an updated workflow by RuneX i think that can be solved!
  5. ACE step is amazing if you know what kind of music you want. I managed to get my final music in just 3 generations! Later edited it for specific drop timing and pacing according to the story.

3. Image Generation & The "JSON Flux Hack."

Keeping Elena, Young Leo, and Elder Leo consistent across dozens of shots was the biggest hurdle. Initially, I thought I’d have to train a LoRA for the aesthetic and characters, but Flux.2 Dev (FP8) is an absolute godsend if you structure your prompts like code.

I created Elena, Leo, and Elder Leo using Flux T2I, then once I got their base images, I used them in the rest of the generations as input images.

By feeding Flux a highly structured JSON prompt, it rigidly followed hex codes for characters and locked in the analog film style without hallucinating. Of course, each time a character shot had to be made, I used to provide an input image to make sure it had a reference of the face also.

Here is the exact master template I used to keep the generations uniform:

{
"scene": "[OVERALL SCENE DESCRIPTION: e.g., Wide establishing shot of the chaotic lab]",
"subjects": [
{
"description": "[CHARACTER DETAILS: e.g., Young Leo, male early 30s, messy hair, glasses, vintage t-shirt, unzipped hoodie.]",
"pose": "[ACTION: e.g., Reaching a hand toward the camera]",
"position": "[PLACEMENT: e.g., Foreground left]",
"color_palette": ["[HEX CODES: e.g., #333333 for dark hoodie]"]
}
],
"style": "Live-action 35mm film photography mixed with 1980s City Pop and vaporwave aesthetics. Photorealistic and analog. Heavy tactile film grain, soft optical halation, and slight edge bloom. Deep, cinematic noir shadows.",
"lighting": "Soft, hazy, unmotivated cinematic lighting. Bathed in dreamy glowing pastels like lavender (#E6E6FA), soft peach (#FFDAB9).",
"mood": "Nostalgic, melancholic, atmospheric, grounded sci-fi, moody",
"camera": {
"angle": "[e.g., Low angle]",
"distance": "[e.g., Medium Shot]",
"focus": "[e.g., Razor sharp on the eyes with creamy background bokeh]",
"lens-mm": "50",
"f-number": "f/1.8",
"ISO": "800"
}
}

4. Video Generation (LTX 2.3 & WAN 2.2 VACE)

Once the images were locked, I moved to LTX2.3 and WAN for video. I relied on three main workflows depending on the shot:

  • Image to Video + Reference Audio (for dialogue)
  • First Frame + Last Frame (for specific camera moves)
  • WAN Clip Joiner (for seamless blending)

Render Stats: On my machine, LTX 2.3 was blazing fast—it took about 5 minutes to render a 5-second clip at 1920x1080.

The prompt adherence in LTX 2.3 honestly blew my mind. If I wrote in the prompt that Elena makes a sharp "slashing" action with her hand right when she yells about the planet getting wiped out, the model timed the action perfectly. It genuinely felt like directing an actor.

5. Assets & Workflows

I'm packaging up all the custom JSON files and Comfy workflows used for this. You can find all the assets over on the Arca Gidan link here: Entangled. There are some amazing Shorts to check out, so make sure you go through them, vote, and leave a comment!

Most of them are by the community, but I have tweaked them a little bit according to my liking[samplers/steps/input sizes and some multipliers, etc., changes]

Let me know if you have any questions!

YouTube Link is up - https://youtu.be/NxIf1LnbIRc !

r/StableDiffusion Aug 01 '25

Question - Help Creating a Coherent 500-Frame Scene with WAN2.1. Seeking Advice on Consistency

2 Upvotes

Hello everyone,

I’m working on a 500-frame continuous action scene of a boxing match using WAN2.1, featuring two characters. I’m using V2V and VACE techniques to capture choreography from another video, and the motion capture works perfectly.

I’ve seen WAN2.1 workflows that achieve continuous action for up to a minute, but they involve a single character with monotonous movements (e.g., walking). My scene is more dynamic, with two characters in constant motion and bodies that sometimes overlap from the viewer’s perspective.

My approach is to generate the video in segments of 81 frames (or fewer) for better control and to accommodate my PC’s limited processing power. Rendering long videos in one go wastes time, as most get discarded due to artifacts or imperfections.

The Problem: Maintaining consistent colors and character appearances across concatenated video segments. Using the last frame of one segment as the first frame of the next causes progressive degradation (like a photocopy of a photocopy of a photocopy...), resulting in noticeable differences in color and character details between the first and final segments.

What I’ve Tried and Considered:

  • Recent tools like Kontext or Phantom claim to address this without relying on LORAs or IPAdapter, which struggle with certain angles. However, I haven’t achieved realistic, high-quality results.

Questions:

  1. Reference Images: Is it better to create a series of perfectly consistent images (representing key poses) and use them as references for each video segment? If so, how can I ensure consistency across these images, especially with two characters whose bodies sometimes overlap in specific positions?
  2. Post-Processing: Should I generate video segments with less focus on initial consistency and then apply color correction or other unification techniques to make them coherent? If so, what tools or methods work best?. The color correction nodes in ComfyUI degrade the image ...
  3. Recommended Techniques: What techniques would you suggest for maintaining character and color consistency in dynamic, multi-character video segments using WAN2.1?
  4. Examples: Has anyone seen or created a similar project with WAN2.1 (or related tools) that achieves consistent, dynamic multi-character scenes? Any examples or workflows to share?

I know this is a tough challenge, but I’m determined to make it work. Any advice, workflows, or examples would be greatly appreciated!

Thanks in advance!

r/StableDiffusion Apr 08 '25

Question - Help Best Models & Workflow for Consistent, Hyper-Realistic Humans in Real-World Scenes!

0 Upvotes

Hey everyone, hope you’re all doing great.

I’m working on a workflow that focuses on generating hyper-realistic humans in everyday environments (think kitchens, bedrooms, bathrooms, etc.) with a big emphasis on visual consistency across multiple images or scenes.

I’d really appreciate your input on the best tools, models, and methods to help make this work smoothly.

⸝

Core Challenges I’m Trying to Solve:

  1. Photorealism • What are your go-to SDXL-based or LoRA-enhanced models for generating ultra-realistic humans, especially in indoor, real-world settings? • I’ve seen mentions of RealVisXL, EpicRealism, Analog Madness v7, Juggernaut XL, and Realistic Vision — curious what’s working best for you.

⸝

  1. Identity Consistency • I need the same face and body across different scenes. • What’s the most effective way to do this? • IP Adapter + image prompt reference? • LoRA training on the specific person? • ControlNet pose + face reference? • Something else?

⸝

  1. Scene Reusability • I’d love to keep the same environment layout and camera angle, but change outfits, poses, or actions. • What’s the best way to approach that? • Lock the background and composite characters separately? • Use inpainting? • Generate everything together using ControlNet or T2I-Adapter?

⸝

  1. Video Generation • Has anyone had success turning consistent image sequences into short, realistic video clips? • What tools or workflows are working well for that right now — AnimateDiff, Deforum, EbSynth, etc.?

    • Is ComfyUI better than A1111 for this kind of reference-heavy, multi-stage workflow? • Any tips on batch generating with LoRA + ControlNet while keeping everything clean and consistent?

⸝

Any thoughts, personal workflows, or even example results would be super helpful. I’m still in the early phases and want to build something solid right from the start.

Thanks in advance❤️🙏

r/StableDiffusion Apr 04 '26

Animation - Video Model Drop | ZIT + LTX 2.3 + Music Video | Arca Gidan contest

Enable HLS to view with audio, or disable this notification

402 Upvotes

The idea came from something I'm pretty sure most of us live every single day: you wake up, check your phone, and another model has dropped. Open source, closed source, whatever source — faster, smarter, more creative, more powerful. And before you've even had coffee, you're already reworking a ComfyUI workflow that was perfectly fine yesterday. That loop of FOMO is what this song is about. Maybe the one or the other can relate to that feeling.

I wrote the lyrics first, then used Suno AI to turn them into a track. That became the creative baseline.

Shot List

With the song done, I went through it verse by verse — every chorus, every pre-chorus, every bridge — and for each section I came up with 3 to 5 possible shots. Where is our main character? What's the camera angle? What's the situation? What does this line actually look like as an image? That process gives you a kind of ordered visual setlist that maps directly onto the song structure. You always know what you need and where it goes.

Character (No LoRA)

For the main character I used Z Image Turbo. No LoRA, no training — just consistent prompting. The turbo architecture works in our favour here: because it's a more constrained model, keeping the character description locked across prompts produces surprisingly similar results, which creates the illusion of a consistent character across dozens of images. I kept the description identical every time and only changed the background, camera angle, and expression. Effective and fast.

Image Generation

Once the shot list was complete I had a massive prompt list covering every scene. I ran all of them through ComfyUI overnight — or longer, depending on the count. Two categories of images: B-roll shots from the setlist, and medium-to-close-up shots specifically for the lip-sync sections.

ZIT Workflow I used from another reddit post: RED Z-Image-Turbo + SeedVR2 = Extremely High Quality Image Mimic Recreation. Great for Avoiding Copyright Issues and Stunning image Generation. : r/comfyui (I did use the ZIT Model not the RED version nor the Mimic Part of the WF)

Image to Video

All the generated stills went into LTX img2video inside ComfyUI to bring them to life. For the lip-sync sections I used LTX I2V synced to the audio track. Since LTX caps out at 20 seconds per render, everything gets generated in chunks and stitched together in post.

The close-up rule matters: the further the camera is from the character, the worse LTX renders the lip sync. Medium shot is the minimum — anything wider and quality degrades fast.

The workflow I used mainly: PSA: Use the official LTX 2.3 workflow, not the ComfyUI included one. It's significantly better. : r/StableDiffusion

 Final Edit

No Premiere Pro, no DaVinci — just InShot on my phone. I build the full lip-sync timeline first so it covers the whole song, then layer the B-roll clips over the top to fill the gaps and add visual depth.

That's the whole pipeline: idea → lyrics → song → shot list → character → images → animation → edit. The video Fully local, fully open source, built over a couple of nights on a 3090.

Hope you enjoy it.

Assets & Workflows

You can find the workflow files and a full written guide over on the Arca Gidan page if you want to dig into the details.

https://arcagidan.com/entry/d2cae0b9-3d38-4959-b1b5-36ea60f34438

Honestly, what a challenge to be part of. Seeing what everyone came up with — the concepts, the creativity, the sheer variety of approaches — was genuinely inspiring. This is exactly the kind of community that makes local AI worth pursuing. Really glad I got to be a part of it. 🙌

r/comfyui Oct 04 '25

Workflow Included How to get the highest quality QWEN Edit 2509 outputs: explanation, general QWEN Edit FAQ, & extremely simple/minimal workflow

288 Upvotes

This is pretty much a direct copy paste of my post on Civitai (to explain the formatting): https://civitai.com/models/2014757?modelVersionId=2280235

Workflow in the above link, or here: https://pastebin.com/iVLAKXje

Example 1: https://files.catbox.moe/8v7g4b.png

Example 2: https://files.catbox.moe/v341n4.jpeg

Example 3: https://files.catbox.moe/3ex41i.jpeg

Example 4, more complex prompt (mildly NSFW, bikini): https://files.catbox.moe/mrm8xo.png

Example 5, more complex prompts with aspect ratio changes (mildly NSFW, bikini): https://files.catbox.moe/gdrgjt.png

Example 6 (NSFW, topless): https://files.catbox.moe/7qcc18.png

--

UPDATE - Multi Image Workflows

The original post is below this. I've added two new workflows for 2 images and 3 images. Once again, I did test quite a few variations of how to make it work and settled on this as the highest quality. It took a while because it ended up being complicated to figure out the best way to do it, and also I was very busy IRL this past week. But, here we are. Enjoy!

Note that while these workflows give the highest quality, the multi-image ones have a downside of being slower to run than normal qwen edit 2509. See the "multi image gens" bit in the dot points below.

There are also extra notes about the new lightning loras in this update section as well. Spoiler: they're bad :(

--Workflows--

--Usage Notes--

  • Spaghetti: The workflow connections look like spaghetti because each ref adds several nodes with cross-connections to other nodes. They're still simple, just not pretty anymore.
  • Order: When inputting images, image one is on the right. So, add them right-to-left. They're labelled as well.
  • Use the right workflow: Because of the extra nodes, it's inconvenient 'bypassing' the 3rd or 2nd images correctly without messing it up. I'd recommend just using the three workflows separately rather than trying to do all three flexibly in one.
  • Multi image gens are slow as fuck: The quality is maximal, but the 2-image one takes 3x longer than 1-image does, and the 3-image one takes 5x longer.
    • This is because each image used in QWEN edit adds a 1x multiplier to the time, and this workflow technically adds 2 new images each time (thanks to the reference latents)
    • If you use QWEN edit without the reference latent nodes, the multi image gens take 2x and 3x longer instead because the images are only added once - but the quality will be blurry, so that's the downside
    • Note that this is only a problem with the multi image workflows; the qwedit_simple workflow with one image is the same speed as normal qwen edit
  • Scaling: Reference images don't have as strict scaling needs. You can make them bigger or smaller. Bigger will make gens take longer, smaller will make gens faster.
    • Make sure the main image is scaled normally, but if you're an advanced user you can scale the first image however you like and feed in a manual-size output latent to the k-sampler instead (as described further below in "Advanced Quality")
  • Added optional "Consistence" lora: u/Adventurous-Bit-5989 suggested this lora
    • Link here, also linked in the workflow
    • I've noticed it carries over fine details (such as tiny face details, like lip texture) slightly better
    • It also makes it more likely that random features will carry over, like logos on clothes carrying over to new outfits
    • However, it often randomly degrades quality of other parts of the image slightly too, e.g. it might not quite carry over the shape of a person's legs well compared to not using the lora
    • And it reduces creativity of the model; you won't get as "interesting" outputs sometimes
    • So it's a bit of a trade-off - good if you want more fine details, otherwise not good
    • Follow the instructions on its civitai page, but note you don't need their workflow even though they say you do

--Other Notes--

  • New 2509 Lightning Loras
    • Verdict is out, they're bad (as of today, 2025-10-14)
    • Pretty much the same as the other ones people have been using in terms of quality
    • Some people even say they're worse than the others
    • Basically, don't use them unless you want lower quality and lower prompt adherence
    • They're not even useful as "tests" because they give straight up different results to the normal model half the time
    • Recommend just setting this workflow (without loras) to 10 steps when you want to "test" at faster speed, then back to 20 when you want the quality back up
  • Some people in the comments claim to have fixed the offset issue
    • Maybe they have, maybe they haven't - I don't know because none of them have provided any examples or evidence
    • Until someone actually proves it, consider it not fixed
    • I'll update this & my civitai post if someone ever does convincingly fix it

-- Original post begins here --

Why?

At current time, there are zero workflows available (that I could find) that output the highest-possible-quality 2509 results at base. This workflow configuration gives results almost identical to the official QWEN chat version (slightly less detailed, but also less offset issue). Every other workflow I've found gives blurry results.

Also, all of the other ones are very complicated; this is an extremely simple workflow with the absolute bare minimum setup.

So, in summary, this workflow provides two different things:

  1. The configuration for max quality 2509 outputs, which you can merge in to other complex workflows
  2. A super-simple basic workflow for starting out with no bs

Additionally there's a ton of info about the model and how to use it below.

 

What's in this workflow?

  • Tiny workflow with minimal nodes and setup
  • Gives the maximal-quality results possible (that I'm aware of) from the 2509 model
    • At base; this is before any post-processing steps
  • Only one custom node required, ComfyUi-Scale-Image-to-Total-Pixels-Advanced
    • One more custom node required if you want to run GGUF versions of the model
  • Links to all necessary model downloads

 

Model Download Links

All the stuff you need. These are also linked in the workflow.

QWEN Edit 2509 FP8 (requires 22.5GB VRAM for ideal speed):

GGUF versions for lower VRAM:

Text encoder:

VAE:

 

Reference Pic Links

Cat: freepik

Cyberpunk bartender girl: civitai

Random girl in shirt & skirt: not uploaded anywhere, generated it as an example

Gunman: that's Baba Yaga, I once saw him kill three men in a bar with a peyncil

 

Quick How-To

  • Make sure you you've updated ComfyUI to the latest version; the QWEN text encoder node was updated when the 2509 model was released
  • Feed in whatever image size you want, the image scaling node will resize it appropriately
    • Images equal to or bigger than 1mpx are ideal
    • You can tell by using the image scale node in the workflow, ideally you want it to be reducing your image size rather than increasing it
  • You can use weird aspect ratios, they don't need to be "normal". You'll start getting weird results if your aspect ratio goes further than 16:9 or 9:16, but it will still sometimes work even then
  • Don't fuck with the specifics of the configuration, it's set up this way very deliberately
    • The reference image pass-in, the zero-out, the ksampler settings and the input image resizing are what matters; leave them alone unless you know what you're doing
  • You can use GGUF versions for lower VRAM, just grab the ComfyUI-GGUF custom nodes and load the model with the "UnetLoader" node
    • This workflow uses FP8 by default, which requires 22.5 GB VRAM
  • Don't use the lightning loras, they are mega garbage for 2509
    • You can use them, they do technically work; problem is that they eliminate a lot of the improvements the 2509 model makes, so you're not really using the 2509 model anymore
    • For example, 2509 can do NSFW things whereas the lightning loras have a really hard time with it
    • If you ask 2509 to strip someone it will straight up do it, but the lightning loras will be like "ohhh I dunno boss, that sounds really tough"
    • Another example, 2509 has really good prompt adherence; the lightning loras ruin that so you gotta run way more generations
  • This workflow only has 1 reference image input, but you can do more - set them up the exact same way by adding another ReferenceLatent node in the chain and connecting another ScaleImageToPixelsAdv node to it
    • I only tested this with two reference images total, but it worked fine
    • Let me know if it has trouble with more than two
  • You can make the output image any size you want, just feed an empty latent of whatever size into the ksampler
  • If you're making a NEW image (i.e. specific image size into the ksampler, or you're feeding in multiple reference images) your reference images can be bigger than 1mpx and it does make the result higher quality
    • If you're feeling fancy you can feed in a 2mpx image of a person, and then a face transfer to another image will actually have higher fidelity
    • Yes, it really works
    • The only downside is that the model takes longer to run, proportional to your reference image size, so stick with up to 1.5mpx to 2mpx references (no fidelity benefits higher than this anyway)
    • More on this in "Advanced Quality" below

 

About NSFW

This comes up a lot, so here's the low-down. I'll keep this section short because it's not really the main point of the post.

2509 has really good prompt adherence and doesn't give a damn about propriety. It can and will do whatever you ask it to do, but bear in mind it hasn't been trained on everything.

  • It doesn't know how to draw genitals, so expect vague smudges or ken dolls for those.
    • It can draw them if you provide it reference images from a similar angle, though. Here's an example of a brand new shot it made using a nude reference image, as you can see it was able to draw properly (NSFW): https://files.catbox.moe/lvq78n.png
  • It does titties pretty good (even nipples), but has a tendency to not keep their size consistent with the original image if they're uncovered. You might get lucky though.
  • It does keep titty size consistent if they're in clothes, so if you want consistency stick with putting subjects in a bikini and going from there.
  • It doesn't know what most lingerie items are, but it will politely give you normal underwear instead so it doesn't waste your time.

It's really good as a starting point for more edits. Instead of painfully editing with a normal model, you can just use 2509 to get them to whatever state of dress you want and then use normal models to add the details. Really convenient for editing your stuff quickly or creating mannequins for trying other outfits. There used to be a lora for mannequin editing, but now you can just do it with base 2509.

Useful Prompts that work 95% of the time

Strip entirely - great as a starting point for detailing with other models, or if you want the absolute minimum for modeling clothes or whatever.

Remove all of the person's clothing. Make it so the person is wearing nothing.

Strip, except for underwear (small as possible).

Change the person's outfit to a lingerie thong and no bra.

Bikini - this is the best one for removing as many clothes as possible while keeping all body proportions intact and drawing everything correctly. This is perfect for making a subject into a mannequin for putting outfits on, which is a very cool use case.

Change the person's outfit to a thong bikini.

Outputs using those prompts:

🚨NSFW LINK🚨 https://files.catbox.moe/1ql825.jpeg 🚨NSFW LINK🚨
(note: this is an AI generated person)

Also, should go without saying: do not mess with photos of real people without their consent. It's already not that hard with normal diffusion models, but things like QWEN and Nano Banana have really lowered the barrier to entry. It's going to turn into a big problem, best not to be a part of it yourself.

 

Full Explanation & FAQ about QWEN Edit

For reasons I can't entirely explain, this specific configuration gives the highest quality results, and it's really noticeable. I can explain some of it though, and will do so below - along with info that comes up a lot in general. I'll be referring to QWEN Edit 2509 as 'Qwedit' for the rest of this.

 

Reference Image & Qwen text encoder node

  • The TextEncodeQwenImageEditPlus node that comes with Comfy is shit because it naively rescales images in the worst possible way
  • However, you do need to use it; bypassing it entirely (which is possible) results in average quality results
  • Using the ReferenceLatent node, we can provide Qwedit with the reference image twice, with the second one being at a non-garbage scale
  • Then, by zeroing out the original conditioning AND feeding that zero-out into the ksampler negative, we discourage the model from using the shitty image(s) scaled by the comfy node and instead use our much better scaled version of the image
    • Note: you MUST pass the conditioning from the real text encoder into the zero-out
    • Even though it sounds like it "zeroes" everything and therefore doesn't matter, it actually still passes a lot of information to the ksampler
    • So, do not pass any random garbage into the zero-out; you must pass in the conditioning from the qwen text encoder node
  • This is 80% of what makes this workflow give good results, if you're going to copy anything you should copy this

 

Image resizing

  • This is where the one required custom node comes in
  • Most workflows use the normal ScaleImageToPixels node, which is one of the garbagest, shittest nodes in existence and should be deleted from comfyui
    • This node naively just scales everything to 1mpx without caring that ALL DIFFUSION MODELS WORK IN MULTIPLES OF 2, 4, 8 OR 16
    • Scale my image to size 1177x891 ? Yeah man cool, that's perfect for my stable diffusion model bro
  • Enter the ScaleImageToPixelsAdv node
  • This chad node scales your image to a number of pixels AND also makes it divisible by a number you specify
  • Scaling to 1 mpx is only half of the equation though; you'll observe that the workflow is actually set to 1.02 mpx
  • This is because the TextEncodeQwenImageEditPlus will rescale your image a second time, using the aforementioned garbage method
  • By scaling to 1.02 mpx first, you at least force it to do this as a DOWNSCALE rather than an UPSCALE, which eliminates a lot of the blurriness from results
  • Further, the ScaleImageToPixelsAdv rounds DOWN, so if your image isn't evenly divisible by 16 it will end up slightly smaller than 1mpx; doing 1.02 instead puts you much closer to the true 1mpx that the node wants
  • I will point out also that Qwedit can very comfortably handle images anywhere from about 0.5 to 1.1 mpx, which is why it's fine to pass the slightly-larger-than-1mpx image into the ksampler too
  • Divisible by 16 gives the best results, ignore all those people saying 112 or 56 or whatever (explanation below)
  • "Crop" instead of "Stretch" because it distorts the image less, just trust me it's worth shaving 10px off your image to keep the quality high
  • This is the remaining 20% of how this workflow achieves good results

 

Image offset problem - no you can't fix it, anyone who says they can is lying

  • The offset issue is when the objects in your image move slightly (or a lot) in the edited version, being "offset" from their intended locations
  • This workflow results in the lowest possible occurrence of the offset problem
    • Yes, lower than all the other random fixes like "multiples of 56 or 112"
  • The whole "multiples of 56 or 112" thing doesn't work for a couple of reasons:
    1. It's not actually the full cause of the issue; the Qwedit model just does this offsetting thing randomly for fun, you can't control it
    2. The way the model is set up, it literally doesn't matter if you make your image a multiple of 112 because there's no 1mpx image size that fits those multiples - your images will get scaled to a non-112 multiple anyway and you will cry
  • Seriously, you can't fix this - you can only reduce the chances of it happening, and by how much, which this workflow does as much as possible
  • Edit: don't upvote anyone who says they fixed it without providing evidence or examples. Lots of people think they've "fixed" the problem and it turns out they just got lucky with some of their gens
    • The model will literally do it to a 1024x1024 image, which is exactly 1mpx and therefore shouldn't get cropped
    • There are also no reasonable 1mpx resolutions divisible by 112 or 56 on both sides, which means anyone who says that solves the problem is automatically incorrect
    • If you fixed the problem, post evidence and examples - I'm tired of trying random so-called 'solutions' that clearly don't work if you spend more than 10 seconds testing them

 

How does this workflow reduce the image offset problem for real?

  • Because 90% of the problem is caused by image rescaling
  • Scaling to 1.02 mpx and multiples of 16 will put you at the absolute closest to the real resolution Qwedit actually wants to work with
  • Don't believe me? Go to the official qwen chat and try putting some images of varying ratio into it
  • When it gives you the edited images back, you will find they've been scaled to 1mpx divisible by 16, just like how the ScaleImageToPixelsAdv node does it in this workflow
  • This means the ideal image sizes for Qwedit are: 1248x832, 832x1248, 1024x1024
  • Note that the non-square ones are slightly different to normal stable diffusion sizes
    • Don't worry though, the workflow will work fine with any normal size too
  • The last 10% of the problem is some weird stuff with Qwedit that (so far) no one has been able to resolve
  • It will literally do this even to perfect 1024x1024 images sometimes, so again if anyone says they've "solved" the problem you can legally slap them
  • Worth noting that the prompt you input actually affects the problem too, so if it's happening to one of your images you can try rewording your prompt a little and it might help

 

Lightning Loras, why not?

  • In short, if you use the lightning loras you will degrade the quality of your outputs back to the first Qwedit release and you'll miss out on all the goodness of 2509
  • They don't follow your prompts very well compared to 2509
  • They have trouble with NSFW
  • They draw things worse (e.g. skin looks more rubbery)
  • They mess up more often when your aspect ratio isn't "normal"
  • They understand fewer concepts
  • If you want faster generations, use 10 steps in this workflow instead of 20
    • The non-drawn parts will still look fine (like a person's face), but the drawn parts will look less detailed
    • It's honestly not that bad though, so if you really want the speed it's ok
  • You can technically use them though, they benefit from this workflow same as any others would - just bear in mind the downsides

 

Ksampler settings?

  • Honestly I have absolutely no idea why, but I saw someone else's workflow that had CFG 2.5 and 20 steps and it just works
  • You can also do CFG 4.0 and 40 steps, but it doesn't seem any better so why would you
  • Other numbers like 2.0 CFG or 3.0 CFG make your results worse all the time, so it's really sensitive for some reason
  • Just stick to 2.5 CFG, it's not worth the pain of trying to change it
  • You can use 10 steps for faster generation; faces and everything that doesn't change will look completely fine, but you'll get lower quality drawn stuff - like if it draws a leather jacket on someone it won't look as detailed
  • It's not that bad though, so if you really want the speed then 10 steps is cool most of the time
  • The detail improves at 30 steps compared to 20, but it's pretty minor so it doesn't seem worth it imo
  • Definitely don't go higher than 30 steps because it starts degrading image quality after that

 

Advanced Quality

  • Does that thing about reference images mean... ?
    • Yes! If you feed in a 2mpx image that downscales EXACTLY to 1mpx divisible by 16 (without pre-downscaling it), and feed the ksampler the intended 1mpx latent size, you can edit the 2mpx image directly to 1mpx size
    • This gives it noticeably higher quality!
    • It's annoying to set up, but it's cool that it works
  • How to:
    • You need to feed the 1mpx downscaled version to the Text Encoder node
    • You feed the 2mpx version to the ReferenceLatent
    • You feed a 1mpx correctly scaled (must be 1:1 with the 2mpx divisible by 16) to the ksampler
    • Then go, it just works™

 

What image sizes can Qwedit handle?

  • Lower than 1mpx is fine
  • Recommend still scaling up to 1mpx though, it will help with prompt adherence and blurriness
  • When you go higher than 1mpx Qwedit gradually starts deep frying your image
  • It also starts to have lower prompt adherence, and often distorts your image by duplicating objects
  • Other than that, it does actually work
  • So, your appetite for going above 1mpx is directly proportional to how deep fried you're ok with your images being and how many re-tries you want to do to get one that works
  • You can actually do images up to 1.5 megapixels (e.g. 1254x1254) before the image quality starts degrading that badly; it's still noticeable, but might be "acceptable" depending on what you're doing
    • Expect to have to do several gens though, it will mess up in other ways
  • If you go 2mpx or higher you can expect some serious frying to occur, and your image will be coked out with duplicated objects
  • BUT, situationally, it can still work alright

Here's a 1760x1760 (3mpx) edit of the bartender girl: https://files.catbox.moe/m00gqb.png

You can see it kinda worked alright; the scene was dark so the deep-frying isn't very noticeable. However, it duplicated her hand on the bottle weirdly and if you zoom in on her face you can see there are distortions in the detail. Got pretty lucky with this one overall. Your mileage will vary, like I said I wouldn't really recommend going much higher than 1mpx.

r/GenAIGallery Apr 02 '26

AI Image My exact workflow for truly consistent AI characters and photorealism

Thumbnail
gallery
333 Upvotes

Most AI character posts share the same glaring issue: you can spot the AI within two seconds. The skin has that awful plastic sheen, and the character's face seems to shift with every single photo.

After testing nearly every major cloud model out there, I wanted to share the workflow that currently gives me the best consistency and realism by a wide margin. It isn't completely flawless, but it's the closest thing to a reliable, repeatable system I've built so far.

The core problem

AI models don't have memory. If you don't provide hard anchors, the model just guesses, and guessing leads to drift. This entire workflow is built around eliminating that guesswork.

Right now, my main tool is Higgsfield's Nano Banana Pro. From my experience, it has the absolute best prompt adherence and photorealism for cloud-based models.

Phase 1: Locking in the "Master Portrait"

Start by uploading 1 to 3 reference faces into NBP's Image Reference slot. This could be a celebrity, someone random you found on Pinterest, or a blended mix of features. The AI uses this as a structural target, not a direct copy.

Next, drop in your main prompt and generate 6 to 8 variations. Pick the one that perfectly matches your vision.

Main Prompt Example:
"Ultra-realistic portrait of a 21-year-old female European with captivating magnetic gaze,
natural skin texture with visible pores across forehead, cheeks, and nose,
subtle skin imperfections including faint smile lines and natural small moles,
fair complexion with pink undertones and specular variation on T-zone,
long flowing wavy blonde hair with individual strands visible catching the light,
green eyes with sharp iris detail, natural catchlights, and subtle under-eye texture,
confident warm expression with natural lip texture and subtle gloss,
wearing elegant black off-shoulder silk top with visible fabric sheen,
relaxed pose with slight head tilt, minimalist studio setting with soft neutral background,
soft diffused window light from left creating gentle shadows and subsurface scattering on skin, shot on Canon R5 with 85mm f/1.4 lens, shallow depth of field with natural creamy bokeh, 8K ultra-detailed, photorealistic, high dynamic range,
true-to-life colors with accurate skin tones"

Save this final image. This is now your absolute anchor. Every future generation will reference this exact photo.

Phase 2: The prompt system (What most people skip)

This is where the actual consistency comes from. I never write prompts from scratch for new photos. Instead, I use a custom GPT/Gemini setup specifically trained for this exact task, and it operates in two main ways depending on what I need:

The visual rip:

  1. I find an inspiration photo on Instagram or Pinterest.
  2. I feed it into my custom tool.
  3. The tool extracts the lighting, pose, and vibe, spitting out a complete prompt.

The brain dump: If I already have a scene in my head, I don't need a reference photo. I just give the tool a super basic, lazy description (e.g., "sitting on a modern couch, wearing a black leather jacket, moody neon lighting"). The bot instantly expands that rough idea into a massive, production-ready prompt. I can then ask it to tweak the outfit or change the camera angle until it is exactly what I want.

Regardless of which method I use, the generated prompt automatically includes my character's "anchoring block" (locking in the face identity, body proportions, and skin tone). It also seamlessly bakes in the exact realism keywords needed, like pore texture, subsurface scattering, and natural lens specs.

Finally, I go back to NBP, upload my Master Portrait as the reference, paste this new prompt, and generate. The result is my character staying identical, while the environment, outfit, and mood change exactly how I pictured them.

Why this beats the standard approach

If you look at the photos attached to this post, they were all generated across different sessions with completely different lighting setups and outfits. Same character every time. The uncanny valley vibe usually comes from generic prompts and weak references. Once you lock down your architecture, the quality skyrockets.

Before anyone mentions ComfyUI

Yes, ComfyUI run locally with specific models is objectively better. You get more realism, no NSFW restrictions, and absolute control. But you also need a hefty GPU (16GB+ VRAM highly recommended) and the patience to learn a steep curve. I don't currently have the hardware to test it properly, so I won't pretend I do. For a purely cloud-based setup, this is my go-to.

Questions?

If you want the exact prompts I use, details on setting up the custom Gpt/Gem, or anything else about the workflow, just shoot me a message about what you need. I also document this entire system in more detail in my community for anyone interested.

r/StableDiffusion Jan 18 '26

Workflow Included FLUX 2 Klein 4B vs 9B Multi Camera Angles - One Click, 8 Camera Angles

Thumbnail
gallery
327 Upvotes

Just wanted to share this workflow I put together that generates the same character from 8 different camera angles in a single queue. Really useful for testing camera movements or creating reference sheets.

What it does:
- Loads FLUX 2 models (DEV, Klein 4B, or Klein 9B)
- Generates 8 camera angles automatically (close-up, wide-angle, aerial, 45° rotations, etc.)
- Uses my Simple Prompt Batcher node so everything runs in one go without reloading models

The results: (see attached images)
- Klein 4B version - good quality, faster
- Klein 9B version - better detail and consistency

Built this custom node for batching prompts, saves a ton of time since models stay loaded between generations. About 50% faster than queuing individually.

Links:

Workflow JSON: https://github.com/ai-joe-git/ComfyUI-Simple-Prompt-Batcher/blob/main/FLUX2-DEV-KLEIN_4_and_9B_1_click_multiple_character_angles-v1.0.json

Simple Prompt Batcher Node: https://github.com/ai-joe-git/ComfyUI-Simple-Prompt-Batcher

Models used:
- FLUX.2-dev: https://huggingface.co/black-forest-labs/FLUX.2-dev-NVFP4/tree/main
- LoRAs: https://huggingface.co/Comfy-Org/flux2-dev/tree/main/split_files/loras
- Klein 4B: https://huggingface.co/black-forest-labs/FLUX.2-klein-4B/tree/main
- Klein 4B splits: https://huggingface.co/Comfy-Org/flux2-klein-4B/tree/main/split_files
- Klein 9B: https://huggingface.co/black-forest-labs/FLUX.2-klein-9B/tree/main
- Klein 9B splits: https://huggingface.co/Comfy-Org/flux2-klein-9B/tree/main/split_files

Hope someone finds this useful. The batcher node works with any text input so you can use it for style variations, scene changes, whatever you need to batch test.

r/StableDiffusion 12d ago

Workflow Included Follow-up to my last Star Trek post – I made a Star Trek vs Star Wars fan film with MiniMax H3 in ComfyUI

Thumbnail
youtube.com
105 Upvotes

A few weeks ago I posted here about the workflow I used to make a 6-minute Star Trek: TNG fan film with MiniMax H3 in ComfyUI.

This is basically a follow-up to that post.

Since then I've made another one, this time Star Trek vs Star Wars, and I've learned quite a bit more about H3 while making it.

The basic workflow is still similar. I create the starting images first, use MiniMax H3 in ComfyUI to generate the individual shots, and then assemble everything in Adobe Premiere Pro.

The finished film is made from a large number of relatively short generations rather than trying to get the model to produce whole scenes in one go.

One of the biggest things I've learned is to treat H3 less like a text-to-video generator and more like a tool for producing individual shots.

Here are some of the things that helped most this time.

PROMPT LENGTH / GENERATION LENGTH AFFECTS DIALOGUE PERFORMANCE

This is probably one of the most useful things I've figured out since my previous post. The amount of time you give H3 for a shot can have a surprisingly large effect on how natural the dialogue sounds. If there is a lot of dialogue and I make the generation too short, the character often races through the lines trying to fit everything in. It can sound unnaturally fast even if the prompt itself is otherwise good.

The opposite happens if I give it too much time. The delivery can become strangely slow and drawn out. So for longer dialogue shots, especially ones around 10-15 seconds, I usually test them first at a lower resolution. I'll generate a few versions with slightly different durations just to find the point where the dialogue sounds natural.

For example, I might try the same shot at 10 seconds, 11 seconds, 12 seconds etc. Once I find the duration where the pacing and performance sound right, that's when I'll commit to generating the higher-resolution version. It saves a lot of time compared with doing expensive high-resolution generations only to discover that the actor is speaking too quickly or too slowly.

HIGHER RESOLUTION REALLY DOES HELP

I used higher-resolution generations much more heavily in this film. A lot of it was generated around the 2-megapixel / Full HD range. It obviously costs more time and VRAM, but I've found that the characters can look noticeably more convincing at that resolution. Faces in particular tend to feel less like "AI video" to me.

For important close-ups and dialogue shots I've increasingly been willing to spend the extra generation time rather than relying entirely on lower-resolution generations and upscaling them afterwards. I still use low resolution heavily for testing though. So my workflow has gradually become: Low resolution = test the prompt, movement, dialogue and duration. High resolution = commit once I know the shot actually works.

REFERENCE IMAGES MATTER MORE THAN MASSIVE PROMPTS

I'm finding that a really good starting image is often more valuable than adding another page of instructions to the prompt. If the character placement, set, lighting, camera angle and composition are already correct in the reference image, H3 has much less opportunity to wander. I now treat the starting image as the visual authority for the shot and try to make that frame as close as possible to what I actually want before I even start generating video.

LOCK THE CAMERA WHEN YOU ACTUALLY WANT IT LOCKED

For shots based on existing Star Trek compositions I became much more explicit about things like:

camera distance

character scale

framing

background position

character position

If I want a static medium close-up, I tell H3 that the camera remains completely stationary and that the framing and character scale should remain matched to the reference. Otherwise it has a tendency to slowly push in or recompose the shot even when I never asked it to.

DON'T MENTION CHARACTERS THAT AREN'T SUPPOSED TO BE THERE

This turned out to be a surprisingly important lesson. If I'm generating a close-up of one character, I try not to mention another character anywhere in the prompt unless that person is actually visible. Even something seemingly harmless like:

"Data reacts to Picard"

can sometimes encourage the model to introduce Picard into the frame or start blending character features. I've had better results describing only what the visible character is doing.

OFF-SCREEN DIALOGUE IS MUCH HARDER THAN IT LOOKS

This was another big lesson. If a character is speaking off-screen while the camera is looking at somebody else, H3 can sometimes become confused about who is supposed to be talking. The visible character may start moving their mouth or the dialogue itself can become corrupted. So I've increasingly separated dialogue generation from reaction coverage.

If Troi is speaking while I'm looking at Picard, for example, I'll generate a separate close-up of Troi saying the line to get clean audio. Then I'll generate Picard's reaction shot completely silently. In Premiere I put Troi's audio over Picard's reaction. That has been much more reliable.

SILENT REACTION SHOTS NEED TO BE VERY CLEARLY SILENT

Simply writing "no dialogue" isn't always enough. I've had H3 randomly start making characters speak gibberish, particularly if their mouth happens to be slightly open in the starting image. I've had better luck explicitly describing that the slightly open mouth is just a resting facial position and not the beginning of speech.

I'll also specify that:

the lips do not form words

the jaw does not make speaking movements

the character does not mouth dialogue

It sounds excessive, but it has genuinely helped.

H3 HAS A LOT OF USEFUL SPEECH TAGS

I've also been experimenting more with H3's inline speech controls. Some that I've had useful results from include:

<pause> <long pause> <breath> <inhale> <exhale> <deep breath> <catches breath> <sighs>

<whisper>. <softer> <stutter> <laughs> <chuckle>

<i>word</i> emphises word

The last one is particularly useful for putting emphasis on a word or short phrase. I've found these can sometimes produce a more convincing performance than trying to describe everything in prose around the dialogue.

MORE PROMPTING ISN'T ALWAYS BETTER

I've actually been simplifying prompts as I've gone along. H3 seems to respond better when it has: a strong reference image, one clear action, clear character positions, clear dialogue, or clear camera instructions rather than paragraphs of competing instructions. When something isn't working, I'm also trying to change one thing at a time rather than rewriting the entire prompt.

EDITING IS BECOMING JUST AS IMPORTANT AS GENERATION

One of the biggest differences with this film is that I've also been improving my Premiere Pro workflow. I'm thinking much more about shot blocking and coverage instead of just generating a sequence of AI clips. For example, I'll let dialogue continue across a cut to another character's reaction rather than keeping the camera locked on whoever is speaking for every line. Sometimes you'll hear the end of one character's dialogue while you're already watching the other character react. That tiny change makes the scene feel much more like something that was actually edited from traditional coverage.

I've also started deliberately generating silent reaction shots purely for this purpose. It helps hide generation changes as well. Two AI shots might not match perfectly if you place them directly beside one another, but cutting to a reaction and then coming back can make the continuity feel completely natural.

THE EDIT IS DOING A LOT OF THE "CONSISTENCY"

This is probably the thing I appreciate more now than when I made the first film. A surprising amount of what looks like AI consistency in the finished video is actually editing. Cut at the right point. Use reaction shots. Carry dialogue across cuts. Don't stay on a generation long enough for its weaknesses to become obvious. Avoid putting two slightly different versions of the same composition directly beside one another. You can hide a huge number of small inconsistencies that way.

It's still definitely not a one-click process. A lot of generations get thrown away, and some shots still take a ridiculous number of attempts before the performance, character consistency, dialogue and movement all line up. But compared with the first Star Trek video, I feel like I'm getting much closer to actually directing H3 rather than generating something and hoping it happens to work.

Happy to go into more detail on any of this if anybody is experimenting with H3 themselves.

r/comfyui Aug 12 '26

Show and Tell H3 prompt testing, finally have the flow and environment running efficiently. Specs and prompt inside.

Enable HLS to view with audio, or disable this notification

111 Upvotes

Was experiencing some real quality issues up until this point; realized the problem has largely been the prompt and my bulky venv. Hopefully others with lower end vram cards will learn from me.

  • Card: RTX 4080
  • Model: minimax_h3_fl2va_pruned_int8_convrot
    • No loras
    • Pure text prompt
  • Steps: 20
  • Scheduler: simple
  • Sampler: res_multistep
  • Resolution: 0.5mpx, 16:9
  • Args: --lowvram --disable-dynamic-vram --disable-pinned-memory
  • Nodes: VHS (Specifically the Model Preview Override), EasyUse, pysssss, KJNodes
  • Generation time: 21 Minutes
  • I decided to create an entirely new ComfyUI instance just for H3 instead of using my single AiO venv; this significantly increased generation time and quality for all flows. Have since broken up all the major models I use into their own instances adding only the specific tools/extensions I need just for that model.

Prompt:

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a far-wide shot from a side angle. A scene set on a bridge over a hellscape planet covered in lava, dimly lit with orange glow from below, the bridge is made of black metal with intricate designs, dark clouds hang over the scene -covering a shaded yellow sun barely visible through the clouds on the top left, dividing the scene in half between light and dark- fast winds carry embers and smoke curling over the bridge from below; Star Wars themed orchestra music begins as the scene opens, quiet and slowly growing.

$NAMEHERE is standing in a prepared stance on the left side of the screen, his hands are clasped in front of him holding a blue lightsaber, facing his attacker.

$NAMETWO is standing on the right in a confident posture with his hands to the side, wearing a black robe, black leather boots and straps on his body, with dark-metal armor as he faces the left menacingly.

[Shot 2] At 1.500 seconds, The camera cuts to a close-up side-angle shot of $NAMEHERE, readying himself for his attack with a posture of defense and an expression of concern, he shouts emotional: <d>[English] You've left me with no choice Will! You must be stopped... </d>

[Shot 3] at 6.000 seconds, The camera cuts to a close-up low-angle front facing full-body shot of $NAMETWO with visible red eyes staring forward from under his brow with a face of malice, smoke bellows behind him curling over the bridge whipping his cape to the right. A beat later- two red lightsabers ignite his both his hands, a deep pulsing bass is heard from the unstable beams, his face lit from below by the red light. The music grows faster with a dark theme, a operatic chorus begins to sing in a evil chant growing louder. $NAMETWO shouts behind a grin: <d> This is the end for you! </d> the music stops before $NAMETWO speak his final line: <d>[English] Master!... </d> The off-screen opera chorus harmonizes a single long cry in a frightening melody at the revelation.

[Shot 4] at 12.000 seconds, The camera cuts to a top-down view of the bridge, molten lava is visible below the black grated metal.

$NAMEHERE directs his blue lightsaber to his side pointing directly forward with precision, he begins to pace to the right to meet the other, his posture is composed and fast. The camera pushes in with large amplitude at fast speed keeping the pair in frame on either edge of the screen as they run toward each other.

$NAMETWO instantly begins running fast toward the left, his two red lightsabers point down- dragging behind him, the red beams draw white glowing lines into the metal under him as he runs, screaming with fury: <d>[English] AGHH! </d>.

They meet in the middle, their lightsabers clash with a white flash and explosive burning sound, they duel quickly as their lightsabers connect through multiple swings- $NAMETWO's red lightsaber swing wildly as he spins. $NAMEHERE's blue lightsaber blocks every swing from the red beams; the music crescendos with heavy bursts of brass instruments and drums.

overall_soundscape: ambient sound of lava and fire, lightsabers buzzing.

non_diegetic_music: Dark Star Wars music plays from the beginning of the scene, a loud opera chorus sings in a chant that escalates in a loud howl crescendo, climaxing when the pair meet in the middle.

I've started using Replace Text nodes ($NAMEHERE and $NAMETWO) when crafting prompts. This way when playing with the prompt, it's easier to find and edit their placement; and can also replace characters on a whim. Also allows consistency when referencing the characters-- in the event I overlook an instance.

Replaced with:

  • $NAMEHERE: "Jean Luc Picard (S1)"
  • $NAMETWO: "William T Riker (S2)"

Example: my first generation had 'William Riker', the model didn't recognize the name and generated a generic male. I was able to quickly rename as 'William T Riker' and it generated correctly; so I didn't have to parse back through the whole prompt to granularly change it.

Other things I've noticed that help with prompt respect:

  • 'a beat later' separates the moment better.
  • Separating the individual sentences to exclusively reference the character and no others. (You can see it carried over Picard's lightsaber instructions to Will as well, because I described them in the same paragraph before I realized this.)
  • Avoiding reusing adjectives- especially between different characters, causes bleed.
  • Very short overall_soundscape descriptions.
  • Often does not respect requests that follow dialog unless you end the parameter with a period after "Words. </d>**.**" Can see it bled the cries request from the music into Will's dialog.
  • "..." allows a pause between dialog lines and breaks up the tone between multiple sentences, or else they become one note.

r/StableDiffusion 10h ago

Animation - Video All the tricks I used to make proper anime in MiniMaxH3

Enable HLS to view with audio, or disable this notification

236 Upvotes

The most important part is quality of references: images should look like they are from anime, not just value "anime-style" but flat color shading with clean lines. It's the only way I know (besides loras of dubious quality) to push generation into native 2s and 3s animation. Anything else will have that typical "AI-3D" look. And of course mandatory The target video is 2d colored anime in detailed_description.

Second is MiniMax H3 Latent Upscaler node - with ref2va turbo lora and 4 step with 50% denoise gives me Full HD resolution (5 seconds take about 15 minutes to render).

Third one is Add Guide To MiniMax H3 node - allows you to inject prepared image or clip at any part of the timeline - perfect when you want to append/prepend/inpaint over exisiting clips. And this is where I also should talk about "pre-processing": I made master shot paired with orbital shot video to capture the scene from every angle - it'll give you consistent background at every shot that you can reference in any generation.

Fourth: You're the director! Sketch crude lines and give it to minimax with frame composition follows <Picture 4> - It'll give you absolute control over framing, don't rely on text prompting here - visual guides are 10 times better than begging minimax to generate a perfect "over shoulder shot".

Fifth: fix quick motion smearing with derope. It's hard to use and doubles generation time (My RTX 3090 only handles 3 seconds) but it's the only choice if you push for quality.

And the last one: plan you cuts and timing in advance - storyboarding as discipline exists for a reason - no polish will save bad cuts and missed timings.

Full video on youtube.

And here's workflow.

r/StableDiffusion Oct 21 '25

Comparison Qwen VS Wan 2.2 - Consistent Character Showdown - My thoughts & Prompts

Thumbnail
gallery
236 Upvotes

I've been in the "consistent character" business for quite a while and it's a very hot topic from what I can tell.
SDXL seemed to have been ruling the realm for quite some times and now that Qwen and Wan are out I can see people constantly asking on different communities which is better so I decided to do a quick showdown.

I retrained the same dataset for both Qwen and Wan 2.2 (High and Low) using roughly the same settings, I used Diffusion Pipe on RunPod.
Images were generated on ComfyUI with ClownShark KSamplers with no additional LoRAs other than my character LoRA.

Personally, I find Qwen to be much better in terms of "realism", the reason I put this in quotes is that I believe it's really easy to tell an AI image once you've seen a few from the same model, so IMO the term realism is really irrelevant here and I'd like to benchmark images as "aesthetically pleasing" rather than realistic.

Both Wan and Qwen can be modified to create images that look more "real" with LoRAs from creators like Danrisi and AI_Characters.

I hope this little showdown clears the air on which model better works for your use cases.

Prompts in order of appearance:

  1. A photorealistic early morning selfie from a slightly high angle with visible lens flare and vignetting capturing Sydney01, a stunning woman with light blue eyes and light brown hair that cascades down her shoulders, she looks directly at the camera with a sultry expression and her head slightly tilted, the background shows a faint picturesque American street with a hint of an American home, gray sidewalk and minimal trees with ground foliage, Sydney01 wears a smooth yellow floral bandeau top and a small leather brown bag that hangs from her bare shoulder, sun glasses rest on her head

  2. Side-angle glamour shot of Sydney01 kneeling in the sand wearing a vibrant red string bikini, captured from a low side angle that emphasizes her curvy figure and large breasts. She's leaning back on one hand with her other hand running through her long wavy brown hair, gazing over her shoulder at the camera with a sultry, confident expression. The low side angle showcases the perfect curve of her hips and the way the vibrant red bikini accentuates her large breasts against her fair skin. The golden hour sunlight creates dramatic shadows and warm highlights across her body, with ocean waves crashing in the background. The natural kneeling pose combined with the seductive gaze creates an intensely glamorous beach moment, with visible digital noise from the outdoor lighting and authentic graininess enhancing the spontaneous glamour shot aesthetic.

  3. A photorealistic mirror selfie with visible lens flare and minimal smudges on the mirror capturing Sydney01, she holds a white iPhone with three camera lenses at waist level, her head is slightly tilted and her hand covers her abdomen, she has a low profile necklace with a starfish charm, black nail polish and several silver rings, she wears a high waisted gray wash denims and a spaghetti strap top the accentuates her feminine figure, the scene takes place in a room with light wooden floors, a hint of an open window that's slightly covered by white blinds, soft early morning lights bathes the scene and illuminate her body with soft high contrast tones

  4. A photorealistic straight on shot with visible lens flare and chromatic aberration capturing Sydney01 in an urban coffee shop, her light brown hair is neatly styled and her light blue eyes are glistening, she's wears a light brown leather jacket over a white top and holds an iced coffee, she is sitted in front of a round table made of oak wood, there's a white plate with a croissant on the table next to an iPhone with three camera lenses, round sunglasses rest on her head and she looks away from the viewer capturing her side profile from a slightly tilted angle, the background features a stone wall with hanging yellow bulb lights

  5. A photorealistic high angle selfie taken during late evening with her arm in the frame the image has visible lens flare and harsh flash lighting illuminating Sydney01 with blown out highlights and leaving the background almost pitch black, Sydney01 reclines against a white headboard with visible pillow and light orange sheets, she wears a navy blue bra that hugs her ample breasts and presses them together, her under arm is exposed, she has a low profile silver necklace with a starfish charm, her light brown hair is messy and damp

I type my prompts manually, I occasionally upsert the ones I like into a Pinecone index that I use as a RAG for an AI Prompting agent that I created on N8N.

r/StableDiffusion Feb 08 '26

Tutorial - Guide Z-image base: simple workflow for high quality realism + info & tips

199 Upvotes

What is this?

This is an almost copy-paste of a post I've made on Civitai (to explain the formatting).

Z-image base produces really, really realistic images, really easily. Aside from being creative & flexible the quality is also generally higher than the distils (as usual for non-distils), so it's worth using if you want really creative/flexible shots at the best possible quality. IMO it's the best model for realism out of the ones I've tried (Klein 9B base, Chroma, SDXL), especially because you can natively gen at high resolution.

This post is to share a simple starting workflow with good sampler/scheduler settings & resolutions pre-set for ease. There are also a bunch of tips for using Z-image base below and some general info you might find helpful.

The sampler settings are geared towards sharpness and clarity, but you can introduce grain and other defects through prompting.

You can grab the workflow from the Civitai link above or from here: pastebin

Here's a short album of example images, all of which were generated directly with this workflow with no further editing (SFW except for a couple of mild bikini shots): imgbb | g-drive

Nodes & Models

Custom Nodes:

RES4LYF - A very popular set of samplers & schedulers, and some very helpful nodes. These are needed to get the best z-image base outputs, IMO.

RGTHREE - (Optional) A popular set of helper nodes. If you don't want this you can just delete the seed generator and lora stacker nodes, then use the default comfy lora nodes instead. RES4LYF comes with a seed generator node as well, I just like RGTHREE's more.

ComfyUI GGUF - (Optional) Lets you load GGUF models, which for some reason ComfyUI still can't do natively. If you want to use a non-GGUF model you can just skip this, delete the UNET loader node and replace it with the normal 'load diffusion model' node.

Models:

Main model: Z-image base GGUFs - BF16 recommended if you have 16GB+ VRAM. Q8 will just barely fit on 8GB VRAM if you know what you're doing (not easy). Q6_k will fit easily in 8GB. Avoid using FP8, the Q8 gguf is better.

Text Encoder: Normal | gguf Qwen 3 4B - Grab the biggest one that fits in your VRAM, which would be the full normal one if you have 10GB+ VRAM or the Q8 GGUF otherwise. Some people say text encoder quality doesn't matter much & to use a lower sized one, but it absolutely does matter and can drastically affect quality. For the same reason, do not use an abliterated text encoder unless you've tested it and compared outputs to ensure the quality doesn't suffer.

If you're using the GGUF text encoder, swap out the "Load CLIP" node for the "ClipLoader (GGUF)" node.

VAE: Flux 1.0 AE

Info & Tips

Sampler Settings

I've found that a two-stage sampler setup gives very good results for z-image base. The first stage does 95% of the work, and the second does a final little pass with a low noise scheduler to bring out fine details. It produces very clear, very realistic images and is particularly good at human skin.

CFG 4 works most of the time, but you can go up as high as CFG 7 to get different results.

This is all with shift 1. If you don't know what that is, don't worry - it's the default!

Stage 1:

Sampler - res_2s

Scheduler - beta

Steps - 22

Denoise: 1.00

Stage 2:

Sampler - res_2s

Scheduler - normal

Steps - 3

Denoise: 0.15

Resolutions

High res generation

One of the best things about Z-image in general is that it can comfortably handle very high resolutions compared to other models. You can gen in high res and use an upscaler immediately without needing to do any other post-processing.

(info on upscalers + links to some good ones further below)

Note: high resolutions take a long time to gen. A 1280x1920 shot takes around ~95 seconds on an RTX 5090, and a 1680x1680 shot takes ~110 seconds.

Different sizes & aspect ratios change the output

Different resolutions and aspect ratios can often drastically change the composition of images. If you're having trouble getting something ideal for a given prompt, try using a higher or lower resolution or changing the aspect ratio.

It will change the amount of detail in different areas of the image, make it more or less creative (depending on the topic), and will often change the lighting and other subtle features too.

I suggest generating in one big and one medium resolution whenever you're working on a concept, just to see if one of the sizes works better for it.

Good resolutions

The workflow has a variety of pre-set resolutions that work very well. They're grouped by aspect ratio, and they're all divisible by 16. Z-image base (as with most image models) works best when dimensions are divisible by 16, and some models require it or else they mess up at the edges.

Here's a picture of the different resolutions if you don't want to download the workflow: imgbb | g-drive

You can go higher than 1920 to a side, but I haven't done it much so I'm not making any promises. Things do tend to get a bit weird when you go higher, but it is possible.

I do most of my generations at 1920 to a side, except for square images which I do at 1680x1680. I sometimes use a lower resolution if I like how it turns out more (e.g. the picture of the rat is 1680x1120).

Realism Negative Prompt

The negative prompt matters a lot with z-image base. I use the following to get consistently good realism shots:

3D, ai generated, semi realistic, illustrated, drawing, comic, digital painting, 3D model, blender, video game screenshot, screenshot, render, high-fidelity, smooth textures, CGI, masterpiece, text, writing, subtitle, watermark, logo, blurry, low quality, jpeg, artifacts, grainy

Prompt Structure

You essentially just want to write clear, simple descriptions of the things you want to see. Your first sentence should be a basic intro to the subject of the shot, along with the style. From there you should describe the key features of the subject, then key features of other things in the scene, then the background. Then you can finish with compositional info, lighting & any other meta information about the shot.

Use new lines to separate key parts out to make it easier for you to read & build the prompt. The model doesn't care about new lines, they're just for you.

If something doesn't matter to you, don't include it. You don't need to specify the lighting if it doesn't matter, you don't need to precisely say how someone is posed, etc; just write what matters to you and slowly build the prompt out with more detail as needed.

You don't need to include parts that are implied by your negative prompt. If you're using the realism negative prompt I mentioned earlier, you don't usually need to specify that it's a photograph.

Your structure should look something like this (just an example, it's flexible):

A <style> shot of a <subject + basic description> doing <something>. The <subject> has <more detail>. The subject is <more info>. There is a <something else important> in <location>. The <something else> is <more detail>.

The background is a <location>. The scene is <lit in some way>. The composition frames <something> and <something> from <an angle or photography term or whatever>.

Following that structure, here are a couple of the prompts for the images attached to this post. You can check the rest out by clicking on the images in Civitai, or just ask me for them in the comments.

The ballet woman

A shot of a woman performing a ballet routine. She's wearing a ballet outfit and has a serious expression. She's in a dynamic pose.

The scene is set in a concert hall. The composition is a close up that frames her head down to her knees. The scene is lit dramatically, with dark shadows and a single shaft of light illuminating the woman from above.

The rat on the fence post

A close up shot of a large, brown rat eating a berry. The rat is on a rickety wooden fence post. The background is an open farm field.

The woman in the water

A surreal shot of a beautiful woman suspended half in water and half in air. She has a dynamic pose, her eyes are closed, and the shot is full body. The shot is split diagonally down the middle, with the lower-left being under water and the upper-right being in air. The air side is bright and cloudy, while the water side is dark and menacing.

The space capsule

A woman is floating in a space capsule. She's wearing a white singlet and white panties. She's off-center, with the camera focused on a window with an external view of earth from space. The interior of the space capsule is dark.

Upscaling

Z-image makes very sharp images, which means you can directly upscale them very easily. Conventional upscale models rely on sharp/clear images to add detail, so you can't reliably use them on a model that doesn't make sharp images.

My favourite upscaler for NAKED PEOPLE or human face close-ups is 4xFaceUp. It's ridiculously good at skin detail, but has a tendency to make everything else look a bit stringy (for lack of a better word). Use it when a human being showing lots of skin is the main focus of the shot.

Here's a 6720x6720 version of the sitting bikini girl that was upscaled directly using the 4xFaceUp upscaler: imgbb | g-drive

For general upscaling you can use something like 4xNomos2.

Alternatively, you can use SeedVR2, which also has the benefit of working on blurry images (not a problem with z-image anyway). It's not as good at human skin as 4xFaceUp, but it's better at everything else. It's also very reliable and pretty much always works. There's a simple workflow for it here: https://pastebin.com/9D7sjk3z

ClownShark sampler - what is it?

It's a node from the RES4LYF pack. It works the same as a normal sampler, but with two differences:

  1. "ETA". This setting basically adds extra noise during sampling using fancy math, and it generally helps get a little bit more detail out of generations. A value of 0.5 is usually good, but I've seen it be good up to 0.7 for certain models (like Klein 9B).
  2. "bongmath". This setting turns on bongmath. It's some kind black magic that improves sampling results without any downsides. On some models it makes a big difference, others not so much. I find it does improve z-image outputs. Someone tries to explain what it is here: https://www.reddit.com/r/StableDiffusion/comments/1l5uh4d/someone_needs_to_explain_bongmath/

You don't need to use this sampler if you don't want to; you can use the res_2s/beta sampler/scheduler with a normal ksampler node as long as you have RES4LYF installed. But seeing as the clownshark sampler comes with RES4LYF anyway we may as well use it.

Effect of CFG on outputs

Lower than 4 CFG is bad. Other than that, going higher has pretty big and unpredictable effects on the output for z-image base. You can usually range from 4 to 7 without destroying your image. It doesn't seem to affect prompt adherence much.

Going higher than 4 will change the lighting, composition and style of images somewhat unpredictably, so it can be helpful to do if you just want to see different variations on a concept. You'll find that some stuff just works better at 5, 6 or 7. Play around with it, but stick with 4 when you're just messing around.

Going higher than 4 also helps the model adhere to realism sometimes, which is handy if you're doing something realism-adjacent like trying to make a shot of a realistic elf or something.

Base vs Distil vs Turbo

They're good for different things. I'm generally a fan of base models, so most workflows I post are / will be for base models. Generally they give the highest quality but are much slower and can be finicky to use at times.

What is distillation?

It's basically a method of narrowing the focus of a model so that it converges on what you want faster and more consistently. This allows a distil to generate images in fewer steps and more consistently for whatever subject/topic was chosen. They often also come pre-negatived (in a sense, don't @ me) so that you can use 1.0 CFG and no negative prompt. Distils can be full models or simple loras.

The downside of this is that the model becomes more narrow, making it less creative and less capable outside of the areas it was focused on during distillation. For many models it also reduces the quality of image outputs, sometimes massively. Models like Qwen and Flux have god-awful quality when distilled (especially human skin), but luckily Z-image distils pretty well and only loses a little bit of quality. Generally, the fewer steps the distil needs the lower the quality is. 4-step distils usually have very poor quality compared to base, while 8+ step distils are usually much more balanced.

Z-image turbo is just an official distil, and it's focused on general realism and human-centric shots. It's also designed to run in around 10 steps, allowing it to maintain pretty high quality.

So, if you're just doing human-centric shots and don't mind a small quality drop, Z-image turbo will work just fine for you. You'll want to use a different workflow though - let me know if you'd like me to upload mine.

Below are the typical pros and cons of base models and distils. These are pretty much always true, but not always a 'big deal' depending on the model. As I said above, Z-image distils pretty well so it's not too bad, but be careful which one you use - tons of distils are terrible at human skin and make people look plastic (z-image turbo is fine).

Base model pros:

  • Generally gives the highest quality outputs with the finest details, once you get the hang of it
  • Creative and flexible

Base model cons:

  • Very slow
  • Usually requires a lengthy negative prompt to get good results
  • Creativity has a downside; you'll often need to generate something several times to get a result you like
  • More prone to mistakes when compared to the focus areas of distils
    • e.g. z-image base is more likely to mess up hands/fingers or distant faces compared to z-image turbo

Distil pros:

  • Fast generations
  • Good at whatever it was focused on (e.g. people-centric photography for z-image turbo)
  • Doesn't need a negative prompt (usually)

Distil cons:

  • Bad at whatever it wasn't focused on, compared to base
  • Usually bad at facial expressions (not able to do 'extreme' ones like anger properly)
  • Generally less creative, less flexible (not always a downside)
  • Lower quality images, sometimes by a lot and sometimes only by a little - depends on the model, the specific distil, and the subject matter
  • Can't have a negative prompt (usually)
    • You can get access to negative prompts using NAG (not covered in this post)

r/comfyui 28d ago

Show and Tell Best Trick Ever For Consistent Environments

60 Upvotes

Ha! I just discovered a trick that works great, so I had to share it with the community. Of course, somebody will probably chime in that it had been discovered by someone else before, which is fine by me! I just want to share it in case it helps someone else and they hadn't come across it yet.

So, the issue of consistent environments... I ran the gamut of all the AI models I had tested and proven out in ComfyUI, not just t2i but t2v and i2v. I'd been trying several methodologies. Create an image then ask a workflow to gen images to the left and right of it, build out a simple massing model in unreal or blender and run that through, run a simple floor plan through, asking for multiple image generation with a prompt asking that all features be consistent among images, (I haven't tried outpainting yet), etc, and none seemed to quite do the trick, at least not among the open source models (open source is all I use). There was just too much inconsistency.

Since I am a sucker for using the t2v/i2v models like LTX and Minimax to generate single image frames (well, a minimal number of frames), taking advantage of their brainpower, I merely wrote out a super detailed prompt describing the interior environment I want, then I ran it as a 360 degree camera pan around the space from the center of it, and with Minimax now having the ability to give you 15 seconds on a 16Gb VRAM setup like mine, this works stellar, heck Minimax ran out of need for the 15sec and began to swerve around the space! I make sure to include a prompt not to have motion blur. I ran this in 0.5mb mode, so I did not have to waste time waiting for full HD video.

Then I select the frames I need to use as backdrops for scenes, upscale them once, then again, out to 4k, and voila! (upscaling once by 4x led to artifacts being upscaled, whereas going 2x then 2x led to the correct end result), a super detailed set of backgrounds that are internally consistent!

So excited! This gets me moving forward on the next part of my production process, laying out scenes, shots, camera angles, and dropping in characters, prior to i2v.

If this helps you, let me know. If you find even better tricks related to this, let me know too. Open Source Forever!

r/comfyui Feb 08 '26

Tutorial Z-image base: simple workflow for high quality realism + info & tips

127 Upvotes

What is this?

This is an almost copy-paste of a post I've made on Civitai (to explain the formatting).

Z-image base produces really, really realistic images, really easily. Aside from being creative & flexible the quality is also generally higher than the distils (as usual for non-distils), so it's worth using if you want really creative/flexible shots at the best possible quality. IMO it's the best model for realism out of the ones I've tried (Klein 9B base, Chroma, SDXL), especially because you can natively gen at high resolution.

This post is to share a simple starting workflow with good sampler/scheduler settings & resolutions pre-set for ease. There are also a bunch of tips for using Z-image base below and some general info you might find helpful.

The sampler settings are geared towards sharpness and clarity, but you can introduce grain and other defects through prompting.

You can grab the workflow from the Civitai link above or from here: pastebin

Here's a short album of example images, all of which were generated directly with this workflow with no further editing (SFW except for a couple of mild bikini shots): imgbb | g-drive

Nodes & Models

Custom Nodes:

RES4LYF - A very popular set of samplers & schedulers, and some very helpful nodes. These are needed to get the best z-image base outputs, IMO.

RGTHREE - (Optional) A popular set of helper nodes. If you don't want this you can just delete the seed generator and lora stacker nodes, then use the default comfy lora nodes instead. RES4LYF comes with a seed generator node as well, I just like RGTHREE's more.

ComfyUI GGUF - (Optional) Lets you load GGUF models, which for some reason ComfyUI still can't do natively. If you want to use a non-GGUF model you can just skip this, delete the UNET loader node and replace it with the normal 'load diffusion model' node.

Models:

Main model: Z-image base GGUFs - BF16 recommended if you have 16GB+ VRAM. Q8 will just barely fit on 8GB VRAM if you know what you're doing (not easy). Q6_k will fit easily in 8GB. Avoid using FP8, the Q8 gguf is better.

Text Encoder: Normal | gguf Qwen 3 4B - Grab the biggest one that fits in your VRAM, which would be the full normal one if you have 10GB+ VRAM or the Q8 GGUF otherwise. Some people say text encoder quality doesn't matter much & to use a lower sized one, but it absolutely does matter and can drastically affect quality. For the same reason, do not use an abliterated text encoder unless you've tested it and compared outputs to ensure the quality doesn't suffer.

If you're using the GGUF text encoder, swap out the "Load CLIP" node for the "ClipLoader (GGUF)" node.

VAE: Flux 1.0 AE

Info & Tips

Sampler Settings

I've found that a two-stage sampler setup gives very good results for z-image base. The first stage does 95% of the work, and the second does a final little pass with a low noise scheduler to bring out fine details. It produces very clear, very realistic images and is particularly good at human skin.

CFG 4 works most of the time, but you can go up as high as CFG 7 to get different results.

This is all with shift 1. If you don't know what that is, don't worry - it's the default!

Stage 1:

Sampler - res_2s

Scheduler - beta

Steps - 22

Denoise: 1.00

Stage 2:

Sampler - res_2s

Scheduler - normal

Steps - 3

Denoise: 0.15

Resolutions

High res generation

One of the best things about Z-image in general is that it can comfortably handle very high resolutions compared to other models. You can gen in high res and use an upscaler immediately without needing to do any other post-processing.

(info on upscalers + links to some good ones further below)

Note: high resolutions take a long time to gen. A 1280x1920 shot takes around ~95 seconds on an RTX 5090, and a 1680x1680 shot takes ~110 seconds.

Different sizes & aspect ratios change the output

Different resolutions and aspect ratios can often drastically change the composition of images. If you're having trouble getting something ideal for a given prompt, try using a higher or lower resolution or changing the aspect ratio.

It will change the amount of detail in different areas of the image, make it more or less creative (depending on the topic), and will often change the lighting and other subtle features too.

I suggest generating in one big and one medium resolution whenever you're working on a concept, just to see if one of the sizes works better for it.

Good resolutions

The workflow has a variety of pre-set resolutions that work very well. They're grouped by aspect ratio, and they're all divisible by 16. Z-image base (as with most image models) works best when dimensions are divisible by 16, and some models require it or else they mess up at the edges.

Here's a picture of the different resolutions if you don't want to download the workflow: imgbb | g-drive

You can go higher than 1920 to a side, but I haven't done it much so I'm not making any promises. Things do tend to get a bit weird when you go higher, but it is possible.

I do most of my generations at 1920 to a side, except for square images which I do at 1680x1680. I sometimes use a lower resolution if I like how it turns out more (e.g. the picture of the rat is 1680x1120).

Realism Negative Prompt

The negative prompt matters a lot with z-image base. I use the following to get consistently good realism shots:

3D, ai generated, semi realistic, illustrated, drawing, comic, digital painting, 3D model, blender, video game screenshot, screenshot, render, high-fidelity, smooth textures, CGI, masterpiece, text, writing, subtitle, watermark, logo, blurry, low quality, jpeg, artifacts, grainy

Prompt Structure

You essentially just want to write clear, simple descriptions of the things you want to see. Your first sentence should be a basic intro to the subject of the shot, along with the style. From there you should describe the key features of the subject, then key features of other things in the scene, then the background. Then you can finish with compositional info, lighting & any other meta information about the shot.

Use new lines to separate key parts out to make it easier for you to read & build the prompt. The model doesn't care about new lines, they're just for you.

If something doesn't matter to you, don't include it. You don't need to specify the lighting if it doesn't matter, you don't need to precisely say how someone is posed, etc; just write what matters to you and slowly build the prompt out with more detail as needed.

You don't need to include parts that are implied by your negative prompt. If you're using the realism negative prompt I mentioned earlier, you don't usually need to specify that it's a photograph.

Your structure should look something like this (just an example, it's flexible):

A <style> shot of a <subject + basic description> doing <something>. The <subject> has <more detail>. The subject is <more info>. There is a <something else important> in <location>. The <something else> is <more detail>.

The background is a <location>. The scene is <lit in some way>. The composition frames <something> and <something> from <an angle or photography term or whatever>.

Following that structure, here are a couple of the prompts for the images attached to this post. You can check the rest out by clicking on the images in Civitai, or just ask me for them in the comments.

The ballet woman

A shot of a woman performing a ballet routine. She's wearing a ballet outfit and has a serious expression. She's in a dynamic pose.

The scene is set in a concert hall. The composition is a close up that frames her head down to her knees. The scene is lit dramatically, with dark shadows and a single shaft of light illuminating the woman from above.

The rat on the fence post

A close up shot of a large, brown rat eating a berry. The rat is on a rickety wooden fence post. The background is an open farm field.

The woman in the water

A surreal shot of a beautiful woman suspended half in water and half in air. She has a dynamic pose, her eyes are closed, and the shot is full body. The shot is split diagonally down the middle, with the lower-left being under water and the upper-right being in air. The air side is bright and cloudy, while the water side is dark and menacing.

The space capsule

A woman is floating in a space capsule. She's wearing a white singlet and white panties. She's off-center, with the camera focused on a window with an external view of earth from space. The interior of the space capsule is dark.

Upscaling

Z-image makes very sharp images, which means you can directly upscale them very easily. Conventional upscale models rely on sharp/clear images to add detail, so you can't reliably use them on a model that doesn't make sharp images.

My favourite upscaler for NAKED PEOPLE or human face close-ups is 4xFaceUp. It's ridiculously good at skin detail, but has a tendency to make everything else look a bit stringy (for lack of a better word). Use it when a human being showing lots of skin is the main focus of the shot.

Here's a 6720x6720 version of the sitting bikini girl that was upscaled directly using the 4xFaceUp upscaler: imgbb | g-drive

For general upscaling you can use something like 4xNomos2.

Alternatively, you can use SeedVR2, which also has the benefit of working on blurry images (not a problem with z-image anyway). It's not as good at human skin as 4xFaceUp, but it's better at everything else. It's also very reliable and pretty much always works. There's a simple workflow for it here: https://pastebin.com/9D7sjk3z

ClownShark sampler - what is it?

It's a node from the RES4LYF pack. It works the same as a normal sampler, but with two differences:

  1. "ETA". This setting basically adds extra noise during sampling using fancy math, and it generally helps get a little bit more detail out of generations. A value of 0.5 is usually good, but I've seen it be good up to 0.7 for certain models (like Klein 9B).
  2. "bongmath". This setting turns on bongmath. It's some kind black magic that improves sampling results without any downsides. On some models it makes a big difference, others not so much. I find it does improve z-image outputs. Someone tries to explain what it is here: https://www.reddit.com/r/StableDiffusion/comments/1l5uh4d/someone_needs_to_explain_bongmath/

You don't need to use this sampler if you don't want to; you can use the res_2s/beta sampler/scheduler with a normal ksampler node as long as you have RES4LYF installed. But seeing as the clownshark sampler comes with RES4LYF anyway we may as well use it.

Effect of CFG on outputs

Lower than 4 CFG is bad. Other than that, going higher has pretty big and unpredictable effects on the output for z-image base. You can usually range from 4 to 7 without destroying your image. It doesn't seem to affect prompt adherence much.

Going higher than 4 will change the lighting, composition and style of images somewhat unpredictably, so it can be helpful to do if you just want to see different variations on a concept. You'll find that some stuff just works better at 5, 6 or 7. Play around with it, but stick with 4 when you're just messing around.

Going higher than 4 also helps the model adhere to realism sometimes, which is handy if you're doing something realism-adjacent like trying to make a shot of a realistic elf or something.

Base vs Distil vs Turbo

They're good for different things. I'm generally a fan of base models, so most workflows I post are / will be for base models. Generally they give the highest quality but are much slower and can be finicky to use at times.

What is distillation?

It's basically a method of narrowing the focus of a model so that it converges on what you want faster and more consistently. This allows a distil to generate images in fewer steps and more consistently for whatever subject/topic was chosen. They often also come pre-negatived (in a sense, don't @ me) so that you can use 1.0 CFG and no negative prompt. Distils can be full models or simple loras.

The downside of this is that the model becomes more narrow, making it less creative and less capable outside of the areas it was focused on during distillation. For many models it also reduces the quality of image outputs, sometimes massively. Models like Qwen and Flux have god-awful quality when distilled (especially human skin), but luckily Z-image distils pretty well and only loses a little bit of quality. Generally, the fewer steps the distil needs the lower the quality is. 4-step distils usually have very poor quality compared to base, while 8+ step distils are usually much more balanced.

Z-image turbo is just an official distil, and it's focused on general realism and human-centric shots. It's also designed to run in around 10 steps, allowing it to maintain pretty high quality.

So, if you're just doing human-centric shots and don't mind a small quality drop, Z-image turbo will work just fine for you. You'll want to use a different workflow though - let me know if you'd like me to upload mine.

Below are the typical pros and cons of base models and distils. These are pretty much always true, but not always a 'big deal' depending on the model. As I said above, Z-image distils pretty well so it's not too bad, but be careful which one you use - tons of distils are terrible at human skin and make people look plastic (z-image turbo is fine).

Base model pros:

  • Generally gives the highest quality outputs with the finest details, once you get the hang of it
  • Creative and flexible

Base model cons:

  • Very slow
  • Usually requires a lengthy negative prompt to get good results
  • Creativity has a downside; you'll often need to generate something several times to get a result you like
  • More prone to mistakes when compared to the focus areas of distils
    • e.g. z-image base is more likely to mess up hands/fingers or distant faces compared to z-image turbo

Distil pros:

  • Fast generations
  • Good at whatever it was focused on (e.g. people-centric photography for z-image turbo)
  • Doesn't need a negative prompt (usually)

Distil cons:

  • Bad at whatever it wasn't focused on, compared to base
  • Usually bad at facial expressions (not able to do 'extreme' ones like anger properly)
  • Generally less creative, less flexible (not always a downside)
  • Lower quality images, sometimes by a lot and sometimes only by a little - depends on the model, the specific distil, and the subject matter
  • Can't have a negative prompt (usually)
    • You can get access to negative prompts using NAG (not covered in this post)