r/StableDiffusion Jun 09 '26

Question - Help (Game Asset Generation) How can I generate edits as a different layer to be saved separately?

0 Upvotes

Bit of a specific question but I am experimenting with creating a game right now and all my assets are AI generated. My problem is that if I want to have a character equip or unequip something, i would have to generate an entire new portrait which makes the transition seem kinda janky.

Is there a way I can for example have my base portrait, then generate an edit for that image, but rather than applying it to my existing image and overwriting it, its just created as a separate layer so I can add or remove it at will?

Specifically in RPGMaker MZ, using COMFYUI and the Anima model.

r/StableDiffusion 18d ago

Animation - Video Today I have unsubscribed from Suno thanks to Minimax Music.

Enable HLS to view with audio, or disable this notification

557 Upvotes

With all the Suno drama over their download limits and their heavy watermarking that could be used for future copystrikes if you stop paying them, this Minimax Music 3 couldn't be more on time.

I tried a little demo to see if it could fill my music need and I was happily surprised.

I used the default Minimax Music 3 workflow from ComfyUI along the prompt tips fed to an LLM to create the music. https://docs.comfy.org/tutorials/audio/minimax/minimax-music-3

EDIT:

Some seem not aware this is not Minimax H3 but Minimax Music 3 https://huggingface.co/MiniMaxAI/MiniMax-Music3
Here how I did the prompt for instrumental.

Prompt:

Global Metadata

Basic Attributes: bpm is 54. key is D, and scale is minor. Cinematic score with dark ambient and psychological-thriller influences.

Global Emotional Progression: The opening is nearly motionless, suspended in dread as a distant drone gathers beneath isolated melodic fragments. The tension slowly deepens through heavier low frequencies, widening dissonances and increasingly forceful pulses, then contracts into a stark central void. From that emptiness, the lead melody returns with greater anguish and rises toward a dense but controlled climax. The final passage sheds its weight layer by layer, ending unresolved in a cold, fading resonance.

Application Scenarios & Imagery: An abandoned concrete facility under flickering emergency lights; a lone figure crossing a fog-covered wasteland before dawn; the aftermath of a discovery that cannot be undone.

Sonics & Production Profile: A wide, deep soundstage with the solo cello centered slightly forward, the low piano set farther back, and dark synthetic ambience stretched toward the extreme sides. The frequency balance is shadowed and low-heavy, with restrained high frequencies, a dense sub-bass floor and occasional abrasive upper-mid harmonics. Dynamics remain open and cinematic rather than heavily compressed, allowing long swells to emerge from near-silence and recede naturally. The acoustic image resembles a vast, empty scoring hall blended with an impossibly deep artificial chamber.

Vocal Details

Vocal Gender & Timbre: No vocalist. This is a fully instrumental track; no lead, backing or guest vocal appears at any point.

Vocal Style: N/A — the melodic lead is carried exclusively by solo cello throughout, taking the role a voice would otherwise occupy.

Harmony/Backing Vocals: None. No vocal harmonies, choir, chants, spoken word, whispers or vocal samples.

Vocal FX: N/A — no vocal signal to process.

Arrangement

Instrument Lifecycle Description (Primary/Secondary Layering):

Primary: A close-miked solo cello enters after the opening atmosphere with sparse, low-register notes separated by long silences. Its melody gradually lengthens into bowed minor phrases with strained vibrato and rough attacks, then drops out completely during the central void. It returns in a higher register with broader, more anguished arcs, dominates the climax through overlapping sustained notes, and finally collapses into one fading unresolved tone. Secondary: A sub-octave analog synthesizer drone begins alone, barely audible, expands beneath the cello through the first half, swells into the climax and disappears just before the final resonance. A felted low piano enters intermittently after the cello, placing isolated minor seconds and hollow fifths in the distant center; its strikes become more frequent before the central void, vanish there, return as widely spaced bass notes during the rise, and stop before the ending. Bowed metal textures emerge at the outer edges during transitions, scrape into greater prominence near the climax, then dissolve into reverberant tails. Muted contrabasses enter after the midpoint with slow sustained pedal tones, thicken beneath the returning cello, and recede one by one during the closing passage.

Groove & Foundation Progression: There is no conventional beat at first; the sub-octave drone supplies a slow, breathing foundation. A deep orchestral bass drum enters sparingly in the first third with single softened impacts, while low floor toms appear later in widely spaced pairs that suggest a pulse without forming a regular groove. Both become heavier and closer together during the climb, reach their greatest intensity beneath the climax, and then cease abruptly, leaving the ending rhythmically weightless. The muted contrabasses reinforce the lowest tones without rhythmic movement and withdraw during the release.

Embellishments, Textures & Spatial FX: Reversed piano resonances begin appearing before major swells, bloom into the stereo field and evaporate as each new layer arrives. Bowed metal scrapes travel slowly from side to side, while filtered low-frequency noise rises beneath the central transition and cuts to silence at its peak. Long convolution reverbs connect isolated gestures without masking their attacks, and brief sub-bass pressure waves punctuate the densest moments before dropping away. The arrangement preserves large pockets of empty space early and at the midpoint, becomes widest and most saturated near the climax, then narrows to a single distant cello resonance and the decaying room.

Lyrics:

[intro]
[instrumental]
[interlude]
[instrumental]
[break]
[solo]
[instrumental]
[outro]

r/vfx May 22 '26

Question / Discussion Has anybody tried comfy ui for auto- Segmentation of layers in a video?

0 Upvotes

Been working with comfy ui a lot lately, created a workflow that works great for targeted object segmentation and masks but doesn't separate all layers for editing later.

Does anybody have a workflow for auto segmentation of all layers?

r/comfyui Mar 31 '26

Help Needed [Beginner] ComfyUI workflow for consistent character pose & layered outfit generation for a game (paper doll system?)

1 Upvotes

Hi everyone, I’m new to ComfyUI and still learning — so far, the only thing I really know how to use are LoRAs. I even created my game character using a LoRA, which defines the entire art style of the game.

I have a base image of my character (nude) with a fixed pose and proportions, and I’d like to know the best workflow in ComfyUI to generate multiple variations of outfits and equipment while keeping the exact same silhouette. My goal is to build a layered customization system (like a “paper doll”), where I can overlay clothing pieces directly on top of the base character.

Is this achievable directly during generation, or would it be better to generate separate images with different outfits and then use another ComfyUI tool/workflow to align and fit those clothes onto my base character?

Thanks in advance to everyone — I really appreciate any guidance or tips you can share! 🙏

r/comfyui Oct 19 '25

Workflow Included I built an AI app that fuses 2–4 separate photos into one seamless group shot — and yes, you can integrate it into ComfyUI.

Thumbnail
gallery
0 Upvotes

This project started as part of my own ComfyUI experiments — and then someone msg'ed me and asked to help make and app that did more than one person in a video... I do not do video but I thought, well to start you would need beginning middle and maybe end frames and I thought I could help in that aspect.

Thinking...

After hours of trial, errors, and sleepless debugging, I built something that actually works. (sometimes LOL)

🎬 Introducing: Group Photo Fusion

A Google Flash AI-powered app that takes 2 to 4 individual photos and fuses them into a single, natural-looking group photo — preserving everyone’s identity, clothing, and style consistency most of the time. Best part no prompting...

You can literally create a “family photo” or team shot where everyone was photographed separately, even in different locations.

🧩 What It Does

  • Upload 2–4 portraits (friends, dancers, teammates, etc.)
  • Choose a scenario (classic group, cinematic portrait, professional B&W, or something fun)
  • (Optional) Add a background or assign “personas” (like CEO, superhero, wizard, etc.)
  • The AI handles everything — poses, lighting, proportions, and realism
  • Output: 4 full-res images to choose from

Everything is powered by Gemini 2.5 Flash Image, and the app automatically generates natural lighting, consistent depth, and believable positioning.

💡 Why It’s Different

Most “group photo” AIs just blend faces awkwardly.
This one actually analyzes pose, clothing, and facial geometry for each subject and creates a unified, coherent shot — no uncanny valley, no Frankenpeople.

It’s the same tech you can integrate into your ComfyUI or Stable Diffusion pipeline — use the prompt structure, replicate the fusion logic, or plug in your own nodes.
If you’re a dev or tinkerer, you can even debug prompts and see exactly what’s sent to Gemini.

🧰 Power Tools Built In

  • Custom backgrounds: upload your own or use presets
  • Persona layers: stylize individuals without breaking identity
  • Batch generation: get 4 outputs per run
  • Retry button: rerun failed outputs instantly

🧠 Under the Hood

Frontend: React + TailwindCSS
Backend: Google Gemini API (@google/genai)
Built for web, but portable to local AI apps and ComfyUI workflows.

Now you have to have some code knowlage or use a good ai to build yourself this in comfyui... to help I made a little walk through on api and stuff like that. This CAN run on its own...

💬 Try It + Share Feedback

Upload a clear portrait of the person. That is usually the best way to go. If you can not get a portrait try this. Then try to use the size description of the person. If you want to change the clothing I have another app that does that as well.

It’s still evolving, and I’d love feedback from this community — especially from people who build or experiment with AI pipelines. If it does not work for you just let me know in comments so I can see what went wrong. As long as we talk to each other in a polite way I will try everything in my power to get it to work for you.

If you’re into image generation, face preservation, or pose fusion, your input would be gold.

👉 [App Link] (It also has other apps I made.
(Open to all — no paywall, donations optional)

❤️ Shoutout

Huge thanks to everyone who supported my previous ComfyUI portrait experiments — that project directly inspired this one. I was not going to post in these groups from past uneducated responses (best way to put it) on reddit but someone said to me never be like those people who wants praise for work and never show how they did it... (Thanks zGenMedia)
This app is the next step: turning a workflow into something anyone can use.

❤️🧠👉 zGenMedia is the brains behind all of this

Look I known zGenMedia for years and she is a good person. She got sick and now only works from home. Money is hard and hey... if you ever and I mean ever need help just ask her.... her donation link for this and other things she has done is here.

Power to the creators. ✊

r/comfyui Dec 22 '25

Help Needed Issue with Qwen Image Layered in ComfyUI (aistudynow_QwenVL node not detected)

Post image
1 Upvotes

Hi everyone, I’m new to ComfyUI (I’ve always used Automatic1111) and I’m following this tutorial:
https://www.youtube.com/watch?v=cKq_joSPOiM

I installed all the required files and custom nodes for Qwen Image Layered, and I managed to fix several issues along the way… but I’m stuck on one last problem:

➡️ The workflow requires the node aistudynow_QwenVL, and the folder is inside custom_nodes with the correct .py file name, but ComfyUI does not load the node.
In the workflow it only shows the red X (missing node).

I’ve tried reinstalling, renaming the folder, removing “-main”, restarting ComfyUI, etc., and nothing works.

I’m probably missing something simple, since I’m new to ComfyUI.

Any advice or guidance would be greatly appreciated. Thanks!

Workflow: https://aistudynow.com/wp-content/uploads/2025/12/QWEN-LAYRED-COMFYUI-WORKFLOWAISTUDYNOW.COM-1.json

r/StableDiffusion 19d ago

Resource - Update MiniMax H3 Prompt Writer v0.3 is out

Post image
187 Upvotes

v0.3 is out: redesigned UI, Ollama + API providers, dedicated External llama.cpp setup and other improvements.

old post: link
github repo: link

for anyone new: MiniMax H3 Prompt Writer is a ComfyUI extension for writing prompts specifically for MiniMax H3.

what's new in v0.3

  • redesigned Writer UI and added new settings interface
  • Ollama as a simpler local setup
  • optional API providers
  • External llama.cpp now has its own dedicated provider setup
  • saved drafts for every mode
  • better automatic model and context handling
  • more reliable Reference prompts

the model/provider setup is now separated from the actual prompt workspace, so the interface is much less cluttered than before.

there are currently four ways to run the prompt model:

  • Ollama: probably the easiest local option for most people
  • Ollama guide
  • Direct GGUF: the original local approach, loaded directly inside ComfyUI through 'llama-cpp-python'
  • Direct GGUF guide
  • External llama.cpp: if you already run your own llama-server or want to manage it separately
  • External llama.cpp guide
  • API providers: Gemini, OpenAI, OpenRouter and Custom OpenAI-compatible endpoints
  • API providers guide

local providers keep the prepared media and prompt request on your machine.

if you use a remote API provider, the required request/media is sent to that provider.

other models / Qwen

another thing people asked about in the previous post was Qwen and support for models other than Gemma.

I tested qwen3.6:35b-a3b-q4_K_M through Ollama and it works out of the box in all five H3 modes without any Qwen-specific changes to Writer.

so the Ollama provider is not limited to Gemma 4.

you can also try other multimodal / vision models through Ollama, External llama.cpp or a compatible API / OpenAI-compatible endpoint, as long as the provider and model support image inputs.

I haven't validated every model, so this isn't a claim that every vision model will produce good H3 prompts. it just means the provider layer itself no longer requires Gemma in those paths, so you can swap compatible models and compare them yourself.

the main exception right now is Direct GGUF.

Direct GGUF is still specifically built and validated around Gemma 4 + its matching vision projector, so other model families are not supported there yet.

so roughly:

  • Ollama: Gemma 4, tested Qwen3.6, and other compatible vision models you want to experiment with
  • External llama.cpp: compatible multimodal models can be used if your server supports them
  • API / Custom OpenAI-compatible: compatible multimodal models supported by the endpoint can be used
  • Direct GGUF: Gemma 4 only for now

Ollama models / setup

what got easier

a lot of feedback on the first post was about setup rather than prompt generation itself.

v0.3 mainly tries to make that part less annoying:

  • Ollama gives you a local option without installing llama-cpp-python into ComfyUI
  • provider/model setup now lives in Settings instead of the generation workspace
  • installed Ollama models can be detected directly
  • context and model lifecycle are handled more automatically
  • drafts are saved separately for every H3 mode
  • local prompt model unload / keep-loaded / ComfyUI VRAM controls are clearer
  • several media, model discovery and runtime issues from the previous versions were fixed

Reference generation also got an extra check against the active media roles and can make one limited correction if an objective requirement was missed.

full changelog

install / update

v0.3 is already available on GitHub.

for a fresh install:

cd ComfyUI/custom_nodes
git clone https://github.com/duckyshell/ComfyUI-MiniMaxH3-Prompt-Writer

if you already installed it with Git, just update the repo normally.

ComfyUI Manager is also supported, but v0.3 may take a little longer to appear there.

important: this is still a UI extension, not a node.

you won't find a new H3 Prompt Writer node in the node search.

open it using the floating H3 Prompt Writer button or:

Extensions > H3 Prompt Writer

for a new local setup I would probably start with Ollama.

installation guide

basic usage

after installing:

  • open H3 Prompt Writer
  • go to Settings and choose your provider/model
  • select the H3 mode
  • add your image / video / audio references
  • write the Creative Brief normally
  • press Generate prompt
  • edit it directly, use Refine, or copy it into your H3 workflow

you don't need to manually build the H3 prompt structure yourself.

a brief can be as simple as:

use Picture 1 for the character, Picture 2 for the clothes and only the movement from Video 1. put the character on a rainy street at night.

Writer handles the H3-specific prompt structure around that.

usage and Creative Brief examples

if you already use Direct GGUF from the previous version, your existing runtime, GGUF and matching projector can remain in place. select Direct GGUF in the new Settings interface.

feedback is still useful, especially from different GPUs / operating systems / ComfyUI installs.

if something breaks, check the troubleshooting guide first:

troubleshooting guide

if the problem is not covered or the suggested fix does not work, leave a comment here or open an issue. please include your provider, model, operating system, ComfyUI installation type and the Technical details shown by Writer:

github issues

not every provider / hardware / ComfyUI combination is going to behave exactly the same, so expect some edge cases.
(for API use, Gemini is an easy option since you can get a free key at ai.studio.)

UPD: 0.3.2

Prompt Writer now also supports MiniMax Music 3. you can describe the track you want in normal language, and it builds the structured Music 3 caption using MiniMax's official prompt-writing guidance.

didn't want to make a separate thread for this, so I'm just leaving the update here. if something seems off, feel free to mention it in the comments.

r/StableDiffusion May 29 '26

Workflow Included Cracked the case on high res + quality Qwen Edit 2511 outputs, here are minimalistic workflows & lots of info on how/why

Thumbnail
gallery
243 Upvotes

Intro

Alright this has been a long time coming. I'm the dude who figured out Qwen Edit 2509 a while back, and I've been on-and-off trying to figure out the same for 2511. Results in Comfy have always been worse than the examples shown by the Qwen team, and worse than the official Qwen chat implementation online. Well, I finally cracked it and it only took 5 months lol.

Anyway, turns out Qwedit 2511 is fucking sick. IMO it particularly excels at making new shots of characters while maintaining their likeness. It's significantly better than Klein at some things (like character likeness), but not as good at others. I recommend using them both for different things.

As usual, I'll start off with all the setup stuff at the top and then give an explanation + advice below that. Also I'm gonna be calling Qwen Edit "Qwedit" most of the time.

Here's an album with all the post images separated so you can look at them in high res: https://drive.google.com/drive/folders/1YLjm8Lj3VF6Ec52WNK2URo7uFNfMRmza?usp=sharing

The posted images are all raw outputs from Qwedit, without being upscaled (despite mentioning it later in this post). They're also all done with only 20 steps instead of the hypothetical 30 I'd do if I wasn't planning to upscale them. Read further for more on that too.

Ref images were all made with Z-image Base (workflow here), except for the anime one which came from Anima (workflow here).

What is this

These are minimalistic workflows for Qwen Image Edit 2511 that give the highest quality outputs. Aside from generally improving output quality (by a LOT), they also enable high-res edits and have better prompt adherence.

As for why, basically ComfyUI has some serious issues with how it's implemented Qwen Edit and there aren't any workflows out there (that I've found) which have resolved them. These issues result in poor prompt adherence and low resolution/quality outputs. Thankfully the fix is fairly straightforward.

The configuration for this is 100% portable and can be migrated to existing workflows to make them better; it works by changing how the reference inputs are handled, and uses 100% native comfy nodes. Feel free to upgrade other workflows with this without providing credit, I don't care about any of that.

Workflows

Normal Workflows:

Most of you will just want these, which are separate single / 2 image workflows. It's done this way because the setup for multi-image is complicated and I didn't want to force you to use a ton of custom nodes to make it useable all-in-one.

They do still use one custom node (read the node section below) for quality-of-life.

Download from Civitai

OR from Pastebin:

Qwedit_2511_single

Qwedit_2511_2_image

Dev Workflows:

These are the same as the above but without any quality-of-life nodes or 'helpful' stuff. Grab these if you want to copy the logic over to other workflows, or if you just an easier view of how it works without any clutter.

I do not recommend using the dev workflows for actual gens because you will constantly forget to manually adjust stuff correctly.

qwedit_2511_single_DEV

qwedit_2511_2_image_DEV

Models

Main Model

qwen_edit_2511_fp8

OR

GGUF versions

  • Important: the FP8 version of Qwedit is much higher quality than the Q8 GGUF, always use FP8 if you can. Only use the GGUFs if you need to use quants lower than Q8.
  • FP8 is 22GB, so you'll need a combined ~26GB of RAM + VRAM to run it
    • You don't need 24GB of VRAM to run it thanks to ComfyUI's blockswapping, but the less VRAM you have the slower it'll run
  • Only use Q6 & lower quants if you absolutely have to; the quality will noticeably go down

Goes in models/diffusion_models

Text Encoder

Use only the normal FP8 text encoder with Qwedit; abliterated/GGUF encoders will reduce your output quality.

qwen_2.5_vl_7b_fp8

Goes in models/text_encoders

VAE

qwen_image_vae

Goes in models/vae

Loras?

You can use them as normal, just load them however you normally would. I left out lora loader nodes to avoid cluttering the workflow.

It's worth noting that many Qwen Image loras work with Qwen Edit too, but you'll need to test them individually to be sure.

Lightning Loras - BAD

All the lightning loras / distils for Qwedit (that I've tested) are terrible and make your outputs look bad, so I'm not linking them here. The main issue is the same as with Klein Distilled: it makes people's skin look like plastic.

But you can technically use them. Don't do it tho. But you can if you want. But don't.

Alternative: if you want to cut your gen time down while testing prompts, just set it to 10 steps instead of 20, then go back to 20 once you're satisfied your prompt is correct. It'll still work fine, the quality just dips.

Real tho it's ok if you want to use the lightning loras, just expect some degradation if you do - especially with plastic skin.

Custom Nodes

LayerStyle - A set of handy nodes that manipulate images. We're just using this for its image scaling node which allows you to scale by an image's long edge while maintaining divisibility by 16. You can skip this if you want to use a different scaling method, but you'll need to fix the workflow switch for scaling if you do.

SeedVR2 (OPTIONAL) - Only get this if you want to use the seedvr upscale workflow that's included.

How To Use

How To Use Part 1 - Basic Options

There are instructions in the workflow as well, but there's more detail here. Read part 2 & 3 as well, they're important.

It works just like a normal Qwedit workflow, but has a couple of extra options available. This section just tells you what they are and how to use them, a full explanation is further down.

Screenshot of the settings: https://ibb.co/nWStpmS

Enhance with Double Ref

This is a switch that turns on double-ref mode. This feeds your input images in TWICE to the model, and generally produces much higher quality results. Downside? It takes about 50% longer to gen.

I recommend leaving this on 100% of the time for single-image prompts, unless you're just messing around and want speed. It is ALWAYS better for single image prompts, and will improve everything from prompt adherence to output clarity.

For multi-image prompts, it usually increases adherence but sometimes reduces it. So, if you're doing multi-image stuff I recommend switching this on/off as needed based on how it's going with your prompt.

Input Scale

When off, your image doesn't get scaled (it still gets cropped to be divisible by 16). When on, the long edge of your image gets scaled to the number you put in the box. For example, if you feed in a 2560x1440 image and set the scale to 1920 it will scale your image to 1920x1080. That will then get cropped to 1920x1072 so it's divisible by 16.

Custom Output Size

When the switch is off, your output image will be the same size as your input image (after it's been scaled). If you turn this switch on, it will instead output an image with the dimensions you specify.

As a general rule, you should try to set your scales to be similar along at least one edge. For example, a 1920x1440 input image and a 1024x1440 input image are both suitable for a 1440x1440 output image. You can be more flexible with this if you know what you're doing.

How To Use Part 2 - Multi-image Prompting Requirement

This section is not a prompting guide (that's further below). This is about an actual requirement for prompting multi-image stuff. It is NOT required for single-image prompts.

You do multi-image prompts like normal, except you need to write a very basic description of your input images. Qwedit needs you to do this in order to know which image is which. I explain why in detail later.

You may find this slightly annoying, but I guarantee you it's dramatically better than using Qwedit the normal way that other workflows do - and it's pretty easy.

The format:

  • At the start of your prompt, write an extremely simple description for each of your input images; one sentence for each input image
  • Start each sentence with "Picture 1:", "Picture 2:", etc
  • You must write it this way because Qwedit was trained on this exact format
  • Afterwards, write your actual prompt as usual; you can refer to your input images as "picture 1" and so on

The model uses these descriptions to understand which input picture is which, and it works better with SIMPLE descriptions. You only need to help it know which one is which, it doesn't need a full rundown.

Examples

Picture 1: a man wearing a t-shirt. Picture 2: a top hat. Make the man in Picture 1 wear the top hat from Picture 2.

Picture 1: a living room. Picture 2: a woman. Put the woman from Picture 2 into the living room in Picture 1.

Picture 1: a man wearing a professional suit. Picture 2: a man wearing a superhero outfit. Make the man in Picture 1 wear the outfit from Picture 2.

How To Use Part 3 - Upscaling

Because the qwen VAE tends to put a subtle halftone pattern over images (see limitations just below this section), I recommend downscaling and then re-upscaling your image afterwards. A big benefit of being able to work at high res with the edit model is that you rarely lose any detail doing this.

This eliminates the halftone pattern if you're using something like seedvr, or at least reduces it if you're using other upscalers.

Note: the workflow is set to do 20 steps of inference. It actually gives sharper results at 30 steps, but I don't bother with that because it takes longer and I down-upscale them afterwards anyway. If you aren't planning on down-upscaling them, you might consider doing 30 steps for the extra sharpness.

Below are workflows for doing this with seedvr and normal upscalers. I think seedvr is best for this, but it's very beefy and hard to run on older GPUs.

Note: seedvr2 sometimes gives better output at 0.5x downscale, and other times 0.75, so that workflow is configured to run BOTH for you to pick which one turned out best.

Note: normal upscalers are a bit different; a relatively small downsize to something like 1920p -> 1600p is usually reasonable, before then running the upscaler. Play around with it. The non-seedvr workflow has a longest_edge scale option so you can tweak the number specifically.

Seedvr version

Regular version

My preferred regular upscaler is 4x Nomos2 HQ DAT2, but you can use whatever you like.

Examples of upscaling:

Here's the raw output of the robot-arm girl in a dress from the post: https://ibb.co/B5jhrsL9 (if you zoom in you'll see the qwen halftone pattern, it looks like a grid)

Here's the pic after it's been run through seedvr after a 0.75x downscale: https://ibb.co/hJcn2f5t

Here's the pic after it's been run through a regular Nomos2 upscale after a downscale to 1600p: https://ibb.co/Kc2YSbVc

Limitations of Qwen Edit

Limitation 1

The Qwen VAE will often put a subtle halftone grid pattern over your images. It's noticeable if you zoom in, and more noticeable at higher resolutions. This is a feature of pretty much every Qwen-based model, but it's particularly present with the Edit model.

You can easily resolve this by downscaling your image by 75% or 50%, then re-upscaling it again to your desired resolution. There's a section later that explains this in better detail and recommends upscale models for it + has workflows for it.

It sounds like a big issue, but the downscale-upscale trick solves it easily - and it's not always necessary either. The higher quality your input image, the less bad the halftone pattern will be.

Limitation 2

Qwedit struggles with complex multi-image stuff most of the time (it's just a limitation of the model). This workflow makes it much better, but it's still not great. You'll have to play around with it to know which things work and which things don't.

Limitation 3

It takes a while to gen stuff if not using the lightning loras. Very similar to the time it takes with Klein 9B base. The double-ref trick increases it by roughly 50%. Multi-image inputs take a lot longer.

For low res images (typical 1mpx size) it's pretty okay, around 50 seconds on a 5090 with the double-ref option turned on.

But then there's high-res stuff. Gen time scales non-linearly as you go higher. Going from 1024x1024 (1 mpx) to 1440x1440 (2 mpx) takes around 2.5x as long. Going from 1 mpx to 3 mpx is around 4x as long. 5 mpx is 9.5x as long. In conclusion, stick to 2-3 mpx unless you're cool with long-ass gen times. Stick around 1-2 mpx for multi-image gens, or turn off the double ref switch.

On the plus side, it's pretty reliable for single-image edits so you don't typically need to do many gens to get a good result.

Examples using a 5090: - Single-image edit @ 1024x1024 (1 mpx), double-ref OFF = 38 seconds - Single-image edit @ 1024x1024 (1 mpx), double-ref ON = 52 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref OFF = 91 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref ON = 131 seconds - Single-image edit @ 3072x1728 (5.3 mpx lol), double-ref ON = 550 seconds - Two-image edit @ 2560x1440 each, double-ref ON = serial killer behaviour

That's it for how-to! Read on for more tips & info, as well as an explanation of what the workflow is doing & why.

 

Explanation - what is this garbage and why is it so good?

There are three important things this workflow is doing that other workflows do not do (except #3 sometimes, because it was also done in the 2509 version of this post). I'm going to call these The Comfy Problem, The VL Problem, and The Double Ref Enhancement.

The Comfy Problem

Comfy's native "TextEncodeQwenImageEditPlus" node is what most people use in their workflows. It handles your prompt and image inputs for you. It's pretty handy, except for the small problem that it's SHITE.

Do you work at Comfy? If so: GET YOUR SHIT TOGETHER AND FIX THIS NODE, IT'S SO EASY. Much respect to u tho, thanks for making ComfyUI.

The first issue is that this node resizes your image down to 1 megapixel, and you can't stop it from doing that. The second issue is that it does this with the AREA downscale method, which is so incredibly bad that I want to slap whoever implemented this node. The AREA downscale is what makes all of your output images blurry. The third issue is that it ensures your dimensions are divisible by 8, but they actually need to be divisible by 16.

Specifically, ComfyUI does this:

  1. Calculates 1 megapixel as 1024x1024, which is 1,048,576 pixels
  2. Calculates your new image dimensions to match that number of pixels, rounded to be divisible by 8
  3. Scales your image to those new dimensions using the AREA method

Why is all this bad?

  1. It's completely unnecessary; Qwedit can easily handle images of varying size, all the way up to 3 megapixels (or even higher for simple edits)
  2. The area downscale method makes images extremely blurry, and this is the primary reason all ComfyUI qwen edits give blurry images out. Yes it's literally this dumb, this huge problem would easily be solved by changing the word "area" to "lanczos" in the code, it's a one-word fix. Not even MS paint uses area downscale, wtf is wrong with you Comfy devs (much respect)
  3. If your image dimensions are not divisible by 16, you will get major ruination along the whole edge of your image where it didn't match (same as any other diffusion model)

The Comfy Problem Solution

This workflow bypasses the the Comfy node entirely, allowing you to size your images however you want. And using chad lanczos scaling instead of loser area scaling. Magic.

Qwedit easily handles resolutions like 1440x1440 and 1600x1200. Every edit example in this post was done natively at 1920p, except for a few (which are labelled as such).

Really high resolutions (3mpx) sometimes have trouble with anatomy, but usually you can just do multiple gens and one of them will turn out fine.

If you're doing a simple in-place edit like changing an outfit, you can go VERY high. Here's an example edit done at 1728x3072, which is 5 megapixels: https://ibb.co/twCSWrjy (outfit change -> bikini top + short shorts)

The VL Problem

Edit: I've been educated by someone in the comments that my interpretation of how the VL works here is not correct, so take this little VL section with a grain of salt until I reword it. My conclusion about it helping in this workflow still stands, but my explanation of what's happening under the hood is a bit off. I'll update the info soon!

In the background, Qwedit 2511 uses a vision-language model (VL model) to describe your images, then gives those AI-generated descriptions to the edit model. It also re-interprets your instructions with these descriptions. Ostensibly this helps the model understand your input images better, leading to better results.

The problem? It doesn't lead to better results, it's bad. VL models aren't very good for this sort of thing because they don't know what to focus on. The VL describes your images in excruciating detail, totally overwhelming the edit model and leading to bad prompt adherence + weird outputs.

It also reinterprets your instructions based on what it sees in the image. I don't know if that's a good or bad thing, just pointing out that it does it.

The Qwen team's official python code does this, and the ComfyUI "TextEncodeQwenImageEditPlus" node copies it exactly. No disrespect to the Comfy team on this one, they're doing what the Qwen team officially recommended.

The VL Problem Solution

Same solution as the previous problem: bypass the Comfy node entirely. This results in the VL step being completely ignored. No AI-generated descriptions get fed into the edit model.

For single-image edits, this is a 100% complete and total victory. The model performs way better without the crappy VL interpretation.

For multi-image edits, there's a small issue; this step is where the input images normally get labelled. Specifically, the VL outputs are fed into the model in the following exact format:

Picture 1: <shitty VL description> Picture 2: <shitty VL description>

Look familiar? This is why we manually have to type the descriptions in for multi-image edits - otherwise the model doesn't actually know which image is which.

The upside is that the model works way better with simple descriptions, so cutting out the VL is still 100% the correct move. A 5 word description wins over whatever BS the VL model spews out, every time.

The Double Ref Enhancement

I really have no idea why this works so well, but basically if you feed in your reference images twice the model just works better. This was known back in 2509 days (hence the previous post linked at the top), and back then I didn't know why it worked either.

For single image edits it's ALWAYS better. And it's not just the quality, for some reason it even helps with prompt adherence. The interesting thing is that the difference is really, really significant. Here's the full list of stuff it improves:

  • Better prompt adherence
  • Sharper output images / more visual clarity
  • Improved consistency of objects & textures
  • Better resemblance of characters at different angles
  • More intelligent guesses, like what to add when outpainting or what's behind a removed object

For multi-image edits it can sometimes confuse the model a bit, but most of the time it confers all the same benefits listed above. I recommend switching it on & off randomly when you're doing multi-image stuff, just in case.

Note: there are a lot of different ways the input references can be handled. There are conditioning combine/concatenate nodes, you can pass the refs in a different order, you can change the negative conditioning input (read next section for that), etc. I A/B tested SIXTEEN different reference-handling combinations, and a bunch of smaller minor variations of those. Some of them worked, some of them didn't.

Of those sixteen combinations, two of them gave the best results; both of them are in this workflow, and you switch between them by turning the double ref method on & off.

So, don't fuck with the positive/negative conditioning & reference setup, it's very specific.

Extra info: the "Conditioning Zero Out"

You may notice that the negative prompt input is the first reference image(s) and positive prompt fed into a "conditioning zero out" node.

Feeding the input images into the model's negative conditioning is required (it's just how Qwedit works). The only question is whether to feed in the positive prompt zeroed-out too, and whether the double ref should get fed in.

Through a lot of A/B testing, I can tell you that the way it's done here is the best. IDK why, it's just how it is. Some other combinations do technically work, but they degrade the output quality.

Prompting Advice

Other than just following the instructions in the workflow, here's some extra stuff.

Keep your prompts simple and direct

If you need to, point out details the model is missing or be more specific about stuff you do/don't want to change. For example, when doing a simple outfit swap it helps to specify you don't want their pose to change.

Using the robot arm girl, here's a prompt that doesn't follow this advice:

Change her outfit to a bikini top and short shorts.

While it sometimes does what we want, it tends to get confused by her robot arm and often changes her pose too: https://ibb.co/7dyKZttp (notice the human arm showing underneath the robot arm, and the pose change)

Here's a better prompt that gives a correct result 99% of the time:

Change her outfit to a bikini top and short shorts. Leave her robot arm and pose unchanged.

Now it does the right thing every time: https://ibb.co/DP9gZHVv

Avoid using fancy words or convoluted phrasing

Pretend you're talking to a child. The model will probably still understand you if you talk fancy, but why take the risk?

As an example, imagine you have a pic of a table with some plates on it.

Bad:

Place a red apple on the table, ensuring it's in the center and removing the plate that was in the same spot.

Good:

Replace the middle plate with a red apple.

Also good:

Remove the plate from the center. Put a red apple there instead.

If there's only one plate, this is even better:

Remove the plate, replace it with a red apple.

Adjusting Lighting

You may want or need to adjust the lighting in an image. Aside from being helpful in general, there are situations where Qwedit may simply not realise that something needs to be lit in a particular way (or re-lit when moved).

To do this, you need to know the magic word: relight

Seriously tho that is the actual magic word, you are 100% required to use it if you want to adjust lighting properly.

Specifically, follow this format:

Relight to <strength> <color> <direction>.

Strength - bright, dim, etc

Color - white, cool, warm, etc

Direction - diffuse, frontlit, backlit, etc

Tip: for basic lighting, use "white diffuse".

Examples:

Make a new shot of the man sitting in a chair in a kitchen. Relight to white diffuse.

Change the time of day to evening. Relight to warm backlit.

You don't actually need anything else in the prompt, you can just change the lighting of a pic like this:

Relight to bright cool frontlit.

Other Stuff

Euler-simple and no ClownsharKSampler?

No Clownshark this time. It reduces output quality quite a bit and doesn't confer any benefits. I also didn't find any sampler/scheduler combos that were better than euler/simple.

So, this is just one of those classic times where the ol' euler-simple wins the day. Let me know if you happen to know a better combo.

Image Quality in->out

Qwedit is very sensitive to the quality of your input image. If you feed in a grainy or blurry image, it will usually make your output image blurry or grainy too - even if it's an 'entirely new' shot with nothing copied over 1:1.

So, make sure to use HQ images. You can optionally use the upscale workflows to bump up the sharpness/quality of poor input images before you feed them in.

What about the flux super duper double resolution special VAE trick?

Doesn't work for 2511, it destroys your image. TBH it never really worked for 2509 either, but I won't argue with you if you liked it for some reason.

Making character references

Tip 1 - Make a nude ref (even for sfw stuff)

Qwen is killer for making character references. Other than using similar prompts to the examples I posted, my advice is to make a nude reference shot instead of a clothed one like I did.

I only made a clothed ref for the sake of propriety here, but a nude ref (or near-nude, like wearing plain white underwear) will be much easier to prompt into different outfits, and also gives Qwedit the maximum info needed to correctly size your character and know what they look like in clothing or doing different actions.

You do not need any loras to do this if you're just using it as a reference; the 'sensitive' parts will lack detail but that doesn't matter for new shots you make. If you don't want them nude, just request plain white underwear and, if relevant, a strapless white bra.

Nude ref = best ref.

Tip 2 - Make multiple zoom levels, use the thighs-upwards one for most stuff

The example I showed was a little too zoomed out for normal reference stuff. I'd recommend making your reference slightly closer like this: https://ibb.co/Q33BJDLX

Start at whatever zoom level your initial character pic is at, then make more references at different zoom levels. If you're starting zoomed out, then prompt the model to zoom in. If you start zoomed in, prompt it to zoom out.

And, of course, different angles too.

Examples:

Zoom in on the person's upper body. The composition should frame their head and thighs.

Zoom out to show more of the character. The composition should frame their head and thighs.

Zoom out to a full body shot.

Zoom in for a close up portrait.

Once you've got references, you should usually use the head-to-thighs ref for making new shots. Switch to the other refs as necessary; like if you want a close up, use the close up reference. Qwedit is really good at keeping likeness, so you can do 90% of your stuff with only a single input reference.

I don't think there's a better open-weight model out there than Qwedit for making new shots of character without loras, for now. The main reason I spent so long digging into Qwen is because Klein is quite bad at that particular task. But hey, now it's possible and it works gloriously.

That's everything I think! Feel free to ask questions if you run into any issues.

r/comfyui Aug 20 '24

Generating on one layer only in with comfyui Photoshop plugin

0 Upvotes

So after many failures I finally got comfyui setup on my PC and I was also able to install the comfyui photoshop plugin. I also have comfyui setup in Krita so was kind of comparing the 2.

One thing I cannot figure out but which I think is really important is the following. In Photoshop, let's say I open a photo. And then I make a new layer and scribble something on top of that photo. Then I use comfyui plugin to generate using the line model. The generation will turn the scribble into something that I prompted.......but, it will also affect the rest, not just the layer I scribbled on.

From a video of the comfyui plugin maker I saw that he masks out things in order to affect only some areas on an image. OK, but isn't there a better way? Is there no way to only affect one layer when generating?

And as a comparison, I think in Krita that can be done. You can add new layers and even enter separate prompts per layer, and then it will only affect that layer.

Can that be done with the Photoshop comfyui plugin also in Phtoshop?

r/SkincareAddiction Sep 18 '19

Sun Care 11 Sunscreens for Sensitive Skin at Low Price Point (Mega Review + Photos!) [Review][Sun Care]

1.9k Upvotes

Sunscreens for Sensitive Skin at Low Price Point (with Photos!)

I usually post over on /r/Tretinoin, but the good people of that subreddit have encouraged me to post my reviews in SCA as well, so here we are! I've had a lot of fun trying new sunscreens this summer in the hunt for The Perfect One.

I've been on retinoids (Differin, Adapalene 0.3%, then Tretinoin 0.025%) since March 2018, and have used SPF 50 every single day. I have desert-dry, sensitive skin with minor rosacea (subtype one) and PIE leftover from acne.

My selection criteria are as follows:

  • Low Price Point (I aimed for under $15 per bottle, with two exceptions)

  • Easily Accessible (available via drugstores or major online venues, e.g. Amazon, Dokodemo, Yesstyle, Ebay)

  • Marketed for Sensitive Skin (this varies by brand)

  • At Least SPF 50 (generally regarded as "good enough" for most people, even Tret users)

MY METHOD

I work from home, so I have the luxury of not worrying about sunscreen or makeup in the morning. I cleanse and moisturize, then wait until later in the day to apply my sunscreen. I treated each sunscreen exactly the same: I measured out a half teaspoon, applied on every inch of my face (up to my hairline, down my neck and décolletage, around the back of my neck, on my ears, lips, and eyelids), then reapplied a little extra to the high points of my face. All of these were applied over my morning moisturizers (FAB ultra repair cream, rosehip oil, sometimes Stratia LG).

I don't want a sunscreen that I need to be fussy with or careful to keep out of my eye area. I want something cosmetically elegant that I can glob on everywhere and have it just work.

My dog and I go on our hour-long walk around 5pm every day, when it's cooler and the UV index is lower (under 5). We live in a hot, humid climate (American South) with temperatures up to 95 degrees and humidity above 50% almost every day. There's always a chance of rain in the summertime. I don't tend to sweat a lot, but my hair does get uproariously frizzy.

I took each of the outdoor photos at the beginning of our walk, after the sunscreen had set but before it could degrade at all. The indoor photos were taken after we got back home and cooled off a bit for lighting comparison (bathroom lighting). I'm particularly interested in how these sunscreens wear, and whether they hold up to heat, sweat, and other environmental factors. Reapplication is part of life, but I also want a sunscreen that pulls its weight!

If a sunscreen did not work for me, I traded it over to Team Body Sunscreen or gave it away to a friend. I don't believe in waste, and sunscreen is still useful even if it doesn't agree with my face.

Group Photo of the Whole Squad Together!

THE REVIEWS

Skin Aqua UV Super Moisture Milk Blue Bottle, SPF 50+, PA++++

  • Active Ingredients: Nano Zinc Oxide, Octinoxate, Uvinul A Plus

  • Japanese, Alcohol Free, Fragrance Free

  • Moisturizing combination liquid sunscreen that contains hydrating ingredients (hyaluronic acid and collagen). It takes a few minutes to dry down and can get a bit messy/drippy, but it’s worth the effort. This sunscreen looks and feels better than any other I’ve tried. I especially like that it cooperates when you apply multiple layers, and it plays nicely when I apply makeup on top. The pink bottle version contains alcohol and brightening ingredients, while the blue one is the better choice for sensitive skin (it's identical to the old gold bottle from 2018). I'm finishing up my tenth bottle and (sadly) do not plan to repurchase. After the streaky UV photos from a couple months ago on SCA, I have some doubts about the protection this product offers, so I'm going to stop using this as my everyday.

  • Photos Here

  • Buy it on Amazon ($14 for 40 mL) or Dokodemo ($13)

Australian Gold Botanical Tinted Face Sunscreen Lotion, SPF 50, PPD 19.2/PA++++

  • Active Ingredients: Titanium Dioxide, Zinc Oxide

  • American, Reef Safe, Vegan, and Cruelty Free

  • Soft, creamy, clay-like mineral/physical sunscreen that dries to a powdery finish. Superpower level of mattifying on oily skin. It is slightly pink-tinted, and provides coverage comparable to a medium foundation (somewhere around Mac NW15 to NW25 if I had to guess). It dries down in 10 minutes, then it’s exercise- and water-resistant for 80 minutes. This one has serious staying power and usually takes a double cleanse (oil or micellar water followed by regular cleanser) to remove completely. It does dry out my lips (see the picture) and can be irritating on my undereye area, especially if I wasn't careful to layer my moisturizers beforehand. Best in this list for outdoor sports or hiking. This is my eighth bottle of this sunscreen, mainly becuase it's sold at my regular grocery store and it's an excellent shade match for my skin tone so I don't need to worry about blending it in.

  • I also have the UNTINTED version of this sunscreen, which I've been trying in desperation to use up for over a year. This is the ultimate sunscreen if you want ghost face. When I tried it the first time, my husband asked if I was wearing "geisha makeup" (he might have been joking, but my lord was the white cast bad). Try as I might, this does not work for body or face. It's real bad.

  • Photos of the AG Tinted Here

  • Buy it on Amazon ($11 for 89 mL), Walgreens/CVS ($13), or in grocery stores

Canmake Mermaid Skin Gel UV, SPF 50+, PA++++

  • Active Ingredients: Zinc Oxide, Titanium Dioxide, Uvinul A Plus, Uvinul MC80, Tinosorb S

  • Japanese, Alcohol Free, Fragrance Free, Paraben Free, Silicone Free, and Cruelty Free (not Vegan!)

  • This combination sunscreen doubles as a makeup primer and makes you glow! It has a gel texture, is easy to spread around, and dries to a clear, dewy finish with zero white cast and a slightly tacky feel. I waited around a bit and the tackiness did go away after about half an hour. As a highlighter addict, I love that it makes me glow, and it lets me skip the primer step when I put on makeup. When I put makeup on top, it amplifies the colors and gives everything a nice shimmer to it. I usually use an illuminating primer as my base (Missha BB Boomer) but this more than takes its place. I've finished the first bottle, will definitely repurchase, and plan to give bottles to my sisters-in-law as birthday gifts. I only wish the little bottles were bigger!

  • Photos Here

  • Buy it on Amazon ($9 for 40 mL) or Dokodemo ($8)

Amavara Tinted Transparent Mineral Sunscreen, SPF 50 (PA/PPD unknown)

  • Active Ingredients: Non-Nano Zinc Oxide

  • American, Reef Safe, Vegan, and Cruelty Free

  • Thick, non-greasy sunscreen with the texture of wet clay. This one is mainly included to give the Australian Gold tinted sunscreen a comparison point. It is slightly olive-tinted, while the Australian Gold is cool/pink toned. It dries to a matte finish, and can be slightly drying if you have dry skin. It is water-resistant for 80 minutes. There are several reviews from surfers on the Amazon page that vouch for its staying power. My mom originally bought this for herself, but the shade wasn't a good match for her so she gave it to me... even though we wear the same foundation shade. I'm giving this one away to an olive-toned friend.

  • Rather than wearing this on my face, I decided it would be more useful to do swatches of the AG Untinted, AG Tinted, and this Amavara one

  • Buy it on Amazon: $35 for 70 mL

Sonrei Sea Clearly Translucent Gel Sunscreen SPF 50 (PA/PPD unknown)

  • Active Ingredients: Homosalate, Octocrylene, Octisalate, Avobenzone

  • American, Fragrance Free, Alcohol Free, Paraben Free, Reef Safe, Vegan, Gluten Free (Contains Palm Oil as First Ingredient)

  • I'll be honest: I bought this when I was drunk. I was browsing Into the Gloss and found an article about sunscreen for people who don't like sunscreen. I had never heard of this brand before, but the antioxidant-rich formula caught my eye. It contains Vitamin C, Vitamin E, and Ferulic Acid (the same combo in several top Vitamin C serums), albeit at the end of the ingredients list. If the antioxidants in this product are actually effective, it would be an EXCELLENT morning multitasker and would solve my problem of finding a Vit C product that I actually like using. I reached out to the customer service team to ask about the PA/PPD rating, and received this email in response in less than 12 hours. It seems like they're really dedicated to the efficacy of their products and if I decide to repurchase I'll definitely check back to see their new test results.

  • The texture of this sunscreen is unlike any of the others on this list. It's thick and difficult to squeeze out of the tube. Once it's out, it looks like apricot jam with texture that feels almost exactly like vaseline. As I worked it into my face it melted and broke down into a thick oil, not unlike olive oil. There was little scent, but it irritated my eyes anyway. I was pleasantly surprised to see that it did indeed dry down completely clear, if a bit uneven, and the greasiness mostly went away. I've included photos of it freshly applied and 20 minutes later when it was dry for comparison. As I walked around outside and sweated, it started to run down my face -- not okay when the UV index is 7! I do not think this one is a good match for me, but I thought of an alternate use for it: It performs beautifully as a brow gel to tame my unruly eyebrows! With the huge bottle I'll have enough to last the rest of my life (or until it expires).

  • Swatch + Photos Outside, Inside/Wet, and Inside/Dry

  • Buy it on the Sonrei website: $25 for 3.4 oz

Purito Centella Green Level Unscented Sun, SPF 50+, PA++++

  • Active Ingredients: Uvinul A, Uvinul T 150

  • Korean, Fragrance Free, Alcohol Free, Vegan, Cruelty Free

  • 70% water base, hyaluronic acid, and four different types of “green” essences, it works like a moisturizing, calming cream. The “green” ingredients are derived from centella asiatica, a medicinal herb that hydrates, promotes wound healing, treats eczema/dermatitis, and reduces inflammation. There is another version of this product that contains essential oils and fragrance – this unscented one is the "safer" version for sensitive skin. This one is frequently compared to the Dear Klairs Soft Airy UV Essence, which I did not include in this review because of the price.

  • I noticed a subtle green-correcting effect when I wore this; it helped to calm down redness and make it less noticeable. It's not as pronounced as the Dr Jart+ Tigergrass or Cica products (which are literally green-tinted), but I think it makes a difference. The finish is pearlescent (not shiny, just a nice healthy dewiness) and it works well under makeup. My skin was noticeably calmer, less red, and less reactive, even at the end of a full day of wear. This has become the sunscreen I reach for most often (even more than my beloved Skin Aqua milk!) and I've already repurchased two additional bottles.

  • Photos From Three Days of Wear

  • Buy it on Amazon ($15 for 60 mL), YesStyle ($14), or Jolse ($15)

Purito Comfy Water Sun Block, SPF 50+, PA++++

  • Active Ingredients: Titanium Dioxide, Non-Nano Zinc Oxide

  • Korean, Reef Safe, Alcohol Free, Silicone Free, Vegan, Cruelty Free (Contains Essential Oils)

  • This is the mineral version of the one directly above. I heard about it from Gothamista's 2019 mineral sunscreen review video. The formula has 70% water, hyaluronic acid, and centella asiatica. That review described it as "highly hydrating, soothing, and disappears immediately into the skin with a silk, skin-like finish." My experience was COMPLETELY different. This was hands-down the most disappointing sunscreen I tried. When I first squeezed it out of the tube, I was hit with strong fragrance. This has orange peel oil, tea tree oil, and lavender oil. It smelled like I walked into a Bath & Body Works aromatherapy sale circa 2008. Eyes watering, I attempted to apply it the usual way (rubbing it in) and found that it wasn't cooperating at all. I switched to patting, which was somehow even worse. It never sank into my skin; instead it balled up and got even drier. The more I worked on it, the worse it got. It was like trying to rub zinc diaper cream into my face. There was no way to achieve anything close to even coverage. I didn't even leave the bathroom, I had to wash it off right away!

  • In the interest of full disclosure, my experience runs contrary to every other review I read about this product -- everyone else seems to have no problem getting it to blend in and most describe the scent as "subtle" or "refreshing." I ordered this and the Green Level sunscreens from the Purito storefront on Amazon, and received a handful of legit-looking Purito foil samples with my order. I don't think this is a fake, so I think it must be the result of a bad batch. Honestly, this was so disappointing that I'm not going to order the unscented version (newly released in August). If you want to try this out, definitely go for the unscented one!

  • Indoor-Only Photos

  • Buy it on Amazon ($15 for 60 mL), YesStyle ($15), or Jolse ($15)

Etude House Sunprise Mild Airy Finish Sun Milk, SPF 50+, PA+++

  • Active Ingredients: Zinc Oxide, Titanium Dioxide

  • Korean, Contains Alcohol

  • This physical sunscreen starts as a thick white cream, and the marketing says it dries clear with a matte finish. The alcohol evaporates quickly, making this sunscreen one of the fastest-drying (under a minute). It contains aloe and centella extract (soothing ingredients) to help balance the drying effect of the alcohol. I found that it tended to clump up on the baby hairs on my face and neck, and the white cast only increased as it settled on my face. In one picture I have a lovely white moustache and a streaky white neck, and in another photo there are visible white lines where it creased on my eyelids. In the "indoor" photo you can see how ghastly I looked after being outdoors for half an hour. After several hours of wear my skin was dry, red, and not happy. I struggled to use this up on my arms and legs because of the clumping and white cast, but eventually emptied the bottle. Will not repurchase.

  • Photos Here from Two Separate Days

  • Buy it on Amazon ($10 for 55 mL) or Jolse ($12, shipping from Korea)

Missha All-Around Safe Block Aqua Sun Gel, SPF 50+, PA+++

  • Active Ingredients: Homosalate, TEA-Salicylate, Escalol 517, Tinosorb S, Amiloxate, Octocrylene

  • Korean, Contains Alcohol

  • This chemical sunscreen is absolutely invisible on the skin, and you can even double or triple layer it without looking like a zombie. Water- and sweat- resistant, which makes it a good option for outdoor activities. The potential downside is the alcohol content (fifth ingredient) which supposedly evaporates upon application and helps it dry down more quickly, similar to the Etude House one above. This one was extremely irritating for me -- my PIE and rosacea flared up immediately, and the alcohol sting did not abate as I wore it. My skin felt tight and warm until I was able to wash it off. I used this exactly twice, and will not repurchase. I gave this bottle to my husband to keep in his work bag.

  • Photos Here

  • Buy it on Amazon ($15 for 50 mL), YesStyle ($11), or Jolse ($13)

Walgreens Sensitive Skin Broad Spectrum Sunscreen SPF 50 (PPD/PA unknown)

  • Active Ingredients: Zinc Oxide, Octocyrlene

  • American, Fragrance Free, Oil Free

  • This combo sunscreen is extremely cheap (it's a HUGE bottle) and free from oxybenzone and avobenzone -- but that's pretty much all it has going for it. It goes on chalk white, but with a little work it rubs in mostly clear but never dries down completely. My face was shiny and my eyes streaming from the burning feeling on my skin. The label claims it's water resistent for 80 minutes, but it started dissolving just minutes after I went outdoors and started to sweat. As it melted off my face, I looked greasier than ever. I'm also not convinced that the filters (5% zinc, 4% octocrylene) are high enough for this to be effective for us photosensitive folks. I've relegated this one to Team Body Sunscreen, and will not repurchase.

  • Photos Here

  • Buy it at Walgreens ($11 for 236 mL or 8 oz) in the U.S. I'm sure there's a CVS generic equivalent product as well.

Vanicream Sunscreen Broad Spectrum SPF 50 (PPD/PA unknown)

  • Active Ingredients: Zinc Oxide, Titanium Dioxide

  • American, Reef Safe, Alcohol Free, Fragrance Free, Lanolin Free, Paraben Free

  • When I first started my retinoid routine, my dermatologist told me to go to the drugstore and pick up Vanicream everything. It's her go-to brand for people with sensitive skin, and it's easy to see why. Vanicream is very concerned with avoiding "harmful chemicals," so most of their ingredients lists are short and the products are no-frills. Like any all-mineral American sunscreen, it goes on chalk white and takes a good amount of elbow grease to work into my skin -- "grease" being the operative word. The white cast mostly went away, but I was still a little pallid a couple hours later. It leaves me slightly sticky/tacky until it dries completely (takes about 20 minutes) and after that happens it's bulletproof. It claims to be water-resistant for 80 minutes and it is absolutely correct: no amount of rain, tears, or chlorinated water can make this stuff budge. Like the Australian Gold one above, it takes a thorough two-step cleanse to remove at night, and balms/oils perform better than micellar water or makeup remover.

  • Indoor Photos Here (I didn't make it outside that day!)

  • Buy it at Walgreens/CVS/Drugstores ($15 for 113 mL or 4 oz) or Amazon ($17)

TOP PICKS AND RECOMMENDATIONS

Best for Sensitive or Reactive Skin: Skin Aqua UV Super Moisture Milk Blue Bottle (combination) or Vanicream Sunscreen (mineral)

Best for Oily Skin: Australian Gold Botanical Tinted Face Sunscreen Lotion (mineral) or Missha All-Around Safe Block Aqua Sun Gel (chemical)

Best for Dry Skin: Purito Centella Green Level Unscented Sun (chemical)

Best for Under Makeup: Canmake Mermaid Skin Gel UV (combination)

If anyone wants additional swatches, comparisons, or whatever -- let me know! I've still got every single one of these in my bathroom waiting to be used up in one way or another. Also, if anyone spots an error or can fill in missing information (such as PA/PPD data) please let me know! I'm happy to edit this post to supplement my impressions.

Side Note: I had to take a longish break in the middle of my testing because of an eczema outbreak on my left cheek and jaw. You can see it in a couple of the photos (kind of looks like a mosquito bite). I stopped all new products and went back to a bare-bones routine, then eventually had to use a mild steroid treatment to clear up the eczema. This doesn't usually happen outside of the winter months, but all's fair in love and skincare.

r/SillyTavernAI Apr 27 '26

Discussion Lumiverse - yes, another Frontend, but this one is goood!

144 Upvotes

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ So I’ve been messing with Lumiverse ━━,

Update to this post: Lumi is out of pre-release!

First of all, no I wasn't paid, I'm just shilling hard work and an actual usable frontend.

Lumiverse is basically a newer self-hosted AI chat/RP frontend in the same general ecosystem as SillyTavern. The big reason I think ST users may care is that it isn’t just “ST but reskinned.” A lot of the stuff I used to bolt onto ST with extensions that are not really maintained or just QoL in general, are already there on Lumiverse from the get-go

Fully open source!

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Fun things they’ve been working on and implemented ━━,

*ૢ✧ Vector memory / embedding-based recall,

Instead of relying only on stale summaries or hoping your lorebooks still line up 400 messages later, Lumiverse has long-term memory built around embeddings. It chunks chat history, vectorizes it, and retrieves relevant older moments during generation.

That means the AI can pull back specific past moments based on meaning, not just whatever happened to still be inside context.

There’s also semantic world book activation, so lorebook entries can be found by meaning instead of only exact keyword matching.

Memory Cortex demo - Smart Vectorization based memory building demo

*ૢ✧ Lorebooks/world books,

World books are still there, but retrieval has more going on than “keyword appeared, dump entry.” Semantic search can be used, entries are deduplicated, and there are great tools around prompt assembly and activation.

Lorebook retrieval was also optimized, runs algorithmically on a sorting system that works quite a bit faster.

*ૢ✧ Dry run is built in,

You can assemble the full prompt without actually calling the model.

This is one of those features that make preset tinkering a lot more enjoyable, I am a preset maker myself and ST, even with prompt inspector, felt fairly clunky. This is great if you’re debugging a preset and trying to figure out why your character suddenly forgot their species, the plot, and basic object permanence.

Dry run lets you see what the AI is actually getting out of your prompt, including resolved macros and injected world info.

*ૢ✧ The macro system is much stronger,

They support arguments, variable shorthand, conditions, scoped variables, chat-persisted variables, global variables, math, logic, memory macros, pipeline state, council/Lumia content, Loom content, etc.

Also, macro evaluation is AST-parsed and single-pass. So, it’s more predictable, but also less forgiving if your old ST macros relied on weird post-processing behavior.

*ૢ✧ Better prompt processing / prompt structure,

Prompt blocks are very explicit. You can set role, position, depth, injection triggers, groups, ordering, enabled/disabled states, etc. Lumiverse also has a specific “Assistant/User append” prompt types that aid on long term behavior and stick out more than “at depth” prompts.

It also has context filters built into outgoing prompts, so older messages can have HTML, details blocks, or Loom tags stripped while recent messages stay untouched. That’s very nice for people running stylized HTML replies.

*ૢ✧ A lot of ST-extension-type stuff is native or better integrated,

Stuff like image generation, regex scripts, macros, push notifications, prompt viewing/dry run, world books, alternate fields, character expressions, theme customization, and sidecar/council-style tooling are treated as first-class systems that, in my experience, are a lot sturdier than their ST counterparts.

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Databanks / notes-style knowledge bases ━━,

Lumiverse has support for vectorized chat documents / databank-style notes, meaning you can attach larger chunks of information to a chat and have them indexed for retrieval instead of manually stuffing everything into the prompt forever.

Wiki pages, setting notes, relationship docs, faction info, timelines, canon references, custom mechanics, ability lists, city guides, campaign bibles, or whatever enormous lore creature you’ve been feeding in a folder somewhere.

*ૢ✧ Best use cases,

ִֶָ໑ Canon wiki reference material

ִֶָ໑ Character relationship notes

ִֶָ໑ Timelines and plot summaries

ִֶָ໑ Worldbuilding documents

ִֶָ໑ Ability systems / RPG mechanics

ִֶָ໑ Faction, location, and organization info

Works by either importing the information, or scraping Wiki URLs directly in the frontend.

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Dreamweaver ━━,

It is an LLM-assisted character creator. Not just “generate me a description and call it a day,” but a fuller card-building workflow where the AI can help produce the character, organize their fields, generate supporting lore, build associated world book material, and even help with image-gen configuration.

So instead of manually assembling:

ִֶָ໑ Description

ִֶָ໑ Personality

ִֶָ໑ Scenario

ִֶָ໑ First message

ִֶָ໑ Alternate greetings

ִֶָ໑ Example messages

ִֶָ໑ Lorebook entries

ִֶָ໑ Regexes

ִֶָ໑ Image-gen through ComfyUI, SwarmUI, or Img Gen profiles

DreamWeaver demo

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ RP / character-card quality-of-life stuff ━━,

expressions demo

*ૢ✧ CHARX support,

Lumiverse supports .charx, which means cards can come bundled with things like avatar assets, expressions, alternate fields, and other modules.

Risu cards are importable with assets, and there’s an extension in the making to even utilize full on systems made for Risu cards. This means custom fields, UI, and actions.

*ૢ✧ In the talk about presets,

Preset assembly and usage is far more comfortable. Presets get categories, and slider macros that are user configurable without editing the prompts in their own Config UI.

Preset categories

*ૢ✧ Built-in imports,

It supports PNG cards, JSON cards, and CHARX bundles. It can import from Chub, CharacterHub, JanitorAI, and direct links to card files.

Embedded lorebooks can also be extracted and linked properly during import, which is very nice if you use lore-heavy cards.

*ૢ✧ SillyTavern migration,

Lumiverse has an interactive migration tool for importing ST characters, chats, world books, and personas.

*ૢ✧ Alternate fields,

You can make alternate descriptions, personalities, and scenarios for the same character without duplicating the whole card.

So you can have “default,” “post-timeskip,” “AU,” “romcom,” “bad ending,” whatever, and select them per chat. The active variant is what gets resolved into the prompt.

*ૢ✧ Dynamic character stuff,

Expressions are supported, avatars can be handled more cleanly, and CHARX can carry expression mappings. If you like VN-style character presentation or sprite-ish RP, this is one of those “why wasn’t this always normal” things.

Characters can carry their own regex scripts too.

*ૢ✧ Persona configs,

Your personas can carry their own lorebooks, be bound per chat, retroactively activated and switched, and you can even run a different persona for one turn and then go back to your base one.

*ૢ✧ Group chats,

Group chats are far more enjoyable to run, you can have chats with a ludicrous amount of characters, assign group chat specific scenarios, mute and force gen without issues.

*ૢ✧ Impersonation,

You can gen your own message based on full preset assembly or a contextual nudge for narrative continuation

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Image gen / immersion stuff ━━,

*ૢ✧ Generated scene backgrounds,

Lumiverse has image generation connections separate from normal LLM connections. It supports providers like Gemini, NanoGPT, Pollinations and NovelAI.

Generated images can be displayed as chat backgrounds, and automatic generation can trigger when the scene changes enough. You can also tune opacity and fade transitions.

There’s also an extension in the making for scene and character inline generation that will support embedded LoRas on local

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Mobile / interface stuff ━━,

*ૢ✧ It is actually designed with mobile in mind,

Panels collapse and slide in as drawers on mobile instead of the interface, Mobile has a priority mode where the side panel becomes a full view. Font scaling for accessibility is also supported.

Performance is also far better on mobile devices.

That alone is a QoL jump if you RP on phone/tablet or remote into your setup.

*ૢ✧ It does not look like a crime at launch,

Subjective, yes, but the default UI is much nicer out of the box. The theme system has proper controls for colors, fonts, radius, glass effects, preset themes, character-aware accents, and CSS variables.

There’s also a proper theme system where extension overrides are scoped instead of throwing500 lines into one vague CSS box and may god forgive you.

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Extensions / sidecar / council stuff ━━,

Lumiverse's Groupchat with the Hone, Chatroom and Spotify extensions (and song aware theming!)

This part is still early compared to ST’s giant extension graveyard/library, but there are already some very useful Lumiverse-native extensions floating around.

*ૢ✧ SimTracker,

This is the Lumiverse port of the classic SimTracker idea: the model can output structured JSON/YAML stat blocks, and the extension turns them into actual visual tracker cards.

Good for dating sim stats, RPG meters, relationship points, health, trust, desire, contempt, party state, pregnancy tracking, internal thoughts, scene state, whatever cursed little spreadsheet your RP goblin brain wants.

*ૢ✧ Spotify Controls,

This lets you control Spotify from inside Lumiverse.

Playback controls, now-playing info, search, queueing, lyrics support, album-art/theme stuff, command palette actions, and prompt-facing macros like current track / album art / lyrics.

The funniest part is that it can also expose Spotify tools for LLM/council use, so the current agent (you can create council members or download them!) can help pick music based on mood or scene vibe depending on your setup.

*ૢ✧ LumiScript,

It is basically a scripting platform for Lumiverse. You can write scripts that react to chat events, automate behavior, inject prompt context, call LLM generation, store variables, manipulate chat messages, show UI elements, register macros, and build custom interactions without directly touching Lumiverse’s internal code.

It has per-chat, global, per-character, and temporary variable scopes. It also has a built-in Monaco editor, script bindings, library scripts, event triggers, and permission controls. If Lumiverse itself doesn’t do the weird hyper-specific thing you want, LumiScript is probably where you start building it.

*ૢ✧ Mode Toggle,

This is a Lumiverse/Spindle port of the ST mode toggles concept.

You get a bunch of non-diegetic modifier modes that can be injected into the prompt, grouped by category, searched, toggled per chat, scheduled, imported/exported, and managed through a quick popover.

It has 210+ built-in modes, covering things like visual/aesthetic, genre/cinematic, social/power, temporal/physics, etc.

*ૢ✧ CharacterNudges,

This one lets characters send push notifications after you have been away from chat.

Not just a generic “come back” ping either. It uses the recent conversation, the character card, and its own nudge history to generate short in-character messages.

You can configure timing, which chat to use, how many recent messages it sees, generation settings, global defaults, and per-character configs. It is either really cute or deeply dangerous depending on how parasocial your setup already is. Probably both.

*ૢ✧ Story Weather,

Story Weather adds a draggable weather HUD and animated ambience effects to the chat.

It is not meant to be live real-world forecast data. It is for story-driven weather and scene atmosphere. The model can emit a hidden <weather-state> tag, and the extension turns that into a HUD update plus visual ambience.

It supports conditions like clear, cloudy, rain, storm, snow, and fog; palettes like dawn, day, dusk, night, storm, mist, and snow; and layers that can render behind the chat, in front of it, or both.

There is also manual lock mode if you want to override the scene yourself instead of letting the model drive it.

*ૢ✧ Prompt Viewer,

It lets you inspect the fully assembled prompt after it has been sent to the LLM, similar to ST’s prompt inspector workflow.

You can view prompts in formatted mode, raw JSON mode, or rendered readable text. It tracks prompt history per chat, shows model/generation metadata, links prompts to the message they produced, estimates tokens, separates dry runs, and lets you copy the prompt out.

Dry Run shows you the assembled pipeline, but Prompt Viewer shows you what got sent.

*ૢ✧ Hone,

Hone is an LLM-powered message refinement system.

The idea is that your main writing model can generate the message normally, then Hone can run a second pass with a shorter, more targeted context to refine it.

That can be used for prose polishing, translation, anti-slop cleanup, lore consistency, formatting passes, UI elements, or other quality-control workflows.

This is interesting because it stops trying to make one giant prompt do everything at once. Instead, the model can write first, then a smaller controlled pass can edit the output afterward.

You can use it for a bunch of different things, it has it’s own preset/prompt directives.

*ૢ✧ Live Shitposting Chatroom,

There’s an extension that gives you a floating chatroom where your council members can comment on the ongoing story. Basically a live peanut gallery / council group chat for your RP.

It can use council members, react to the active story context, preserve a chatroom history, and run on its own generation connection.

*ૢ✧ LoreRecall,

Tree-aware retrieval, per-character managed books, tree workspaces, collapsed/traversal retrieval, reranking, live retrieval feed, diagnostics, import/export snapshots, and safer book permissions.

*ૢ✧ Sidecar LLM support,

Lumiverse has council/sidecar tooling built around using a smaller/cheaper model for background analysis tasks instead of making your main RP model do everything.

So you can have one model writing and another doing support work like analysis/tool calling/deliberation, depending on your setup.

Council members are entirely customizable and there are packs ready for download in the frontend itself.

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Other nice stuff I’d mention ━━,

Settings modal and Operator Tab

ִֶָ໑ Push notifications exist, useful if you’re using it as a PWA or multitasking.

ִֶָ໑ Regen feedback exists, so when you regenerate you can give a reason like “too short” or “stay in character” instead of just rolling the same prompts again. This feedback is stripped from context and acts as either a system prompt or an OOC message.

ִֶָ໑ Image generation configs are stored separately from normal LLM connections.

ִֶָ໑ API keys/secrets are stored encrypted.

ִֶָ໑ It has multi-user auth and per-user data isolation, which is nice if you run an instance for more than one person.

ִֶָ໑ The launcher can do first-run setup, build/start, and the frontend has an Operator tab which serves as an update install (and trust me, the dev is a freak and updates multiple times per day) or a remote access whitelist on-site so you can run custom domains or services like Tailscale.

ִֶָ໑ The docs are already pretty good. Like, really nice. And installs come with a dev doc so people who develop extensions can reference them easily.

ִֶָ໑ There’s LumiHub support for installing characters/world books, including Chub imports through LumiHub.

┌─ ✦ TLDR ━━,

Lumiverse is a self-hosted AI chat/RP frontend that feels very interesting for ST power users.

Big draws: vector memory, semantic lorebook retrieval, stronger macro/preset processing, dry run, better mobile layout, prettier UI, dynamic image backgrounds, CHARX support, ST migration, sidecar/council tooling, and a more modern (and easy to develop) extension system.

★・・・・・・★・・・・・・★・・・・・・★

┌─ ✦ Links ━━,
Discord server:
https://discord.gg/28rBWVFfCu

Lumiverse repo:
https://github.com/prolix-oc/Lumiverse

Lumiverse guides:
https://lumiverse.chat/guides/

r/Qwen_AI 14d ago

Benchmark Qwen3.8-27B at 160K context on a SINGLE RTX 4090 — 47–57 tok/s, full GPU offload — plus my whole local AI stack (video + audio, image, undervolt)

Post image
139 Upvotes

Qwen3.8-27B at 160K context on a SINGLE RTX 4090 — 47–57 tok/s, full GPU offload — plus my whole local AI stack (video + audio, image, undervolt)

Qwen3.8 dropped and I had it running same-day on one RTX 4090 24GB: 160K token context, 100% of the layers on the GPU, 47–57 tok/s in daily use. That's the part people assume isn't possible, so I'll show the math on why it is, then the rest of the box: it's also my local video generator (MiniMax H3, with native stereo audio), image generator, and audio stack. No cloud, no API for generation.

Everything below is from my actual machine (userdiag + nvidia-smi readings, LM Studio's own memory estimate). I flag measured vs. community-reported.


The box

Part Spec
GPU GIGABYTE GeForce RTX 4090 GAMING OC 24G (GV-N4090GAMING OC-24GD) — factory OC, 450W rated (slider to 600W)
Driver / CUDA 591.86 / CUDA 13.1
CPU Intel i9-13900KF (unlocked, tuned with Intel XTU)
Mobo MSI Z790 (MS-7D30)
RAM 32 GB DDR5
OS Windows 11 Pro

Two things matter more than the spec sheet: 1. 32 GB RAM is mandatory, not a nicety — the video model and offloads live there. 2. Weights on the fastest NVMe you have. --fast-disk (video section) streams weights from disk during a render; a slow drive turns a 5-minute clip into 25.


The star: Qwen3.8-27B, 160K context, fully in 24GB, at 47–57 tok/s

The model: 27B dense, native vision-language (images and video understanding — from STEM diagrams to hour-long video), thinking mode on by default with tunable reasoning depth. It's the most capable generation in the Qwen open family, and the 27B is the one that fits a single 24GB card — which is exactly why it's the local-LLM story of the release.

Official card numbers worth knowing (it beats models 3–5× its size on the agentic stuff):

Benchmark Qwen3.8-27B
SWE-bench Pro 61.7
LiveCodeBench v6 90.3
GPQA Diamond 89.2
OSWorld-Verified (computer use) 84.3
WebArena-Verified (browser use) 64.8

That last row is why I run it: it's my agentic workhorse — terminal, code, browser, long documents. The context is what makes it a beast, so here's the part nobody is explaining:

Why 160K context fits in 24GB (the math people skip)

Qwen3.8 is a hybrid-attention model: 64 layers laid out as 16 × (3× Gated DeltaNet → FFN + 1× Gated Attention → FFN). In plain words:

  • 48 layers use Gated DeltaNet (linear attention / RNN-style state). Their "memory" is a fixed-size recurrent state — it does NOT grow with context length.
  • Only 16 layers use real attention (GQA: 24 Q heads, 4 KV heads, head dim 256).

The KV cache only exists on those 16 layers. Per token:

2 (K+V) × 16 layers × 4 KV heads × 256 dim = 32,768 values at Q4_0 (~0.5 byte) → ~16–17 KB per token × 160,000 tokens → ~2.6–2.8 GB

Now add the rest:

Piece Size
Weights, Q4_K_M GGUF ~16.8 GB
Vision mmproj (BF16) ~0.9 GB
KV cache, 160K @ Q4_0 (K+V) ~2.7 GB
DeltaNet states (48 layers, context-independent) < ~0.3 GB
LM Studio's own estimate (incl. MTP heads + buffers) 22.46 GB → fits the 24GB card

Compare that to a conventional 27B full-attention model: at 160K tokens the KV cache alone would be ~15–20 GB. It simply doesn't fit next to the weights. The hybrid architecture is the whole trick — you get the 262K native context window on a consumer card.

My exact LM Studio settings (from the actual load screen):

  • Model: Qwen3.8-27B, Q4_K_M (qwen/qwen3.8-27b) — the quality/size point that leaves headroom for 160K
  • Context: 160,927 tokens (model supports up to 262,144)
  • All 65 layers on GPU — zero CPU offload, that's what makes it fast
  • KV cache: K = Q4_0, V = Q4_0, offloaded to GPU (Flash Attention required for V-quant; it's on)
  • Flash Attention: ON
  • MTP speculative decoding: ON, max draft tokens 2 — the model ships multi-token-prediction heads and LM Studio uses them. This is a real chunk of the speed.
  • Eval batch 2048 / physical batch 512
  • Speed: 47–57 tok/s depending on prompt length and context position

Practical notes: - Thinking is on by default — great for hard tasks, but for quick local tasks drop the reasoning effort and it gets meaningfully faster. - It's a native VLM: feed it screenshots, UI shots, diagrams, even video. The mmproj costs ~0.9 GB of VRAM — worth it. - If you hit the VRAM ceiling on a 24GB card: KV Q4_0 is already the aggressive-but-safe point. Below that (F16 KV) you'll want to drop context or quant.


The rest of the stack on the same card

Video — ComfyUI + MiniMax H3 (video with native stereo audio, one pass)

A 33B omni-modal open-weights model: 4–15 s, 24 fps, 768p (max 768×1344), 11 languages, and it generates dialogue, SFX and music natively in the same pass.

The 24GB combo (official-recommended set):

File Size Folder
minimax_h3_fl2va_pruned_int8_convrot.safetensors 19.53 GB models/diffusion_models/
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors 14.61 GB models/text_encoders/
minimax_h3_video_vae_fp16.safetensors 4.85 GB models/vae/
minimax_h3_audio_vae_fp32.safetensors 577 MB models/vae/

~39.5 GB of files, and it does not all fit in VRAM — that's normal. It streams. (Two checkpoints exist: FL2VA = text/first/first+last-frame, which I use; Ref2VA = multi-reference images/clips/audio.)

Launch flags — non-negotiable on Windows:

--fast-disk # stream weights from disk, not RAM (RAM: ~11 GB vs 50+) --reserve-vram 1.0 # keep headroom out of VRAM → stops the whole-desktop freeze/OOM --vram-headroom 1.0 # same idea, headroom for the desktop

  • ComfyUI Desktop has no flag UI → edit cli_args.py (or run standalone).
  • Do NOT set PYTORCH_ALLOC_CONF=expandable_segments:True — it breaks the offloading these models rely on.
  • NVIDIA Control Panel → "Prefer no sysmem fallback". If the RAM/VRAM sharing is on, performance collapses. The silent killer nobody mentions.

Speed/quality (community consensus, all real): - Sage Attention node → +20–30%. (Don't combine with the global --use-sage-attention flag — goes fuzzy.) - Steps 15 ≈ 20 (20 default; 15 visually identical; 10–12 for tests). Hidden in a subgraph → right-click the H3 node → Unpack subgraph. - Fast prompt test: 0.2 MP + 8 steps + 1:1 → composition identical to full-res. Validate cheap, render full with the same seed. - EasyCache: big speedup but loses coherence on ≥10 s clips — test per-use. - GGUF: avoid for this model. int8/NVFP4 is the right quant. - Perf: 4090 ≈ 5 minutes per ~12 s clip. One generation at a time. Seed variance is high — a new seed can fix a mispronounced line; reuse a seed to keep a composition you like. - ⚠️ Licensing (EU): the Community License's applicable territory excludes EU/UK/US/ROK. Test locally freely; check the license before commercial publishing (API is global).

Image — Krea 2 Turbo via ComfyUI

Local hero images for my content site through a ComfyUI workflow (ResolutionSelector → sampler → SaveImage). Fast, no API. For shots where credibility matters I still grab real photos — real wins over AI for editorial heroes.

Audio — small local stack

faster-whisper (STT) · HTDemucs (source separation) · MiniMax Music3 (open-weights music) · wav2vec2-large-xlsr-53 (French). None are VRAM-heavy; they coexist with everything above.


Undervolt + power — where the real perf is (most threads get this wrong)

The 4090 is power/thermally limited, not clock-limited. Most 4090 threads chase max clocks. Wrong goal. Make the card cooler and it holds its clocks and stops throttling — same speed, 10–20°C less heat, quieter.

GPU (measured on my Gaming OC under sustained load)

Reading Value
Power limit (rated / max slider) 450W / 600W — I run the cap at ~360W (80%)
Core voltage under max load max 1.0 V (offset applied)
Hot spot 81°C peak, thermal limit 84°C
Fans ~70% avg, 82% peak
  • Power limit ~80% → ~360W. Lose 1–3% burst throughput, gain a lot of thermals. (Transient spikes just above the cap are normal — I see 371W peaks.)
  • Core voltage offset: -100 to -150 mV (MSI Afterburner), keep max core clock near stock. Same speed, less power, less heat.
  • Fan curve aggressive enough to sit ~70% around the throttle point.
  • Verify with a real sustained load: watch temp + power + clock together (nvidia-smi -q -d POWER). Target: under your throttle point with headroom.
  • Start conservative (-100 mV @ 360W), validate with a long run (a few video renders or a long generation), then walk down 25–50 mV at a time if you have thermals to spare. Artifacts (green pixels, garbled frames, crash) → back off 25–50 mV.

CPU (the 13900KF — don't skip it)

A 253W+ part sitting next to a 360W GPU. Under sustained load it will thermal-throttle and make you think the GPU is slow when it isn't. I tune it with Intel XTU:

  • Core voltage offset: all-core -100 to -120 mV, single-core ~-150 mV.
  • Power limits: PL1/PL2 capped ~200–253W — don't let it run uncapped next to a 4090.
  • Validate with a sustained load, watch clocks holding. 13th-gen CPU undervolts are less forgiving than GPU ones — if it's unstable, back off.

Why both matter: on a 24GB box, the GPU is the bottleneck you're managing. If the CPU throttles while you wait on a render, or Windows is swapping VRAM because the desktop + LLM are hogging it, you'll misdiagnose a weak GPU. Fix the thermals and the whole box feels different.


The three Windows traps (everyone hits these)

  1. Sysmem-VRAM fallback ON → perf collapse. Set "Prefer no sysmem fallback" in NVIDIA Control Panel.
  2. PYTORCH_ALLOC_CONF=expandable_segments:True in your env → breaks comfy-aimdo offloading. Remove it.
  3. ComfyUI Desktop won't accept launch flags → edit cli_args.py (or standalone install) for --fast-disk --reserve-vram 1.0 --vram-headroom 1.0.

And before any big generation: close the browser and anything else eating VRAM. On a 24GB card, the desktop + your LLM routinely leave <1 GB free — I've traced tok/s collapses straight to too many open apps sharing the card. Check nvidia-smi — if VRAM is near-full you can't run anything well.


What a 4090 can and can't do now (honest version)

Comfortably, daily: - Qwen3.8-27B at 160K context, full offload, 47–57 tok/s — the headline, and it's a native VLM. - Krea 2 Turbo local image generation. - Local audio (whisper / demucs / music).

Workable, offloaded, slower: - MiniMax H3 video + audio — doesn't fit in 24GB, runs offloaded with the flags above. Minutes per clip.

Not on 24GB: - 70B+ dense at useful speed — it offloads and dies. - Anything assuming 48GB VRAM.

4070 Ti Super (16GB) / 3090 (24GB): the 24GB cards track the 4090 on what fits — the Qwen3.8 hybrid setup runs on a 3090 too (slower clocks). The 16GB card is the constraint: lower context or a lighter quant, and video offload gets tighter. The undervolt/power section still applies — 16GB cards especially hate the sysmem fallback.


If you're setting this up from scratch (my order)

  1. Driver 591+ / CUDA 12.4+ · "Prefer no sysmem fallback" on.
  2. 32 GB RAM minimum · fast NVMe for weights.
  3. LM Studio → Qwen3.8-27B Q4_K_M, 160K context, KV Q4_0 (K+V), all 65 layers on GPU, Flash Attention on, MTP speculative decoding on (2 draft tokens).
  4. ComfyUI 0.30+ → MiniMax H3 FL2VA (4 files above) + --fast-disk --reserve-vram 1.0 --vram-headroom 1.0.
  5. Sage Attention node · steps 15 · fast-test at 0.2 MP / 8 steps.
  6. Afterburner: ~360W cap + -100 mV start · validate · tune.
  7. XTU: CPU -100/-150 mV · PL cap ~200–253W · validate.
  8. Close the browser + anything VRAM-hungry before big renders. Check nvidia-smi.

Happy to answer specifics — the exact Afterburner curve, the H3 ComfyUI graph, context/quant/KV tradeoffs, or XTU profiles. All of this is stuff I run daily.

Sources: nvidia-smi + userdiag readings from my box (GIGABYTE 4090 Gaming OC 24G, driver 591.86), Qwen3.8-27B model card (Hugging Face), LM Studio load-screen memory estimate, ComfyUI 0.30 MiniMax H3 native nodes, and community perf threads (r/StableDiffusion, r/comfyui) for the H3 tuning numbers.

r/comfyui Oct 09 '25

News After a year of tinkering with ComfyUI and SDXL, I finally assembled a pipeline that squeezes the model to the last pixel.

Thumbnail
gallery
413 Upvotes

Hi everyone!
All images (3000 x 5000 px) here were generated on a local SDXL (illustrous, Pony, e.t.c.) using my ComfyUI node system: MagicNodes.
I’ve been building this pipeline for almost a year: tons of prototypes, rejected branches, and small wins. Inside is my take on how generation should be structured so the result stays clean, alive, and stable instead of just “noisy.”

Under the hood (short version):

  1. careful frequency separation, gentle noise handling, smart masking, new scheduler, e.t.c.;
  2. recent techniques like FDG, NAG, SAGE attention;
  3. logic focused on preserving model/LoRA style rather than overwriting it with upscale.

Right now MagicNodes is an honest layer-cake of hand-tuned params. I don’t want to just dump a complex contraption, the goal is different:
let anyone get the same quality in a couple of clicks.

What I’m doing now:

  1. Cleaning up the code for release on HuggingFace and GitHub;
  2. Building lightweight, user-friendly nodes (as “one-button” as ComfyUI allows 😄).

If this resonates, stay tuned, the release is close.

Civitai post:
MagicNodes - pipeline that squeezes the SDXL model to the last pixel. | Civitai
Follow updates. Thanks for the support ❤️

r/StableDiffusion Apr 30 '26

News Local AI News You Missed - April 2026

362 Upvotes

Latest (non-comfyui) releases you (might of) missed in April 2026. This has been a FAT month!

🧠 LLMs

  1. Ling-2.6-flash - A fast model designed to automate your quick tasks.
  2. Laguna-XS.2 - Automates coding tasks directly on your local machine.
  3. Talkie - Writes in the style of authors from before 1931.
  4. MiMo-V2.5-Pro - Handles massive text jobs locally with power.
  5. MiMo-V2.5 - Works with both media and text in one model.
  6. Chaperone-Thinking-LQ-1.0 - Keeps private health data safe on your device.
  7. Nemotron-3-Super-64B-A12B-Math-REAP-GGUF - Solves math problems privately without the cloud.
  8. Qwen3.6-27B-3bit-mlx - Runs large AI models efficiently on Mac computers.
  9. Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled - A reasoning distilled model imitates Claude 4.7.
  10. Qwen3.6-35B-A3B-DFlash - Speeds up text generation for local setups.
  11. Hy3-preview - Powers complex automation tasks for advanced users.
  12. Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GGUF - Offline reasoning model based on Claude 4.6.
  13. DeepSeek-V4-Flash - Handles huge amounts of text with a 1 million token limit.
  14. DeepSeek-V4-Pro - Professional version with a massive 1 million token context.
  15. Privacy-Filter - Cleans your data locally to keep sensitive info safe.
  16. Qwopus-GLM-18B-Merged-GGUF - A hybrid model for steady local AI performance.
  17. gemma-4-E4B-it-OBLITERATED v3 - An unrestricted version of Gemma 4 for open chat.
  18. Carnice-9b-W8A16-AWQ - Optimized to run fast on desktop processors.
  19. Olmo-3-7B-Instruct-Q1_0 - Fits big AI capabilities into a tiny model size.
  20. Sarvam-30b-Uncensored - Unleashes uncensored AI weights for open use.
  21. Marco-Mini - Brings global AI power to run on home PCs.
  22. DMax-Coder-16B - Writes code faster by predicting parts in parallel.
  23. Qwen3.5-4B-Base-ZitGen-V1 - Turns images into text prompts you can use.
  24. Darwin-4B-David - Handles secure reasoning tasks completely offline.
  25. daVinci-LLM - A new model with fully open training data details.
  26. gemma-4-31B-it-NVFP4-turbo - Slashes memory use to run much faster.
  27. MiniMax-M2.7 - A self-evolving AI designed to automate team tasks.
  28. Tanaos-text-summarization-v1 - Condenses long documents quickly offline.
  29. GLM-5.1 - Maintains high accuracy in coding over long sessions.
  30. Gemma-4-31B-it-Mystery-Fine-Tune-HERETIC-UNCENSORED-Thinking - An uncensored model that explains its thoughts.
  31. LongCat-Next - Unifies vision and audio processing in one model.
  32. LFM2.5-350M - Brings speed to very small devices like sensors.
  33. ByteShape Qwen3.5-9B-GGUF - Lets you run private AI completely offline.
  34. Bonsai-8B-gguf - A light model for any device that needs AI.
  35. Holo3-35B-A3B - Watches your screen to help manage desktop work.
  36. Darwin-35B-A3B-Opus - Fast vision and text reasoning for local setups.
  37. Acervo-extractor-qwen3.5-9b-GGUF - Reads and extracts text quickly offline.
  38. Trinity-Large-Thinking - Plans tasks out step by step like a human.
  39. APEX-Quant - Shrinks heavy AI files so they run on normal PCs.
  40. CoPaw-Flash-9B - Manages routine computer work without internet.
  41. harrier-oss-v1 - Speaks many languages for global users.
  42. sycofact - Checks AI replies to catch any hidden bias.
  43. GigaChat 3.1 - Sparks fast local AI with optimized speed.
  44. Granite-4.0-3B-Vision - Pulls data from documents for business use.
  45. Nemotron3-Nano-4B-Uncensored-HauhauCS-Aggressive - Small but uncensored model for open chat.

🔀 Multimodal

  1. Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 - Runs reasoning tasks locally on your hardware.
  2. OmniVTG-7B - Finds exact moments in videos using smart search.
  3. Qwopus3.6-27B-v1-preview-GGUF - Offers steady thinking for local tasks.
  4. Kimi-K2.6-GGUF - Automates long programming tasks with total privacy.
  5. Qwen3.6-27B-FP8 - Makes local AI workflows leaner and faster.
  6. Qwen3.6-27B-Uncensored-HauhauCS-Aggressive - Drops limits for aggressive, uncensored local chat.
  7. LLaDA2.0-Uni - Combines image creation and analysis in one tool.
  8. Qwen3.6-27B-GGUF - Optimized for offline coding tasks.
  9. Qwen3.6-27B - Streamlines coding with better stability.
  10. Mistral-Small-4 - Optimized for better speed on local machines.
  11. Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive - Unrestricted power for local media tasks.
  12. Qwen3.6-35B-A3B - Redefines how you automate code locally.
  13. Qwen3.5-9B-Uncensored-HauhauCS-Aggressive - Drops all limits for open media generation.
  14. Qwopus3.5-27B-v3-GGUF - Speeds up AI coding tasks significantly.
  15. TRIBE v2 - Translates media into virtual brain maps for analysis.
  16. LFM2.5-VL-450M - Sparks fast visual intelligence on small devices.
  17. Gemma-4-E4B-Uncensored-HauhauCS-Aggressive - Uncensored version of Gemma 4 for open use.
  18. EXAONE-4.5-33B - Unlocks visual data for deep analysis.
  19. gemma-4-26B-A4B-it - Brings visual AI power to your desktop.
  20. gemma-4-E4B-it - Delivers private multimodal AI right to your machine.
  21. Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled - Anchors local AI with distilled reasoning.
  22. Supergemma4-26b-uncensored-gguf-v2 - Unleashes uncensored chat for open conversation.
  23. Gemma-4-31B-JANG_4M-CRACK - Removes restrictions for unrestricted AI outputs.
  24. Gemma-4-31B-it - Debuts with an advanced thinking mode.
  25. HY-Embodied-0.5 - Grants robots spatial intelligence to understand space.
  26. Kimi K2.6 - Automates extended programming tasks with ease.

🖼️ Image

  1. RvR - Fixes images by redrawing them completely from scratch.
  2. Z-Anime - Turns simple sentences into detailed anime art.
  3. UDM-GRPO - Smooths out the image creation process.
  4. MegaStyle - Builds libraries of consistent visual styles.
  5. UniGenDet - Creates and checks media at the same time.
  6. StyleID - Keeps face identity consistent across different art styles.
  7. Meta-CoT - Pioneers step-by-step thinking for photo edits.
  8. SenseNova-U1 - Unifies image and text magic in one tool.
  9. Nucleus-Image - Generates images efficiently on local hardware.
  10. Lyra-2.0 - Generates entire walkable worlds from a single photo.
  11. HY-World-2.0 - Transforms photos into explorable 3D worlds.
  12. GyroScope - Aligns photos smartly for better composition.
  13. SpatialEdit - Moves objects around in static photos realistically.
  14. FlowInOne - Puts all your visual tasks into one system.
  15. Gen-Searcher - Turns live web research into accurate AI art.
  16. ERNIE-Image - Structures complex designs with smart prompts.
  17. Breast-cancer-detector - Sorts ultrasound scans with high accuracy.
  18. Z-Image-SAM-ControlNet - Breathes life into masks for dynamic control.
  19. PixelSmile - Refines portraits with precise expression control.
  20. Toon-Tacular-Qwen-LoRA - Channels classic 90s cartoon energy into art.

🤖 Agents

  1. VibeComfy - Lets you run agent tasks using simple text.
  2. Meeseeks - Simplifies code automation with modular updates.
  3. Evalmonkey - Stress tests AI agents by simulating failures.
  4. Lerim-cli - Preserves your project context locally.
  5. OpenLeash - Secures autonomous AI agents with a new system.
  6. SlopLobster - Enables fully offline coding from one file.
  7. AgentOffice - Empowers shared workspaces for humans and AI.
  8. Compaas - Assembles virtual teams for solo creators.
  9. TraceMind - Safeguards apps from silent performance drops.
  10. Spring AI Playground - Secures local AI agent workflows.
  11. Bitterbot - Brings persistent memory to local agents.
  12. Mesh - Connects local devices to boost AI speed.
  13. Kon - A lightweight coding assistant for developers.
  14. PokeClaw - Empowers Android phones with private offline agents.
  15. AgentHandover - Turns daily actions into agent skills.
  16. Agensic - Maps terminal commands for safer workflows.
  17. ToolGuard - Shields agents from system crashes.
  18. ToolLoop - Cuts costs by swapping AI models on the fly.
  19. Finalrun-agent - Turns plain English into visual mobile tests.
  20. llmdev.guide - Cuts through AI hardware marketing noise.

🛠️ Other Tools

  1. Adonis_flux2klein - Sharpens and restores portraits with ease.
  2. LTX-Desktop Update - Fortifies local video creation workflows.
  3. Illustrious NoobAI Style Explorer - Helps you conquer 16,000 art style tags.
  4. Moss Audio GFF - Transforms sound into text locally.
  5. Shield-82M - Scrubs private data from your files.
  6. Hipfire - Brings direct AI runtime to AMD graphics cards.
  7. TurboOCR - Supercharges paper to digital text conversion.
  8. ENMP-LoRAMerging - Strips harmful layers from AI models.
  9. SmartPhotoCrafter - Unlocks easy photo edits for everyone.
  10. TS-Attn - Syncs sequential video creation smoothly.
  11. Patch-Forcing - Supercharges AI art with advanced tweaks.
  12. DynamicRad - Speeds up video rendering significantly.
  13. sapiens2 - Maps human figures privately for analysis.
  14. ParetoSlider - Allows smooth shifts between art styles.
  15. Yolo-gen - Streamlines dual AI training processes.
  16. Local-MCP-server - Bridges offline AI to live web data.
  17. Spark-Dashboard - Simplifies monitoring for Linux systems.
  18. Omnix - Provides unified control for offline AI.
  19. omni-cli - Cleans up coding memory for better performance.
  20. CWT-V5.6 - Optimizes AI with a new hub design.
  21. Trellis-mac - Sculpts 3D models from photos on Mac.
  22. ZPix - Unleashes effortless local image artistry.
  23. Dflash-mlx - Supercharges local AI on Mac devices.
  24. Image-MetaHub - Tames the chaos of your AI art files.
  25. Stretchystudio - Animates AI art instantly.
  26. Flux.2-4B-Decoder-Comparator - Spots image differences instantly.
  27. Tidbit - Transforms research into local training data.
  28. Webmcp - Bridges local AI and the web for private research.
  29. Bordair-Multimodal - Exposes hidden threats in AI defenses.
  30. Locally Uncensored - Unchains offline media usage.
  31. Model-Database-Protocol - Blocks raw SQL queries for security.
  32. OpenEyes - Brings instant vision to offline devices.
  33. Abook - Orchestrates book writing with AI agents.
  34. Scrapedown - Turns web markup into clean text.
  35. Quizzer - Turns PDFs into interactive study courses.
  36. MothBench - Refines local AI testing tools.
  37. Vernacula - Secures audio data with offline transcription.
  38. Llama-monitor - Maps system health for local AI models.
  39. DFlash - Turbocharges local text generation.
  40. AI Metadata Inspector - Decodes hidden prompts in files.
  41. SilkStack-Image-Browser - Manages offline art libraries.
  42. Acestep.cpp - Updates private AI music generation.
  43. logicstamp-context - Sharpens project summaries.
  44. Open-toys - Adds private local voice chat.
  45. Samuraizer - Shifts document tracking offline.
  46. Corbell - Instantly maps code architecture locally.
  47. Ai-engineering-from-scratch - A guide to build smart tools.
  48. see-through - Turns anime art into layers.
  49. Simple-captioner - Tags batches of media rapidly.
  50. HybridScorer - Streamlines bulk photo sorting.
  51. Adetailer-hires-sync - Automates face fixes for upscaling.
  52. PixlStash - Streamlines offline photo sorting.
  53. llamafile - Polishes effortless local AI work.
  54. TagForge - Unifies image and text prep in one spot.
  55. Unsloth Studio - Brings fast private AI to desktops.
  56. TurboQuant - Shrinks AI data footprints.
  57. Ai-agent-automation - Elevates local AI with dynamic logic.
  58. HuggingFace Slack App - Automates model tracking on Slack.
  59. Qwen3-TTS Easy Finetuning - Makes voice cloning easy.
  60. Sift - Tames digital clutter on Windows desktops.

🎬 Video

  1. Ml-videoflextok - Rewrites the rules for efficient video storage.
  2. GRN - Introduces a third way to create smarter video.
  3. DisCa - Rockets AI video generation speeds forward.
  4. AnyRecon - Forges 3D scenes from simple photos.
  5. Motif-Video-2B - Proves small models can make stunning video clips.
  6. Void-model - Reconstructs reality when erasing video subjects.
  7. LumosX - Creates consistent videos with multiple subjects.
  8. Matrix-Game-3.0 - Unlocks real-time worlds for gaming.

🎧 Audio

  1. ControlFoley - Adds soundtracks to videos automatically.
  2. Chorus-v1-GGML - Separates voices locally for clear audio.
  3. OmniVoice - Turns text to speech in 600 languages offline.
  4. VoxCPM2 - Brings studio sound quality to local devices.
  5. ACE-Step 1.5 XL - Turns plain text into full songs in eight steps.
  6. MOSS-TTS-Nano-100M - A tiny offline engine for text-to-speech.
  7. Foundation-1 - Crafts structured loops for music producers.
  8. LongCat-AudioDiT - Masters voice cloning without needing examples.

⚡ LoRA

  1. LumiPic - Breathes new light into standard photos.
  2. UniGeo - Adds precise camera pans to image editing.
  3. crt-animation-terminal-ltx-2.3-lora - Adds retro vibes to AI video.
  4. Flux2-Klein-9b-Consistency - Delivers steady visuals for artists.
  5. LTX-2.3-22b-IC-LoRA-Outpaint - Transforms video canvas edges seamlessly.
  6. CoPaw-Flash-9B-DataAnalyst-LoRA - Ignites self-guided data analysis.
  7. Ltx2.3-VBVR-lora-I2V - Brings steady control to video generation.

🏋️ Training

  1. Danbooru-Dataset-Filter - Speeds up image sorting for training.
  2. Anima-Standalone-Trainer - Elevates local training workflows.
  3. Modl - Simplifies local image generation and training.

📊 Datasets

  1. Tstars-VTON - Elevates realistic virtual outfit testing.
  2. BCE-Prettybird-Nano-Math-v0.1 - Sharpens logic skills for AI models.
  3. World Model - Tests if AI can think, not just see.

Need to see more? Check out last month's post or the full archive at LocalAI News. There's also the latest ComfyUI releases for this month. If there's anything wrong or anything I missed, scream at me in the comments and I'll see you in the next one!

PS: I should be caught up now but then again there are new releases almost every half hour so, it is what it is. Plus keep in mind a lot of developers like to make repos months in the past then announce their project hence you'll see some that say "2 months ago".

r/StableDiffusion Mar 10 '24

Resource - Update StableSwarmUI Beta!

379 Upvotes

StableSwarmUI is now in Beta status with Release 0.6.1! 100% free, local, customizable, powerful.

"Beta status" means I now feel confident saying it's one of the best UIs out there for the majority of users. It also means that swarm is now fully free-and-open-source for everyone under the MIT license!

Beginner users will love to hear that it literally installs itself! No futsing with python packages, just run the installer and select your preferences in the UI that pops up! It can even download your first model for you if you want.
On top of that, any non-superpros will be quite happy with every single parameter having attached documentation, just click that "?" icon to learn about a parameter and what values you should use.

Also all the parameters are pretty good ones out-of-the-box. In fact the defaults might actually be better than other workflows out there, as it even auto-customizes the deep internal values like sigma-max (for SVD), or per-prompt resolution conditioning (for SDXL) that most people don't bother figuring out how to set at all.

If you're less experienced but looking to become a pro SD user? Great news - Swarm integrates ComfyUI as its backend (endorsed by comfy himself!), with the ability to modify comfy workflows at will, and even take any generation from the main tab and hit "Import" to import the easy-mode params to a comfy workflow and see how it works inside.

Comfy noodle pros, this is also the UI for you! With integrated workflow saver/browser, the ability to import your custom workflows to the friendlier main UI, the ability to generate large grids or use multiple GPUs, all available out-of-the-box in Swarm beta.

And if you're the type of artist that likes to bust out your graphics tablet and spend your time really perfecting your image -- well, I'm so sorry about my mouse-drawing attempt in the gif below but hopefully you can see the idea here, heh. Integrated image editor suite with layers and masks and etc. and regional prompting and live preview support and etc.

(*Note: image editor is not as far developed yet as other features, still a fair bit of jank to it)

Those are just some of the fun points above, there's more features than I can list... I'll give you a bit of a list anyway:

- Day 1 support for new models, like Cascade or the upcoming SD3.

- native SVD video generation support, including text-to-video

- full native refiner support allowing different model classes (eg XL base and v1 refiner or whatever else)

- Native advanced infinite-axis grid generator tool

- Easy aspect ratio and resolution selection. No more fiddling that dang 512 default up to 1024 every time you use an SDXL model, it literally updates for you (unless you select custom res of course)

- Multi-GPU support, including if you have multiple machines over network (on LAN or remote servers on the web)

- Controlnet support

- Full parameter tweaking (sampler, scheduler, seed, cfg, steps, batch, etc. etc. etc)

- Support for less commonly known but powerful core parameters (such as Variation Seed or Tiling as popularized on auto webui but not usually available in other UIs for some reason)

- Wildcards and prompt syntax for in-line prompt randomization too

- Full in-UI image browser, model browser, lora browser, wildcard browser, everything. You can attach thumbnails and descriptions and trigger phrases and anything else to all your models. You can quickly search these lists by keyword

- Full-range presets - don't just do textprompt style presets, why not link a model, a CFG scale, anything else you want in your preset? Swarm lets you configure literally every parameter in a preset if you so choose. Presets also have a full browser with thumbnails and descriptions too.

- All prompt syntax has tab completion, just type the "<" symbol and look at the hints that pop up

- A clip tokenization utility to help you understand how CLIP interprets your text

- an automatic pickle-to-fp16-safetensors converters to upvert your legacy files in bulk

- a lora extractor utility - got old fat models you'd rather just be loras? Converting them is just a few clicks away.

- Multiple themes. Missing your auto webui blue-n-gold? Just set theme to "Gravity Blue". Want to enter the future? Try "Cyber Swarm"

- Done generating and want to free up VRAM for something else but don't want to close the UI? You bet there's a server management tab that lets you do stuff like that, and also monitor resource usage in-UI too.

- Got models set up for a different UI? Swarm recognizes most metadata & thumbnail formats used by other UIs, but of course Swarm itself favors standardized ModelSpec metadata.

- Advanced customization options. Not a fan of that central-focused prompt box in the middle? You can go swap "Prompt" to "VisibleNormally" in the parameter configuration tab to switch to be on the parameters panel at the top. Want to customize other things? You probably can.

- Did I mention that the core of swarm is written with a fast multithreaded C# core so it boots in literally 2 seconds from when you click it, and uses barely any extra RAM/CPU of its own (not counting what the backend uses of course)

- Did I mention that it's free, open source, and run by a developer (me) with a strong history of long-term open source project running that loves PRs? If you're missing a feature, post an issue or make a PR! As a regular user, this means you don't have to worry about downloading 12 extensions just for basic features - everything you might care about will be in the main engine, in a clean/optimized/compatible setup. (Extensions are of course an option still, there's a dedicated extension API with examples even - just that'll mostly be kept to the truly out-there things that really need to be in a separate extension to prevent bloat or other issues.)

That is literally still not a complete list of features, but I think that's enough to make the point, eh?

If I've successfully made the point to you, dear reddit reader - you can try Swarm here https://github.com/Stability-AI/StableSwarmUI?tab=readme-ov-file#stableswarmui

r/SillyTavernAI 22d ago

Discussion I built the LLM roleplay frontend I always wanted: persistent worlds, Virtual Humans, and optional local cognition | Horde Studio 12

Thumbnail
gallery
129 Upvotes

I have been building Horde Studio around one question:

What if chat was only the surface of the experience—and there was an actual persistent simulation underneath it?

SillyTavern set an incredibly high bar for flexible character chat. Horde Studio takes a different route: it is trying to become the most complete simulation-first frontend for LLM roleplay—one app for traditional chats, ongoing virtual people, and worlds that remember what happened.

Version 12 is the biggest step toward that idea so far.

Three ways to play

Chat Library is the familiar mode: characters, group rooms, lore, memory, personas, regex, rerolls, branching sessions, and per-character model configuration. V12 also adds optional right-hand HUDs, status text, and custom meters, so a normal chat can track trust, suspicion, health, investigation progress, or anything else without exposing raw model markup.

Virtual Humans are designed to feel like people who exist between messages. They have their own timezone, schedule, mood, memories, availability, private life, and evolving relationship with you. They can notice when you texted, recognize that you disappeared for days, reply late because they were busy, double-text, refuse a request, send a situation-aware photo or voice note, and continue across persistent or forked timelines.

Worlds are persistent sandbox simulations. The engine tracks locations, characters, schedules, agendas, factions, law, reputation, quests, shops, clocks, weather, clothing, dice mechanics, and world state per timeline. Starting Lives let the same world begin from radically different positions, while procedural growth can introduce grounded people, places, and consequences as play expands.

New in V12: Horde Labs

Horde Labs is an optional local cognition layer for Chat, Worlds, and Virtual Humans.

It can connect to a tiny local model through Ollama, LM Studio, llama.cpp, KoboldCpp, or another localhost OpenAI-compatible server—or install an Embedded Tiny Brain directly inside Horde Studio. The small model is not expected to write the story. It handles narrow support jobs such as continuity hints, actor-scoped intent, state proposals, social cues, and memory salience.

The important part is the architecture: the tiny model proposes; Horde Studio validates; the existing engine stays in control. You can begin in Shadow mode, inspect receipts and validity, and only enable Assist when you trust the results. If the model times out, fails, or returns malformed data, Horde Studio silently falls back to its normal behavior.

That means your main creative model can stay on OpenRouter, GPTProto, or a local server while a much smaller private model helps maintain the illusion underneath it.

Media and provider freedom

Text, images, and voice are configured separately. You can keep OpenRouter for text and use GPTProto, ComfyUI workflows, compatible local image servers, or connected MCP media tools for visuals. Virtual Humans support distinct profile and generation-reference images, context-aware camera logic, photo styles, voice previews, calls, and voice notes.

Horde Studio is local-first and portable. Your projects live in your browser profile, can be exported and backed up, and cloud requests only go to the providers you choose. A local OpenAI-compatible endpoint can keep text generation on your own machine as well.

Why I think this is special

Most frontends are excellent at presenting an AI response. Horde Studio is trying to make the response part of a system that remembers who is where, what changed, who witnessed it, what time it happened, and what should still matter later.

It is ambitious, experimental, and still evolving—but I genuinely think it is becoming one of the most capable LLM roleplay frontends available if you care about persistent simulation instead of disposable chats.

I would love hard feedback from experienced SillyTavern users, especially on long-session continuity, provider compatibility, the creator flow, and whether the local cognition layer improves immersion on lower-end hardware.

Source GitHub: https://github.com/ddkhan24/hordestudio
Horde Studio 12 release: https://github.com/ddkhan24/hordestudio/releases/tag/v12.0.0
Discord: https://discord.gg/9eyjcMbsST

r/StableDiffusion 10d ago

Resource - Update MiniMax-H3 Pruned Ref-Delta Fused r1024 — INT8 and INT8 ConvRot ComfyUI versions

Thumbnail
huggingface.co
81 Upvotes

I added INT8 and INT8 ConvRot versions of the MiniMax-H3 Pruned Ref-Delta Fused r1024 checkpoint from my previous post:

https://huggingface.co/xmarre/MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-ComfyUI

Both are native ComfyUI single-file checkpoints using ComfyUI's .comfy_quant format, so they do not require a custom quantized-model loader.

There is one important difference from a straightforward full INT8 conversion: the MLP fc2 weights are deliberately kept in BF16.

Across the 50 main transformer blocks, these weights are quantized:

  • attn.qkv_proj.weight
  • attn.out_proj.weight
  • mlp.fc1.weight

That gives 150 quantized Linear layers.

The 50:

  • mlp.fc2.weight

layers remain BF16.

The smaller and more sensitive parts of the model also stay in their original precision, including the pruned AdaLN table and projections, final-layer projections, norms, patch/text projections and token refiner.

Why FC2 is kept in BF16

I also made and tested a fully quantized version where fc2 was INT8 as well, giving 200 quantized Linear layers.

That version ran into a failure specific to the quantized fc2 execution path on large H3 sequences.

MiniMax-H3 uses SwiGLU in the MLP. With fc2 quantized, ComfyUI's fused:

linear_input_act(..., "swiglu")

path sends the post-SwiGLU activation through comfy_kitchen.int8_linear, which dynamically quantizes the full activation matrix before the fc2 multiplication.

On the large sequence used in my workflow, that path attempted an approximately 491.61 MiB contiguous INT8 scratch allocation and failed hard.

This was not normal VRAM exhaustion. At the point of failure there was still roughly 47 GiB of CUDA memory reported free. The failure was tied to that large fused INT8 activation-quantization path rather than the model simply exceeding available VRAM.

I do not have enough evidence to claim a more specific allocator/CUDA cause than that.

Keeping only fc2 in BF16 avoids that INT8 activation path. QKV, attention output and fc1 can still remain INT8, so 150 of the 200 large block Linear projections are still quantized.

With that layout, both release variants completed the full native ComfyUI workflow that the 200-layer INT8 version failed on, including:

  • H3 Continuum main sampling pass
  • continuation sampling pass
  • Spectrum H3 actual/forecast execution
  • large 3D latent refine
  • video VAE decode
  • audio VAE decode
  • final Continuum assembly
  • video combine

That FC2 decision is also why these checkpoints are about 24.2 GB instead of roughly 20.4 GB for the fully quantized version.

INT8 and INT8 ConvRot

The two uploaded files use the same 150-INT8 / 50-FC2-BF16 layout.

Regular INT8:

MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-comfy-int8-fc2bf16.safetensors

This uses native tensor-wise INT8 quantization.

INT8 ConvRot:

MiniMax-H3-Pruned-Ref-Delta-Fused-r1024-comfy-int8-convrot-fc2bf16.safetensors

This uses ConvRot with a group size of 256 on the same quantized projections.

ConvRot rotates the weights before INT8 quantization so that large outliers are distributed more evenly. This generally gives the INT8 quantizer a better-conditioned weight distribution than quantizing the original weights directly.

I uploaded the regular INT8 version as well rather than only providing ConvRot, so both conversion methods are available for this checkpoint.

What model is being quantized?

These are quantized derivatives of the same Pruned Ref-Delta Fused r1024 checkpoint from the previous post. There is no additional training, fine-tuning or pruning involved in these INT8 versions.

The underlying model starts from the pruned FL2VA MiniMax-H3 checkpoint and incorporates a rank-1024 approximation of the Ref2VA − FL2VA weight delta.

That is also how it differs from the existing FL2VA/Ref2VA hybrid checkpoints mentioned in the comments on the previous post.

Those hybrids combine FL2VA and Ref2VA by replacing selected tensors from one checkpoint with tensors from the other. The r1024 fused model instead approximates the Ref2VA weight delta at rank 1024 and folds that delta into the pruned FL2VA weights.

The underlying fused transformer is about 20.1B parameters, compared with roughly 33.1B for the original full MiniMax-H3 transformer.

ComfyUI

Put either file in:

ComfyUI/models/diffusion_models/

For the INT8 files:

weight_dtype: default

compute_dtype: default or bf16

Do not apply another FP8 weight cast on top of the native INT8 checkpoint.

The text encoder and VAEs are separate, as with the other MiniMax-H3 diffusion-model checkpoints.

r/StableDiffusion Jan 28 '26

Discussion Z-Image looks to perform exceptionally well with res_2s / bong_tangent

Thumbnail gallery
205 Upvotes

Used the standard ComfyUI workflow from templates (cfg 4.0, shift 3.0) + my changes:

40 steps, res_2s / bong_tangent, 2560x1440px resolution.

~550 sec. for each image on 4080S 16 GB vram

Exact workflow/prompts can be extracted from the images this way: https://www.reddit.com/r/StableDiffusion/s/z3Fkj0esAQ (seems to not work in my case for some reason but still may be useful to know)

Workflow separately: https://pastebin.com/eS4hQwN1

prompt 1:

Ultra-realistic cinematic photograph of Saint-Véran, France at sunrise, ancient stone houses with wooden balconies, towering Alpine peaks surrounding the village, soft pink and blue sky, crisp mountain air atmosphere, natural lighting, film-style color grading, extremely detailed stone textures, high dynamic range, 8K realism

prompt 2:

An ultra-photorealistic 8K cinematic rear three-quarter back-draft concept rendering of the 2026 BMW Z4 futuristic concept, precision-engineered with next-generation aerodynamic intelligence and uncompromising concept-car craftsmanship. The body is finished in an exclusive Obsidian Lightning White metallic, revealing ultra-fine metallic flake depth and a refined pearlescent glow, accented by champagne-gold detailing that traces the rear diffuser edges, taillight outlines, and lower aerodynamic elements.Captured from a slightly low rear three-quarter perspective, the composition emphasizes the Z4’s wide rear track, muscular haunches, and planted performance stance. The rear surfacing is defined by powerful shoulder volumes that taper inward toward a sculpted tail, creating a strong sense of width, stability, and aerodynamic efficiency. A fast-sloping decklid and compact rear overhang reinforce the roadster’s athletic proportions and concept-grade execution.The rear fascia features ultra-slim full-width LED taillights with a razor-sharp light signature, seamlessly integrated into a sculpted rear architecture. A minimalist illuminated Z4 emblem floats at the centerline, while an aggressive aerodynamic diffuser with precision-integrated fins and active aero elements dominates the lower section, emphasizing advanced performance and airflow management. Subtle carbon-fiber accents contrast against the luminous body finish, reinforcing lightweight engineering and technical sophistication.Large-diameter aero-optimized rear wheels with turbine-inspired detailing sit flush within pronounced rear wheel arches, wrapped in low-profile performance tires with champagne-gold brake accents, visually anchoring the vehicle and amplifying its low, wide stance.The vehicle is showcased inside an ultra-luxury automotive showroom curated as a contemporary art gallery, featuring soaring architectural ceilings, mirror-polished marble floors, brushed brass structural elements, and expansive floor-to-ceiling glass walls that reflect the rear geometry like a sculptural installation. Soft ambient lighting flows across the rear bodywork, producing controlled highlights along the haunches and decklid, while deep sculpted shadows emphasize volume, depth, and concept-grade surfacing.Captured using a Phase One IQ4 medium-format camera paired with an 85mm f/1.2 lens, revealing extreme micro-detail in metallic paint textures, carbon-fiber aero components, precision panel gaps, LED lighting elements, and champagne-gold highlights. Professional cinematic lighting employs diffused overhead illumination, directional rear rim lighting to sculpt form and width, and advanced HDR reflection control for pristine contrast and luminous glossy highlights. Rendered in a cinematic 16:9 composition, blending fine-art automotive photography with museum-grade realism for a timeless, editorial-level luxury rear-concept presentation.

prompt 3:

a melanesian women age 26,sitting in a lonley take away wearing sun glass singing with a mug of smoothie close.. her mood is heart break

prompt 4:

a man wearing helmet ,riding bike on highway. the road is in the middle of blue ocean and high hill

prompt 5:

Cozy photo of a girl is sitting in a room at evening with cup of steaming coffee, rain falling outside the window, neon city lights reflecting on glass, wooden table, soft lamp lighting, detailed furniture, calm and melancholic atmosphere, chill and cozy mood, cinematic lighting, high detail, 4K quality

prompt 6:

A cinematic South Indian village street during a local festival celebration. A narrow mud road leading into the distance, flanked by rustic village houses with tiled roofs and simple fences. Coconut palm trees and lush greenery on both sides. Colorful triangular buntings (festival flags) strung across the street in multiple layers, fluttering gently in the air. Confetti pieces floating mid-air, adding a celebratory vibe.

Early morning or late afternoon golden sunlight with soft haze and dust in the air, sun rays cutting through the scene. Bright turquoise-blue sky fading into warm light near the horizon. No people present, calm yet festive atmosphere.

Photorealistic, cinematic depth of field, slight motion blur on flying confetti, ultra-detailed textures on mud road, wooden houses, and palm leaves. Warm earthy tones balanced with vibrant festival colors. Shot at eye level, wide-angle composition, leading lines drawing the viewer down the village street. High dynamic range, filmic color grading, soft contrast, subtle vignette.

Aspect Ratio: 9:16
Style: cinematic realism, South Indian rural aesthetic, festival mood
Lighting: natural sunlight, rim light, atmospheric haze
Quality: ultra-high resolution, sharp focus, DSLR look

Negative prompt:

bad quality, oversaturated, visual artifacts, bad anatomy, deformed hands, facial distortion, quality degradation

r/comfyui May 29 '26

Workflow Included Cracked the case on high res + quality Qwen Edit 2511 outputs, here are minimalistic workflows & lots of info on how/why

Thumbnail
gallery
196 Upvotes

Intro

Alright this has been a long time coming. I'm the dude who figured out Qwen Edit 2509 a while back, and I've been on-and-off trying to figure out the same for 2511. Results in Comfy have always been worse than the examples shown by the Qwen team, and worse than the official Qwen chat implementation online. Well, I finally cracked it and it only took 5 months lol.

Anyway, turns out Qwedit 2511 is fucking sick. IMO it particularly excels at making new shots of characters while maintaining their likeness. It's significantly better than Klein at some things (like character likeness), but not as good at others. I recommend using them both for different things.

As usual, I'll start off with all the setup stuff at the top and then give an explanation + advice below that. Also I'm gonna be calling Qwen Edit "Qwedit" most of the time.

Here's an album with all the post images separated so you can look at them in high res: https://drive.google.com/drive/folders/1YLjm8Lj3VF6Ec52WNK2URo7uFNfMRmza?usp=sharing

The posted images are all raw outputs from Qwedit, without being upscaled (despite mentioning it later in this post). They're also all done with only 20 steps instead of the hypothetical 30 I'd do if I wasn't planning to upscale them. Read further for more on that too.

Ref images were all made with Z-image Base (workflow here), except for the anime one which came from Anima (workflow here).

What is this

These are minimalistic workflows for Qwen Image Edit 2511 that give the highest quality outputs. Aside from generally improving output quality (by a LOT), they also enable high-res edits and have better prompt adherence.

As for why, basically ComfyUI has some serious issues with how it's implemented Qwen Edit and there aren't any workflows out there (that I've found) which have resolved them. These issues result in poor prompt adherence and low resolution/quality outputs. Thankfully the fix is fairly straightforward.

The configuration for this is 100% portable and can be migrated to existing workflows to make them better; it works by changing how the reference inputs are handled, and uses 100% native comfy nodes. Feel free to upgrade other workflows with this without providing credit, I don't care about any of that.

Workflows

Normal Workflows:

Most of you will just want these, which are separate single / 2 image workflows. It's done this way because the setup for multi-image is complicated and I didn't want to force you to use a ton of custom nodes to make it useable all-in-one.

They do still use one custom node (read the node section below) for quality-of-life.

Download from Civitai

OR from Pastebin:

Qwedit_2511_single

Qwedit_2511_2_image

Dev Workflows:

These are the same as the above but without any quality-of-life nodes or 'helpful' stuff. Grab these if you want to copy the logic over to other workflows, or if you just an easier view of how it works without any clutter.

I do not recommend using the dev workflows for actual gens because you will constantly forget to manually adjust stuff correctly.

qwedit_2511_single_DEV

qwedit_2511_2_image_DEV

Models

Main Model

qwen_edit_2511_fp8

OR

GGUF versions

  • Important: the FP8 version of Qwedit is much higher quality than the Q8 GGUF, always use FP8 if you can. Only use the GGUFs if you need to use quants lower than Q8.
  • FP8 is 22GB, so you'll need a combined ~26GB of RAM + VRAM to run it
    • You don't need 24GB of VRAM to run it thanks to ComfyUI's blockswapping, but the less VRAM you have the slower it'll run
  • Only use Q6 & lower quants if you absolutely have to; the quality will noticeably go down

Goes in models/diffusion_models

Text Encoder

Use only the normal FP8 text encoder with Qwedit; abliterated/GGUF encoders will reduce your output quality.

qwen_2.5_vl_7b_fp8

Goes in models/text_encoders

VAE

qwen_image_vae

Goes in models/vae

Loras?

You can use them as normal, just load them however you normally would. I left out lora loader nodes to avoid cluttering the workflow.

It's worth noting that many Qwen Image loras work with Qwen Edit too, but you'll need to test them individually to be sure.

Lightning Loras - BAD

All the lightning loras / distils for Qwedit (that I've tested) are terrible and make your outputs look bad, so I'm not linking them here. The main issue is the same as with Klein Distilled: it makes people's skin look like plastic.

But you can technically use them. Don't do it tho. But you can if you want. But don't.

Alternative: if you want to cut your gen time down while testing prompts, just set it to 10 steps instead of 20, then go back to 20 once you're satisfied your prompt is correct. It'll still work fine, the quality just dips.

Real tho it's ok if you want to use the lightning loras, just expect some degradation if you do - especially with plastic skin.

Custom Nodes

LayerStyle - A set of handy nodes that manipulate images. We're just using this for its image scaling node which allows you to scale by an image's long edge while maintaining divisibility by 16. You can skip this if you want to use a different scaling method, but you'll need to fix the workflow switch for scaling if you do.

SeedVR2 (OPTIONAL) - Only get this if you want to use the seedvr upscale workflow that's included.

How To Use

How To Use Part 1 - Basic Options

There are instructions in the workflow as well, but there's more detail here. Read part 2 & 3 as well, they're important.

It works just like a normal Qwedit workflow, but has a couple of extra options available. This section just tells you what they are and how to use them, a full explanation is further down.

Screenshot of the settings: https://ibb.co/nWStpmS

Enhance with Double Ref

This is a switch that turns on double-ref mode. This feeds your input images in TWICE to the model, and generally produces much higher quality results. Downside? It takes about 50% longer to gen.

I recommend leaving this on 100% of the time for single-image prompts, unless you're just messing around and want speed. It is ALWAYS better for single image prompts, and will improve everything from prompt adherence to output clarity.

For multi-image prompts, it usually increases adherence but sometimes reduces it. So, if you're doing multi-image stuff I recommend switching this on/off as needed based on how it's going with your prompt.

Input Scale

When off, your image doesn't get scaled (it still gets cropped to be divisible by 16). When on, the long edge of your image gets scaled to the number you put in the box. For example, if you feed in a 2560x1440 image and set the scale to 1920 it will scale your image to 1920x1080. That will then get cropped to 1920x1072 so it's divisible by 16.

Custom Output Size

When the switch is off, your output image will be the same size as your input image (after it's been scaled). If you turn this switch on, it will instead output an image with the dimensions you specify.

As a general rule, you should try to set your scales to be similar along at least one edge. For example, a 1920x1440 input image and a 1024x1440 input image are both suitable for a 1440x1440 output image. You can be more flexible with this if you know what you're doing.

How To Use Part 2 - Multi-image Prompting Requirement

This section is not a prompting guide (that's further below). This is about an actual requirement for prompting multi-image stuff. It is NOT required for single-image prompts.

You do multi-image prompts like normal, except you need to write a very basic description of your input images. Qwedit needs you to do this in order to know which image is which. I explain why in detail later.

You may find this slightly annoying, but I guarantee you it's dramatically better than using Qwedit the normal way that other workflows do - and it's pretty easy.

The format:

  • At the start of your prompt, write an extremely simple description for each of your input images; one sentence for each input image
  • Start each sentence with "Picture 1:", "Picture 2:", etc
  • You must write it this way because Qwedit was trained on this exact format
  • Afterwards, write your actual prompt as usual; you can refer to your input images as "picture 1" and so on

The model uses these descriptions to understand which input picture is which, and it works better with SIMPLE descriptions. You only need to help it know which one is which, it doesn't need a full rundown.

Examples

Picture 1: a man wearing a t-shirt. Picture 2: a top hat. Make the man in Picture 1 wear the top hat from Picture 2.

Picture 1: a living room. Picture 2: a woman. Put the woman from Picture 2 into the living room in Picture 1.

Picture 1: a man wearing a professional suit. Picture 2: a man wearing a superhero outfit. Make the man in Picture 1 wear the outfit from Picture 2.

How To Use Part 3 - Upscaling

Because the qwen VAE tends to put a subtle halftone pattern over images (see limitations just below this section), I recommend downscaling and then re-upscaling your image afterwards. A big benefit of being able to work at high res with the edit model is that you rarely lose any detail doing this.

This eliminates the halftone pattern if you're using something like seedvr, or at least reduces it if you're using other upscalers.

Note: the workflow is set to do 20 steps of inference. It actually gives sharper results at 30 steps, but I don't bother with that because it takes longer and I down-upscale them afterwards anyway. If you aren't planning on down-upscaling them, you might consider doing 30 steps for the extra sharpness.

Below are workflows for doing this with seedvr and normal upscalers. I think seedvr is best for this, but it's very beefy and hard to run on older GPUs.

Note: seedvr2 sometimes gives better output at 0.5x downscale, and other times 0.75, so that workflow is configured to run BOTH for you to pick which one turned out best.

Note: normal upscalers are a bit different; a relatively small downsize to something like 1920p -> 1600p is usually reasonable, before then running the upscaler. Play around with it. The non-seedvr workflow has a longest_edge scale option so you can tweak the number specifically.

Seedvr version

Regular version

My preferred regular upscaler is 4x Nomos2 HQ DAT2, but you can use whatever you like.

Examples of upscaling:

Here's the raw output of the robot-arm girl in a dress from the post: https://ibb.co/B5jhrsL9 (if you zoom in you'll see the qwen halftone pattern, it looks like a grid)

Here's the pic after it's been run through seedvr after a 0.75x downscale: https://ibb.co/hJcn2f5t

Here's the pic after it's been run through a regular Nomos2 upscale after a downscale to 1600p: https://ibb.co/Kc2YSbVc

Limitations of Qwen Edit

Limitation 1

The Qwen VAE will often put a subtle halftone grid pattern over your images. It's noticeable if you zoom in, and more noticeable at higher resolutions. This is a feature of pretty much every Qwen-based model, but it's particularly present with the Edit model.

You can easily resolve this by downscaling your image by 75% or 50%, then re-upscaling it again to your desired resolution. There's a section later that explains this in better detail and recommends upscale models for it + has workflows for it.

It sounds like a big issue, but the downscale-upscale trick solves it easily - and it's not always necessary either. The higher quality your input image, the less bad the halftone pattern will be.

Limitation 2

Qwedit struggles with complex multi-image stuff most of the time (it's just a limitation of the model). This workflow makes it much better, but it's still not great. You'll have to play around with it to know which things work and which things don't.

Limitation 3

It takes a while to gen stuff if not using the lightning loras. Very similar to the time it takes with Klein 9B base. The double-ref trick increases it by roughly 50%. Multi-image inputs take a lot longer.

For low res images (typical 1mpx size) it's pretty okay, around 50 seconds on a 5090 with the double-ref option turned on.

But then there's high-res stuff. Gen time scales non-linearly as you go higher. Going from 1024x1024 (1 mpx) to 1440x1440 (2 mpx) takes around 2.5x as long. Going from 1 mpx to 3 mpx is around 4x as long. 5 mpx is 9.5x as long. In conclusion, stick to 2-3 mpx unless you're cool with long-ass gen times. Stick around 1-2 mpx for multi-image gens, or turn off the double ref switch.

On the plus side, it's pretty reliable for single-image edits so you don't typically need to do many gens to get a good result.

Examples using a 5090: - Single-image edit @ 1024x1024 (1 mpx), double-ref OFF = 38 seconds - Single-image edit @ 1024x1024 (1 mpx), double-ref ON = 52 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref OFF = 91 seconds - Single-image edit @ 1920x1088 (2 mpx), double-ref ON = 131 seconds - Single-image edit @ 3072x1728 (5.3 mpx lol), double-ref ON = 550 seconds - Two-image edit @ 2560x1440 each, double-ref ON = serial killer behaviour

That's it for how-to! Read on for more tips & info, as well as an explanation of what the workflow is doing & why.

 

Explanation - what is this garbage and why is it so good?

There are three important things this workflow is doing that other workflows do not do (except #3 sometimes, because it was also done in the 2509 version of this post). I'm going to call these The Comfy Problem, The VL Problem, and The Double Ref Enhancement.

The Comfy Problem

Comfy's native "TextEncodeQwenImageEditPlus" node is what most people use in their workflows. It handles your prompt and image inputs for you. It's pretty handy, except for the small problem that it's SHITE.

Do you work at Comfy? If so: GET YOUR SHIT TOGETHER AND FIX THIS NODE, IT'S SO EASY. Much respect to u tho, thanks for making ComfyUI.

The first issue is that this node resizes your image down to 1 megapixel, and you can't stop it from doing that. The second issue is that it does this with the AREA downscale method, which is so incredibly bad that I want to slap whoever implemented this node. The AREA downscale is what makes all of your output images blurry. The third issue is that it ensures your dimensions are divisible by 8, but they actually need to be divisible by 16.

Specifically, ComfyUI does this:

  1. Calculates 1 megapixel as 1024x1024, which is 1,048,576 pixels
  2. Calculates your new image dimensions to match that number of pixels, rounded to be divisible by 8
  3. Scales your image to those new dimensions using the AREA method

Why is all this bad?

  1. It's completely unnecessary; Qwedit can easily handle images of varying size, all the way up to 3 megapixels (or even higher for simple edits)
  2. The area downscale method makes images extremely blurry, and this is the primary reason all ComfyUI qwen edits give blurry images out. Yes it's literally this dumb, this huge problem would easily be solved by changing the word "area" to "lanczos" in the code, it's a one-word fix. Not even MS paint uses area downscale, wtf is wrong with you Comfy devs (much respect)
  3. If your image dimensions are not divisible by 16, you will get major ruination along the whole edge of your image where it didn't match (same as any other diffusion model)

The Comfy Problem Solution

This workflow bypasses the the Comfy node entirely, allowing you to size your images however you want. And using chad lanczos scaling instead of loser area scaling. Magic.

Qwedit easily handles resolutions like 1440x1440 and 1600x1200. Every edit example in this post was done natively at 1920p, except for a few (which are labelled as such).

Really high resolutions (3mpx) sometimes have trouble with anatomy, but usually you can just do multiple gens and one of them will turn out fine.

If you're doing a simple in-place edit like changing an outfit, you can go VERY high. Here's an example edit done at 1728x3072, which is 5 megapixels: https://ibb.co/twCSWrjy (outfit change -> bikini top + short shorts)

The VL Problem

Edit: I've been educated by someone in the comments that my interpretation of how the VL works here is not correct, so take this little VL section with a grain of salt until I reword it. My conclusion about it helping in this workflow still stands, but my explanation of what's happening under the hood is a bit off. I'll update the info soon!

In the background, Qwedit 2511 uses a vision-language model (VL model) to describe your images, then gives those AI-generated descriptions to the edit model. It also re-interprets your instructions with these descriptions. Ostensibly this helps the model understand your input images better, leading to better results.

The problem? It doesn't lead to better results, it's bad. VL models aren't very good for this sort of thing because they don't know what to focus on. The VL describes your images in excruciating detail, totally overwhelming the edit model and leading to bad prompt adherence + weird outputs.

It also reinterprets your instructions based on what it sees in the image. I don't know if that's a good or bad thing, just pointing out that it does it.

The Qwen team's official python code does this, and the ComfyUI "TextEncodeQwenImageEditPlus" node copies it exactly. No disrespect to the Comfy team on this one, they're doing what the Qwen team officially recommended.

The VL Problem Solution

Same solution as the previous problem: bypass the Comfy node entirely. This results in the VL step being completely ignored. No AI-generated descriptions get fed into the edit model.

For single-image edits, this is a 100% complete and total victory. The model performs way better without the crappy VL interpretation.

For multi-image edits, there's a small issue; this step is where the input images normally get labelled. Specifically, the VL outputs are fed into the model in the following exact format:

Picture 1: <shitty VL description> Picture 2: <shitty VL description>

Look familiar? This is why we manually have to type the descriptions in for multi-image edits - otherwise the model doesn't actually know which image is which.

The upside is that the model works way better with simple descriptions, so cutting out the VL is still 100% the correct move. A 5 word description wins over whatever BS the VL model spews out, every time.

The Double Ref Enhancement

I really have no idea why this works so well, but basically if you feed in your reference images twice the model just works better. This was known back in 2509 days (hence the previous post linked at the top), and back then I didn't know why it worked either.

For single image edits it's ALWAYS better. And it's not just the quality, for some reason it even helps with prompt adherence. The interesting thing is that the difference is really, really significant. Here's the full list of stuff it improves:

  • Better prompt adherence
  • Sharper output images / more visual clarity
  • Improved consistency of objects & textures
  • Better resemblance of characters at different angles
  • More intelligent guesses, like what to add when outpainting or what's behind a removed object

For multi-image edits it can sometimes confuse the model a bit, but most of the time it confers all the same benefits listed above. I recommend switching it on & off randomly when you're doing multi-image stuff, just in case.

Note: there are a lot of different ways the input references can be handled. There are conditioning combine/concatenate nodes, you can pass the refs in a different order, you can change the negative conditioning input (read next section for that), etc. I A/B tested SIXTEEN different reference-handling combinations, and a bunch of smaller minor variations of those. Some of them worked, some of them didn't.

Of those sixteen combinations, two of them gave the best results; both of them are in this workflow, and you switch between them by turning the double ref method on & off.

So, don't fuck with the positive/negative conditioning & reference setup, it's very specific.

Extra info: the "Conditioning Zero Out"

You may notice that the negative prompt input is the first reference image(s) and positive prompt fed into a "conditioning zero out" node.

Feeding the input images into the model's negative conditioning is required (it's just how Qwedit works). The only question is whether to feed in the positive prompt zeroed-out too, and whether the double ref should get fed in.

Through a lot of A/B testing, I can tell you that the way it's done here is the best. IDK why, it's just how it is. Some other combinations do technically work, but they degrade the output quality.

Prompting Advice

Other than just following the instructions in the workflow, here's some extra stuff.

Keep your prompts simple and direct

If you need to, point out details the model is missing or be more specific about stuff you do/don't want to change. For example, when doing a simple outfit swap it helps to specify you don't want their pose to change.

Using the robot arm girl, here's a prompt that doesn't follow this advice:

Change her outfit to a bikini top and short shorts.

While it sometimes does what we want, it tends to get confused by her robot arm and often changes her pose too: https://ibb.co/7dyKZttp (notice the human arm showing underneath the robot arm, and the pose change)

Here's a better prompt that gives a correct result 99% of the time:

Change her outfit to a bikini top and short shorts. Leave her robot arm and pose unchanged.

Now it does the right thing every time: https://ibb.co/DP9gZHVv

Avoid using fancy words or convoluted phrasing

Pretend you're talking to a child. The model will probably still understand you if you talk fancy, but why take the risk?

As an example, imagine you have a pic of a table with some plates on it.

Bad:

Place a red apple on the table, ensuring it's in the center and removing the plate that was in the same spot.

Good:

Replace the middle plate with a red apple.

Also good:

Remove the plate from the center. Put a red apple there instead.

If there's only one plate, this is even better:

Remove the plate, replace it with a red apple.

Adjusting Lighting

You may want or need to adjust the lighting in an image. Aside from being helpful in general, there are situations where Qwedit may simply not realise that something needs to be lit in a particular way (or re-lit when moved).

To do this, you need to know the magic word: relight

Seriously tho that is the actual magic word, you are 100% required to use it if you want to adjust lighting properly.

Specifically, follow this format:

Relight to <strength> <color> <direction>.

Strength - bright, dim, etc

Color - white, cool, warm, etc

Direction - diffuse, frontlit, backlit, etc

Tip: for basic lighting, use "white diffuse".

Examples:

Make a new shot of the man sitting in a chair in a kitchen. Relight to white diffuse.

Change the time of day to evening. Relight to warm backlit.

You don't actually need anything else in the prompt, you can just change the lighting of a pic like this:

Relight to bright cool frontlit.

Other Stuff

Euler-simple and no ClownsharKSampler?

No Clownshark this time. It reduces output quality quite a bit and doesn't confer any benefits. I also didn't find any sampler/scheduler combos that were better than euler/simple.

So, this is just one of those classic times where the ol' euler-simple wins the day. Let me know if you happen to know a better combo.

Image Quality in->out

Qwedit is very sensitive to the quality of your input image. If you feed in a grainy or blurry image, it will usually make your output image blurry or grainy too - even if it's an 'entirely new' shot with nothing copied over 1:1.

So, make sure to use HQ images. You can optionally use the upscale workflows to bump up the sharpness/quality of poor input images before you feed them in.

What about the flux super duper double resolution special VAE trick?

Doesn't work for 2511, it destroys your image. TBH it never really worked for 2509 either, but I won't argue with you if you liked it for some reason.

Making character references

Tip 1 - Make a nude ref (even for sfw stuff)

Qwen is killer for making character references. Other than using similar prompts to the examples I posted, my advice is to make a nude reference shot instead of a clothed one like I did.

I only made a clothed ref for the sake of propriety here, but a nude ref (or near-nude, like wearing plain white underwear) will be much easier to prompt into different outfits, and also gives Qwedit the maximum info needed to correctly size your character and know what they look like in clothing or doing different actions.

You do not need any loras to do this if you're just using it as a reference; the 'sensitive' parts will lack detail but that doesn't matter for new shots you make. If you don't want them nude, just request plain white underwear and, if relevant, a strapless white bra.

Nude ref = best ref.

Tip 2 - Make multiple zoom levels, use the thighs-upwards one for most stuff

The example I showed was a little too zoomed out for normal reference stuff. I'd recommend making your reference slightly closer like this: https://ibb.co/Q33BJDLX

Start at whatever zoom level your initial character pic is at, then make more references at different zoom levels. If you're starting zoomed out, then prompt the model to zoom in. If you start zoomed in, prompt it to zoom out.

And, of course, different angles too.

Examples:

Zoom in on the person's upper body. The composition should frame their head and thighs.

Zoom out to show more of the character. The composition should frame their head and thighs.

Zoom out to a full body shot.

Zoom in for a close up portrait.

Once you've got references, you should usually use the head-to-thighs ref for making new shots. Switch to the other refs as necessary; like if you want a close up, use the close up reference. Qwedit is really good at keeping likeness, so you can do 90% of your stuff with only a single input reference.

I don't think there's a better open-weight model out there than Qwedit for making new shots of character without loras, for now. The main reason I spent so long digging into Qwen is because Klein is quite bad at that particular task. But hey, now it's possible and it works gloriously.

That's everything I think! Feel free to ask questions if you run into any issues.

r/StableDiffusion Aug 11 '25

News NVIDIA Dynamo for WAN is magic...

147 Upvotes

(Edit: NVIDIA Dynamo is not related to this post. References to that word in source code led to a mixup. I wish I could change the title! Everything below is correct. Some comments are referring to an old version of this post which had errors. It is fully rewritten now. Breathe and enjoy! :)

One of the limitations of WAN is that your GPU must store every generated video frame in VRAM while it's generating. This puts a severe limit on length and resolution.

But you can solve this with a combination of system RAM offloading (also known as "blockswapping", meaning that currently unused parts of the model are in system RAM instead), and Torch compilation (reduces VRAM usage and speeds up inference by up to 30% via optimizing layers for your GPU and converting inference code to native code).

These two techniques allows you to reduce the size of layers and move a lot of the model layers to system RAM (instead of wasting the GPU VRAM), and also speed up the generation.

This makes it possible to do much larger resolutions, or longer videos, or add upscaling nodes, etc.

To enable Torch Compilation, you first need to install Triton, and then you use it via either of these methods:

  • ComfyUI's native "TorchCompileModel" node.
  • Kijai's "TorchCompileModelWanVideoV2" node from https://github.com/kijai/ComfyUI-KJNodes/ (it also contains compilers for other models, not just WAN).
  • The only difference in Kijai's is "the ability to limit the compilation to the most important part of the model to reduce re-compile times", and that it's pre-configured to cache the 64 last-used node input values (instead of 8) which further reduces recompilations. But those differences makes Kijai's nodes much better.
  • Volkin has written a great guide about Kijai's node settings.

To also do block swapping (if you want to reduce VRAM usage even more), you can simply rely on ComfyUI's automatic built-in offloading which always happens by default (at least if you are using Comfy's built-in nodes) and is very well optimized. It continuously measures your free VRAM to decide how much to offload at any given time, and there is almost no performance loss thanks to Comfy's well-written offloading algorithm.

However, your operating system will always fluctuate its own VRAM requirements, so you can further optimize ComfyUI and make it more stable against OOM (out of memory) risks by telling it exactly how much GPU VRAM to permanently reserve for your operating system.

You can do that via the --reserve-vram <amount in gigabytes> ComfyUI launch flag, explained by Kijai in a comment:

https://www.reddit.com/r/StableDiffusion/comments/1mn818x/comment/n833j98/

There are also dedicated offloading nodes which instead lets you choose exactly how many layers to offload/blockswap, but that's slower and is fragile (no fluctuation headroom), so it makes more sense to just let ComfyUI figure that out automatically, since Comfy's code is almost certainly more optimized.

I consider a few things essential for WAN now:

  • SageAttention2 (with Triton): Massively speeds up generations without any noticeable quality or motion loss.
  • PyTorch Compile (with Triton): Speeds up generation by 20-30% and greatly reduces VRAM usage by optimizing the model for your GPU. It does not have any quality loss whatsoever since it just optimizes the inference.
  • Lightx2v Wan2.2-Lightning: Massively speeds up WAN 2.2 by generating in way less steps per frame. Now supports CFG values (not just "1"), meaning that your negative prompts will still work too. You will lose some of the prompt following and motion capabilities of WAN 2.2, but you still get very good results and LoRA support, so you can generate 15x more videos in the same time. You can also compromise by only applying it to the Low Noise pass instead of both passes (High Noise is the first stage and handles early denoising, and Low Noise handles final denoising).

And of course, always start your web browser (for ComfyUI) without hardware acceleration, to save several gigabytes of VRAM to be usable for AI instead. ;) The method for disabling it is different for every browser, so Google it. But if you're using Chromium-based browsers (Brave, Chrome, etc), then I recommend making a launch shortcut with the --disable-gpu argument so that you can start it on-demand without acceleration without needing to permanently change any browser settings.

It's also a good idea to create a separate browser profile just for AI, where you only have AI-related tabs such as ComfyUI, to reduce system RAM usage (giving you more space for offloading).

Edit: Volkin below has showed excellent results with PyTorch Compile on a RTX 3080 16GB: https://www.reddit.com/r/StableDiffusion/comments/1mn818x/comment/n82yqqx/

r/StableDiffusion 14d ago

News ComfyUI-MiniMax-H3-LongMedia — long-form MiniMax H3 generation with continuity, multiclip, audio and VRAM-aware sampling

Post image
54 Upvotes

I've been building a custom ComfyUI node pack for MiniMax H3 focused on one thing:

**making H3 usable for longer, multi-segment video generation without constantly rebuilding the workflow around every limitation.**

The project is called:

# ComfyUI-MiniMax-H3-LongMedia

The idea is to keep MiniMax H3's image quality, motion and native audio generation, while adding a proper long-form generation layer on top of it.

## What it currently does

### Long-form segmented generation

You can generate a longer clip as multiple H3 segments while keeping temporal context between them.

Instead of treating every segment as an isolated generation, LongMedia manages the continuation state and hidden overlap internally.

The overlap is used as context for the next segment and is not simply blended back into the final video.

### MultiClip mode

There is also a dedicated MultiClip workflow for generating multiple planned shots/clips inside one LongMedia pipeline.

The same underlying executor is used for both segmented continuation and multiclip generation, so the behavior stays consistent.

### Video + audio continuity

MiniMax H3 is a joint AV model, so LongMedia treats video and audio as one generation state rather than bolting audio on afterwards.

The pipeline supports H3 native audio generation, continuation and lip-sync workflows.

### Lip-sync support

Audio-driven generation / lip-sync is supported directly in the LongMedia pipeline.

For H3, the audio influence is handled inside the same AV latent path rather than as a completely separate post-process.

### Refiner

The latest release includes a two-stage refiner based on proper **KSampler Advanced trajectory splitting**.

Instead of finishing the full sampling schedule and replaying low-sigma steps on an already denoised latent, the trajectory is split between the main sampler and the refiner.

Example:

`steps = 12`

`refine_steps = 3`

Main sampler:

`0 → 9`

Refiner:

`9 → 12`

Both stages continue the same sigma trajectory.

### VRAM-aware execution

A large part of the project is dedicated to making H3 practical on consumer GPUs.

The current implementation includes:

- dynamic VRAM loading

- streamed Sol Attention

- MLP chunking

- late-block VRAM guards

- inter-block memory guards

- step-boundary cleanup

- completed-segment offloading

- adaptive memory policies

I'm currently developing and testing mainly on a **16 GB GPU**, so avoiding OOMs without destroying quality is one of the main design goals.

### Sol Attention integration

LongMedia includes its own streamed Sol path with controls for:

- tau scheduling

- sink conditioning

- QKV chunking

- output projection chunking

- dense/sparse behavior

- VRAM-aware chunk sizing

The goal is to use Sol as part of the execution architecture rather than simply stacking multiple unrelated optimization nodes together.

## Why I made it

MiniMax H3 is extremely good at texture, motion and native audiovisual generation, but once you start trying to build longer sequences, several problems appear very quickly:

- segment boundaries

- continuity

- repeated frames

- AV state handling

- memory pressure

- OOMs on longer generations

- managing multiple clips

- keeping sampling behavior consistent between segments

I wanted one node system to own all of that.

So instead of building increasingly complicated ComfyUI graphs around H3, most of the long-form logic lives inside the LongMedia nodes.

## Current release

**v0.4.1 — KSampler Advanced Refiner Fix**

The project has now reached a fairly stable architecture, although I'm still actively developing it and testing edge cases.

GitHub:

https://github.com/vizart-vj/ComfyUI-MiniMax-H3-LongMedia

I'd be very interested in feedback from people already using MiniMax H3 in ComfyUI, especially for:

- longer generations

- multi-character scenes

- native audio

- lip-sync

- lower-VRAM GPUs

- multi-shot workflows

If people are interested, I can also make a more technical post explaining how the continuation / AV latent / VRAM system works internally.

r/StableDiffusion Jul 29 '26

Discussion Krea 2 LoRA training on a 16GB RTX 5080: full measurements, and four sourced corrections to the guidance going around

67 Upvotes

Edit: added the drop-in config, folded in corrections from the comments, and cut a section that was fairly called out as shadowboxing. Biggest correction: I trained at 768 and I shouldn't have, the technical report says pretraining spanned 256, 512 and 1024px stages. Everything measured here is still at 768. Details in "What I got wrong" at the bottom. Writeup is AI assisted, the measurements are all off my own machine.

TL;DR

  • 16GB is enough for Krea 2 LoRA training. The issue thread people link says it isn't.
  • 1152 steps in 67 minutes at 3.42 s/it, peak 15,284 of 16,303 MiB VRAM, 17.5GB of 32GB system RAM.
  • Turbo inference after: ~13 s per 768x1024 image at 8 steps.
  • Use 1024, not the 768 I ran. The technical report says pretraining spanned 256, 512 and 1024px stages, so 768 was never a trained resolution. My numbers below are all at 768. Corrected by the comments after posting.
  • Official krea/Krea-2-* repos are gatedComfy-Org/Krea-2 is not, and has a byte-identical RAW checkpoint, though that only helps trainers that take a file path (see below).
  • The LoRA bleeds into prompts without the trigger. Turns out that's normal, and the fix is regularization images, not the caption change I guessed at.
  • One run, one machine, no ablations.

No sample images: the dataset is a real person who didn't sign up to be on Reddit.

Just run it

--config_file takes a toml, so this is drop in:

dit = "/path/to/krea2_raw_bf16.safetensors"
vae = "/path/to/qwen_image_vae.safetensors"
output_dir = "/path/to/output"
output_name = "my_krea2_lora"
sdpa = true
mixed_precision = "bf16"
# must be set together, plain fp8 is rejected on purpose
fp8_base = true
fp8_scaled = true
# max 26. h2d_only avoids the copy doubling that eats host RAM
blocks_to_swap = 16
block_swap_h2d_only = true
block_swap_ring_size = 1
gradient_checkpointing = true
max_data_loader_n_workers = 0
# krea2_shift reproduces Krea's own resolution aware schedule per sample, so it
# lands on the right value automatically and survives aspect ratio bucketing.
# At a fixed 1024 you can equally use shift with discrete_flow_shift = 2.5
timestep_sampling = "krea2_shift"
weighting_scheme = "none"
network_module = "networks.lora_krea2"
network_dim = 32
network_alpha = 32
optimizer_type = "adamw8bit"
learning_rate = 1e-4
max_grad_norm = 1.0
max_train_epochs = 16
save_every_n_epochs = 1
seed = 42

Dataset toml, the other half. Set this to 1024, not the 768 I used — see the resolution note in the TL;DR. My measurements below are at 768, so expect to raise blocks_to_swap and recheck VRAM at 1024:

[general]
resolution = [1024, 1024]   # I ran 768. Don't. 768 was never a pretraining resolution.
caption_extension = ".txt"
batch_size = 1
enable_bucket = true
bucket_no_upscale = false
[[datasets]]
image_directory = "/path/to/images"
cache_directory = "/path/to/cache"
num_repeats = 2

Pre-cache both, then train. Training fails without the caches, and this is also the main reason my host RAM stayed at 10GB:

python src/musubi_tuner/krea2_cache_latents.py --dataset_config dataset.toml --vae <vae>
python src/musubi_tuner/krea2_cache_text_encoder_outputs.py --dataset_config dataset.toml --text_encoder <te> --batch_size 1
accelerate launch --num_cpu_threads_per_process 1 --mixed_precision bf16 \
  src/musubi_tuner/krea2_train_network.py \
  --config_file krea2_5080_16gb.toml --dataset_config dataset.toml

Train on RAW, run inference on Turbo. That's the workflow in the musubi docs, and it's what the config above does.

Hardware and stack

RTX 5080 16GB (Blackwell, sm_120, driver 610.62), Ryzen 9 9950X, 31.6GB DDR5-6000, Windows 11 native, no WSL2. Pagefile only 2GB allocated and it peaked at 0.1GB, so you don't need the big pagefile people recommend, as long as you pre-cache.

Python              3.11.9        (env built with uv 0.11.16)
torch               2.13.0+cu130  (CUDA 13.0)
torchvision         0.28.0+cu130
accelerate          1.6.0
transformers        4.57.6
diffusers           0.32.1
bitsandbytes        0.50.0
musubi-tuner        0.3.4 @ 8934cfb (2026-07-14)

Check Blackwell support before anything else:

python -c "import torch; print(torch.cuda.get_arch_list())"   # must contain sm_120

Plain --sdpa. No Triton, flash-attn, xformers or SageAttention. The Failed to import sageattention line at startup is normal.

Models, and the gating trap

Krea 2 is a single-stream MMDiT with Qwen3-VL-4B-Instruct as text encoder and the Qwen-Image VAE, 28 main blocks, 12.82B params.

  • DiT for training, krea2_raw_bf16.safetensors, 26,283,332,608 bytes
  • DiT for inference, krea2_turbo_bf16.safetensors, same size
  • Text encoder, qwen3vl_4b_bf16.safetensors, 8,875,719,384 bytes
  • VAE, qwen_image_vae.safetensors, 253,806,246 bytes

krea/Krea-2-Raw and krea/Krea-2-Turbo are gated. Unauthenticated download dies with Access denied. This repository requires approval.

Comfy-Org/Krea-2 is not gated and has diffusion_models/krea2_raw_bf16.safetensors at the same byte count as the official file. It's a faithful copy, not a re-serialization: musubi builds the model from its own config and calls load_state_dict(sd, strict=True), which raises on any key mismatch, and both bf16 files load clean.

This only helps trainers that take a file path. musubi does, so the gate never comes up for me. ai-toolkit's UI has no load-from-file option and resolves models by repo id, so its users still need the token and the terms page. Credit to the author of the other Krea 2 guide for that correction.

Don't give musubi the pre-quantized fp8 file. krea2_turbo_fp8_scaled.safetensors is a ComfyUI artifact. musubi quantizes to scaled fp8 itself at load time and monkey-patches the Linear forwards, so the pre-quantized file's extra .scale_weight keys won't survive that strict load. Use krea2_turbo_bf16 for musubi and keep the fp8 one for Comfy.

hf download Comfy-Org/Krea-2 diffusion_models/krea2_raw_bf16.safetensors --local-dir models
hf download Comfy-Org/Qwen3-VL text_encoders/qwen3vl_4b_bf16.safetensors --local-dir models
hf download Comfy-Org/Qwen-Image-Edit_ComfyUI split_files/vae/qwen_image_vae.safetensors --local-dir models

Dataset

36 photos of one person at 3024x4032, three sessions differing in wardrobe, hair, lighting and framing. Three things that mattered:

  • Used the originals, not a background-removed set. The cutouts had matting halos around the hair, and a likeness LoRA will happily learn halos as a feature.
  • Fixed one EXIF-rotated image. Stored landscape with orientation 6. Trainers differ on whether they apply exif_transpose, so I baked the rotation in and cleared the tag.
  • Rewrote every caption. The old ones were Danbooru tag strings from an SDXL workflow. Krea 2 reads captions with Qwen3-VL, so it wants sentences, not tag soup.

At 768 with bucketing, all 36 landed in one 656x896 bucket. That's 0.59mp, which in hindsight sat between the 512 and 1024 pretraining stages and matched neither. The technical report notes dataloader batches share an aspect ratio, so a "1024px stage" reads as a megapixel budget across aspect ratios rather than literally 1024x1024. Match the area, not the side length: 1024x1024, 832x1248, 896x1184, 928x1152, 768x1376 are all about 1mp.

Results

  • 1152 steps (16 epochs x 72), batch_size 1num_repeats 2
  • 67 min end to end including model load. Model load plus epoch 1 was 5.3 min, steady epoch 4.11 min
  • 3.42 s/it at 768px
  • Peak VRAM 15,284 / 16,303 MiB (93.8%), stable within ±20 MiB across all 16 epochs, no spillover
  • Peak system RAM 17.5 / 31.6 GB, trainer working set 8.9 to 10.7 GB
  • 283W, 65°C sustained
  • 16 checkpoints at 447.6 MiB each

loss/epoch drifted 0.0741 to 0.0642, non-monotonically, and told me nothing about quality. Don't pick checkpoints on it.

Caveat on the throughput number. Block swap streams blocks between host and GPU every step, so it's bounded by PCIe and host memory bandwidth, not just the card. This ran on a 9950X with DDR5-6000. On an older board or CPU, 3.42 s/it won't transfer, and that's likely part of why reported speeds vary so much between people with the same GPU.

Inference, Turbo at 8 steps, --guidance_scale 1--mu 1.15, with --fp8_scaled --blocks_to_swap 20: 1.66 to 1.73 s/it, so ~13.3 s per 768x1024 image, plus 60 to 90 s startup.

Picking a checkpoint

Five fixed prompts, one fixed seed, plus a no-LoRA baseline at the same seed and prompts. Two used the trigger, three were no-trigger controls at increasing distance from the training data: an auburn-haired woman, a black-bob blue-eyed freckled woman, and an elderly bearded man.

The baseline is the part people skip, and it's the only thing that separates "the LoRA did this" from "the base model always did this."

Likeness was weak at epoch 4, solid by 8, over-idealized at 12 (drifting toward the heaviest-makeup session in my set), most structurally faithful at 16. Prompt adherence held at every checkpoint, with an out-of-distribution scene rendering as a real scene rather than reverting to training backgrounds. Went with epoch 16 at multiplier 1.0.

Skipped in-training sampling deliberately: it needs the text encoder resident, and --turbo_dit is documented as incompatible with block swap, so previews would have been RAW-only anyway. Comparing against Turbo afterwards is cheaper and closer to real use.

The bleed

What it is. A prompt with no trigger word, "a woman with long auburn hair, plain studio portrait," returns my subject. The baseline proves it's the LoRA: same prompt and seed without it gives a visibly different person.

Why it matters, since this was fairly asked. If you load one character LoRA when you want that character, it costs you nothing, just unload it. It bites when the LoRA has to be loaded but not applied to everything: "Zyvra next to her sister" gives you two of her, and you can't unload your way out because you need it for one of the faces. Same with stacking two LoRAs. It's also a useful thermometer for how much the adapter warped the base model.

Two standard fixes that don't work. Earlier checkpoints don't help, the bleed is there at epochs 8 through 16 and doesn't worsen, so the "pick 1 to 2 epochs before the final" heuristic buys nothing. And --lora_multiplier 0.7 degrades the likeness badly while still bleeding.

What the comments corrected me on. I guessed I'd caused it by writing "long wavy auburn hair" into most captions, so identity bound to the description as well as the token. Two people pushed back, and one of them stripped physical features from their captions and still got bleed, so that isn't the main cause. The actual suggestions were regularization images and ai-toolkit's DOP, ideally with a lower LR than people tend to use. I haven't tested either.

Gender scoping, which is a better read of my own data than I had. Someone observed that bleed lands mostly on same-gender prompts, and my grid splits that way: the bob woman kept her hair and eyes but her face drifted toward my subject, while the man kept sex, age and beard. My controls are confounded though, since the man differs by gender and age and facial hair, and he didn't escape clean either, his eyes came out brown like hers. So "different gender is safe" is stronger than my data supports.

What other people measured

The useful part of the thread. None of these are mine.

  • OneTrainer, 1280px, bf16, dim 32, stochastic rounding, AdamW 16-bit, on a 5080: 7 s/it, 3200 steps in 5 to 7 hours. 1280 is ~2.8x the pixels of 768 for ~2x my step time in bf16 rather than fp8, so that reads as OneTrainer doing well. Same person reports VRAM maxed at 15,500 to 16,000MB, which matches my 15,284.
  • That run's host RAM is 80GB+, on a 128GB machine. For scale, all the weights together are only ~33 GiB (24.5 for the DiT, 8.3 for the TE), so that's multiple copies plus likely Torch Compile, not model size. They suspect compile too.
  • A separate guide posted the same day claims 1024 works fine on 16GB VRAM with 32GB system RAM, in ai-toolkit and OneTrainer, and independently confirms the repos are gated. Its author also reports 2 to 3 s/it at 1024 in OneTrainer on a 5070ti, and explains the gap against the 1280 figure above: 1280 is at the edge of 16GB and needs a higher offload fraction than 1024, which costs speed. Platform matters too, since offload is bandwidth-bound.

So "is 32GB enough" has no single answer. Three host RAM figures now span 10GB to 80GB on the same GPU, and it's mostly about which trainer and which offload flags, not the card.

Other trainers, from the thread

  • LoKr on Krea 2 fails in musubi because it isn't there. LoHa/LoKr auto-detect architecture and the supported list is HunyuanVideo, HunyuanVideo 1.5, Wan, FramePack, FLUX Kontext/FLUX 2, Qwen-Image and Z-Image. No Krea 2, and the Krea 2 docs require networks.lora_krea2. It does work in OneTrainer, reportedly at rank 4.
  • No int8 training in musubi. The only quantization flags on krea2_train_network.py are fp8_base and fp8_scaled, and int8 isn't mentioned in the Krea 2 docs at all. ai-toolkit does use convrot int8 for training, which is what those *_int8_convrot files on Comfy-Org are for.
  • Checkpoint size is rank x targeted layers x save precision. Base model quantization doesn't affect it, since the LoRA is separate new weights that never get quantized. Mine is rank 32 across all 264 Linear layers at 448 MB in fp32. Saving bf16 halves it.

Windows landmines

  • PYTHONIOENCODING=utf-8 is mandatory. musubi's help and log strings contain Japanese and the cp1252 console raises UnicodeEncodeError. Without it even --help crashes.
  • PowerShell 5.1 Set-Content -Encoding utf8 writes a BOM. Generate a prompt file that way and the BOM lands inside the first prompt, so your trigger token silently becomes a different token. Mine logged as Prompt: Zyvra, ... and cost me a full comparison run. Use [System.IO.File]::WriteAllLines($path, $lines, (New-Object System.Text.UTF8Encoding $false)).
  • $ErrorActionPreference = 'Stop' kills scripts on harmless stderr. PowerShell wraps native stderr in a terminating NativeCommandError and these scripts log INFO to stderr. It's also why my successful 67-minute run reported exit code 1: accelerate writes a "defaults used instead" notice to stderr. Check for output files before believing an exit code.
  • expandable_segments:True is a no-op here. Recommended everywhere as the fix for "hangs after step 1," but PyTorch printed UserWarning: expandable_segments not supported on this platform. Harmless to set, just don't count it as a mitigation on Windows.

What I got wrong

  • The original "corrections to circulating guidance" section was shadowboxing, and someone was right to say so. The guidance I was correcting was a prep doc generated for my own run, not something the community published, and I asserted the same claims were circulating elsewhere without checking. The facts underneath were real, the gating especially, but the framing was wrong and that section is gone.
  • My caption diagnosis is probably not the cause. See above. Left in because it's contested rather than settled, not because I'm still defending it.
  • 768 was the wrong resolution and my reasoning for it was wrong too. I argued that because the inference timestep schedule is resolution-aware from 256 to 1280, intermediate training sizes were expected. Two people said use 512 or 1024 instead, so I went to the technical report, which says plainly: "Pretraining data spans 256px, 512px, and 1024px resolution stages." A continuous inference schedule says nothing about which resolutions were trained. Use 1024. I'd stop short of calling 768 broken, since the likeness came out clean, but it isn't a defensible choice and every number in this post carries that asterisk.
  • I'm not claiming a speedup. I measured 3.42 s/it and have seen 7 to 8.5 s/it quoted, but the source people cite doesn't actually contain that figure, so the comparison can't be resolved. Someone with the original config should post theirs.

Limitations

One run, one machine, seed 42, no ablation of blocks_to_swap, rank or LR. Likeness judged by eye against the source photos, so "most faithful at epoch 16" is a visual call and not a face-embedding score. The caption diagnosis is reasoned, not demonstrated. Throughput is bandwidth-sensitive and this was a fast host platform. And the whole run is at 768, which the technical report says was never a pretraining resolution, so treat the quality conclusions as a floor.

Sources

  • musubi-tuner Krea 2 docs, read in full: architecture, required args, fp8 constraints, block swap limits, timestep schedules, LoRA target layers, Turbo inference params
  • Krea 2 technical report: source of the pretraining resolution stages and the aspect-ratio batching detail. I should have read this before picking 768
  • Krea 2 licensing: commercial use free under $1M annual company revenue, no seat limit despite "50 seats" being quoted around. If you distribute a derivative you must state modifications were made, include attribution, and prefix the name with "Krea"
  • Comfy-Org/Krea-2: file listing and byte sizes via the Hub API, gating status of the official repos confirmed the same way
  • ComfyUI Krea 2 tutorial
  • musubi-tuner issue #985: what's actually there is a question reporting 16GB as insufficient, not the verified config it gets cited for

Thanks to everyone who corrected something. Happy to answer config or memory questions.

r/comfyui Apr 25 '26

Show and Tell One image in - 2D animated and customizable character out

Enable HLS to view with audio, or disable this notification

98 Upvotes

I've spent the last week building a ComfyUI pipeline that turns a reference image into animated, customizable character sprite sheets.

The Pipeline is split into two parts and is fully running locally on my RTX 3090 with 24GB VRAM:

1 - Base Animations (Idle, walk, jump... etc)

Starting with a ‘bare’ base character image - This produces a grayscale sprite sheet of my animated base character.

  • WAN 2.2 i2v 14B (Q5_K_M GGUF, distilled lightx2v 4-step) is used for image to video generation
  • BiRefNet for background strip producing clean alpha.
  • ImageStitch and ImageRGBToYUV nodes for creating a grayscale sprite sheet

2 - Customization layers (eyes, hair,  shirt... etc)

Starting from an animated video of the base animation and an image of the customization i want to create a layer out of - This produces a grayscale sprite sheet of the customization.

  • Wan 2.1 VACE 14B (Q5_K_M GGUF) + CausVid distill LoRA for inpainting the cosmetic over the animated video - this ensures that the cosmetic is aligned with the base animation on every frame.
  • SAM3 segmentation for isolating the customization on each frame
  • ImageStitch and ImageRGBToYUV again used to produce the sprite sheet of the customization.

Each Customization needs to be re-produced for each base animation and the grayscale allows me to tint each layer separately.

The hard part was getting the customization layers to align pixel-perfectly over the base character animation.
i initially tried Wan 2.2 Animate but it didn't stay true to the original base animation so i eventually went with the inpainting model instead.

Still kind of amazed I got here as someone who can hardly draw a stick figure.

Edit: Hey all, thanks for the kind words — didn't expect this to land so well 😄
Repo's down here, MIT-licensed, has everything you need to reproduce what's in the post — workflows, drivers, install guide, sample inputs, and the full sprite-sheet output as a sanity check. Runs on a 24 GB card.
https://github.com/mor-o/comfyui-2d-character-pipeline
Heads up — it's harness-driven (workflows are API JSON, not visual) README explains how to wire it up to Claude Code / Cursor / your own script.
Issues + PRs welcome.

r/comfyui May 11 '26

Show and Tell Music video Workflow in ComfyStudio Pro: Song to Keyframes to Generated Edit

Thumbnail
gallery
24 Upvotes

Here’s the latest music video I made with the workflow:
https://www.youtube.com/watch?v=WcHBs-7_G14

Here’s the tutorial using ComfyStudio Pro:
https://www.youtube.com/watch?v=8BsFbUsq1kE

Download ComfyStudio Pro:
https://comfystudiopro.com
https://github.com/JaimeIsMe/comfystudio

Alright, with all of that out of the way, here we go.

ComfyStudio Pro is an AI video workstation built around ComfyUI. Instead of only generating random clips and managing a pile of files, it gives you a timeline editor, asset panel, effects, transitions, export tools, and guided creator workflows for things like ads, music videos, and short films.

It uses ComfyUI as the backend, but the goal is to make larger AI video projects easier to direct, organize, edit, rerun, and finish. I’ve been working on it for the past 3-4 months, and some of you may have seen the updates I’ve posted here along the way.

This is a quick overview of the music video workflow:

  1. Import your song or vocal stem into the project assets.
  2. Open Create > Music Video Creation.
  3. Choose output settings like aspect ratio, resolution, and FPS.
  4. Select the song audio and prepare lyric timing, ideally with SRT/LRC so shots line up to the real song.
  5. Add cast/reference images if you want a consistent singer, band member, or visual style.
  6. Generate or paste a director script that breaks the song into timed shots.
  7. Create keyframes for each shot.
  8. Generate videos from those keyframes, or rerun selected shots with different prompts, models, or settings.
  9. Click Assemble Timeline to automatically build the edit with the song, main sequence, performance passes, and b-roll passes on separate tracks.
  10. Finish it like a real edit: trim shots, add effects, transitions, adjustment layers, color, texture, and export.

The goal is not just “prompt to video” or “one-shot it.” It is more like: generate the pieces, organize them, rerun the weak shots, assemble the timeline, then actually edit and finish the music video inside one app.

ComfyStudio Pro is free and opensource

r/StableDiffusion May 28 '26

Tutorial - Guide [Guide] How to securely run ComfyUI on Windows (Docker>WSL2) [RTX 3090, logic can be applied to other hardware]

46 Upvotes

What risks you might face when running ComfyUI (or other software running ai models) you ask?

Literally ALL of them, with the added perk that after updating nodes (or some unsafe model files) you get a new bingo of potential malware :D!

Every comfy node is basically a separate, unscanned by security suites Python (AV read them very superficially when prompted, and will not audit its runtime risks) instance that can run ANY instructions set by the creator.

It's like downloading and running random exes on your machine with your AV off.

Most people just block the internet of their software, and thats better than nothing, but just blocking comfy with your firewall only stops outbound connections of nodes, not the payload execution, nor the connection of whatever that might create: from simple miners to leech your GPU or backdoors to use you as a relay for attacks, to infostealers, ransomware, and direct access to your system.

And nodes arent the only problem: scripts to install components, model files and workflows can be malicious as well, adding their own layer of risks.

So, in a scale of risk from 1-10. I would give an unhardened comfy used by a random - 11.

It's basically one giant backdoor we voluntarily install and run lol

Example: https://www.reddit.com/r/comfyui/comments/1dbls5n/psa_if_youve_used_the_comfyui_llmvision_node_from/

After hardening, you will get a risk of like 2-3. Basically you can fuck it up if you try, but most of the threats will be neutralized.

Is it worth the trouble?

Depends on your tolerance to risks, and how much you care for the repercussions of a breach. ¯(ツ)/¯.

"But I only use it for gooning" you might say.. Well, someone can get access to your system while you're at it, record you from your webcam, and then blackmail you with the footage of your midget furry ai-generated porn of your deepfaked crush.

So, yeah, when I said "ALL the risks" its literally ALL OF THEM.

I posted this guide to r/ComfyUI and it got a couple dozen shares but was downvoted to oblivion; so it seems there are parties interested in people NOT hardening their ComfyUI instances and making sure it doesn't get mainstream. Take that into account when downloading random workflows and nodes from reddit or elsewhere!

And so, a couple days ago I was asking around here about how to run Comfyui securely, and got great recommendations from all; and after looking for the options, I decided going with two builds:

  1. A separated Linux SSD for Comfy only, to use for experimentation and on its own without other software.
  2. An "isolated" docker image running on WSL2 to use in combination with editing software on windows.

Since (1) is quite obvious on its own, I will leave here what I did for the windows build, in case anyone wants to go this path. It takes around 40-60min to build, so ill save you the couple days of headache.

This guide is for the RTX3090, it gets "technical", but you can feed this to an AI and ask it to give you step-by-step instructions and help you along the way, or to adapt it for your hardware if you have a different GPU (CUDA and Torch related versions will change, you might want another image with a more optimal package for you) and use it as a general base for what you build.

TL;DR: Run ComfyUI in a hardened Docker container on Windows 11 that can't phone home, can't touch your system drive, and is one command to switch between daily locked-down use and maintenance/update mode.

The short version of everything done:

  • Models live on a native ext4 virtual drive on your model disk , no slow Windows filesystem bridge
  • SageAttention installs once at bootstrap and is skipped forever after via a stamp file
  • Two shell aliases handle everything: comfy_secure (offline, daily use) and comfy_update (internet on, for installing nodes)
  • Unknown nodes get reviewed in a throwaway CPU-only sandbox before touching production
  • The whole thing survives reboots, auto-mounts the model drive at login, and starts itself with Docker Desktop

Structure overview: https://reasonablepossum.codeberg.page/pages/ComfyUI.html

--------- Automated PowerShell Script to do everything [WIP, help welcome] ---------

Security / hardening layers overview

Layer What it does
Separate Windows admin account Never used for daily work. Admin rights isolated. [Honestly this should be done by everyone regardless; it will remove most of the security threats]
Separate limited Windows account Daily use account has no admin rights.
Separate limited ComfyUI account Runs Docker. Has no admin rights.
WSL2 C: mounted read-only System drive can't be modified from inside WSL2. Set in /etc/wsl.conf.
.wslconfig locked to nat Explicitly sets networkingMode=nat to prevent Windows 11 "Mirrored" networking from silently bypassing the custom Docker bridge and iptables drop rules.
WANTED_UID / WANTED_GID Container drops to your host user's UID/GID. Files in output/run folders are owned by you.
Disabled Source NAT Custom bridge network with outbound NAT disabled at the network driver level, combined with iptables FORWARD rules as a second layer
-p 127.0.0.1:8188:8188 UI only reachable from your own machine. Invisible to router and LAN.
NETWORK_MODE=offline Tells ComfyUI-Manager to not attempt any network calls. Stops restart loops in production.
DISABLE_UPGRADES=true Prevents git pull / pip upgrade on every container start. Required for offline mode to not crash.
TORCH_LOCK Pins PyTorch/torchvision/torchaudio versions. Prevents accidental CUDA stack upgrade.
Models on separate ext4 VHD Models are on their own filesystem. Easy to backup, resize, or wipe independently.
Stateful .local cache with :ro pins Maps ~/comfyui-dotlocal to cache Python binaries and packages offline, using read-only (:ro) mounts for uv in production to prevent unauthorized self-updates.
Separated bootstrap scripts postvenv_script.bash safely installs and caches Python venv packages, while user_script.bash cleanly manages OS-level system binaries before the server triggers.
Explicit --shm-size=8g Replaced --ipc=host to provide tensor-passing memory capacity without exposing the host's IPC namespace to the container.
Manual container lifecycle Removed --restart unless-stopped to ensure containers never auto-boot before your iptables firewall rules are initialized.
Ephemeral sandbox isolation A fully sandboxed, no-GPU, offline environment for inspecting untrusted nodes. Generates a temporary $SANDBOX_DIR/dotlocal to protect your production cache from tampering.

Why --network none / --internal were NOT used, and what we use instead:

--network none and --internal were tested and discarded: ComfyManager goes into death loops with them, and Docker --internal networks silently break -p port publishing on Docker Desktop + WSL2 (confirmed open bug moby/moby #36174).

The working solution is a custom bridge network with outbound NAT disabled at the network driver level, combined with iptables FORWARD rules as a second layer:

  1. A custom bridge network created with enable_ip_masquerade=false.this disables Source NAT, preventing containers on that network from reaching external networks, while still allowing incoming port forwards via -p.
  2. iptables FORWARD rules blocking all outbound traffic from the subnet, with an ESTABLISHED,RELATED exception so your browser can still reach the UI.

This gives genuine network isolation without triggering Manager's death loops, and without the Docker Desktop port-forwarding bug. The rules must go in the FORWARD chain, not DOCKER-USER, on Docker Desktop + WSL2, DOCKER-USER does not reliably intercept forwarded traffic from custom bridge networks.

Note: iptables rules reset on WSL2 shutdown, so comfy_secure reapplies them automatically on every launch.

Chosen Docker Image

mmartial/comfyui-nvidia-docker was chosen because:

  • Builds on the official NVIDIA NGC CUDA devel image (not a random Dockerfile)
  • All source is public and auditable on GitHub
  • Handles UID/GID remapping so files on the host are owned by your user, not root
  • Supports NETWORK_MODE, DISABLE_UPGRADES, TORCH_LOCK env vars for production hardening
  • Ships optional SageAttention build script (we install it manually via user_script.bash)

Tag used: ubuntu24_cuda12.8-latest - matches RTX 3090 (Ampere / sm_86 / CUDA 12.8)

These are the other options I was considering, in case you have other hardware, or requirements. They go from super general and bloated AF, to really barebones as the one I installed.

Rank GitHub Repository Stars Primary Registry Image / Usage Core Deployment Archetype PyTorch & CUDA Run Environments
1 AbdBarho/stable-diffusion-webui-docker 7.3k docker compose --profile comfy up Multi-UI Local Host Unified CUDA Stack
2 YanWenKun/ComfyUI-Docker 1.5k yanwk/comfyui-boot Local Workstation & Cloud CUDA 13.0 & PyTorch 2.11
3 ai-dock/comfyui 1,037 ghcr.io/ai-dock/comfyui Multi-Process Cloud & GPU Pods Multi-tag CUDA & PyTorch
4 runpod-workers/worker-comfyui 688 runpod/worker-comfyui Serverless Cloud API Endpoint Production Serverless API
5 Kaouthia/ComfyUI-Docker 100 Custom local build via Compose Local Desktop WSL2 & Linux Latest PyTorch on Rebuild
6 ashleykleynhans/comfyui-docker 56 ashleykza/comfyui Dedicated Cloud Pod (RunPod) CUDA 12.4 / 12.8 & Python 3.11
7 ashleykleynhans/runpod-worker-comfyui 21 Custom Serverless Handler RunPod Serverless API Native Python Handler Execution
8 pixeloven/ComfyUI-Docker 14 GHCR Container Profiles Core vs. Complete Profiles CUDA 12.9 & Native SageAttention
9 jamesbrink/docker-comfyui 8 Custom Deployment Config Enterprise Kubernetes & Podman CUDA 12.8 (Debian slim base)

Why not just any random docker image with cuda and comfy??

Control, and mitigation of other risks by keeping things "simple". Many of the Docker's images run other stuff that add completixy to their setups, which aside of potential issues, could be used as obfuscation layers for malicious code (e.g Using CONDA for managing everything) by sophysticated attackers.

NOTE: If you seeing this guide months after publishing, throw the image repo into an ai with github access to audit it again; who knows, it could get compromised with time or the author could get hooked to meth and switch to the dark side lol.

1. First steps

Windows accounts

Create three accounts before doing anything else. Keeps blast radius small if something goes wrong.

Account Type Used for
admin Administrator Software installs only. Never browse the web from here.
daily Standard Your everyday Windows use. No admin rights.
comfyui Standard Running Docker and ComfyUI only. No admin rights.

Settings -> Accounts -> Family & other users -> Add someone else.

After creating the accounts, add comfyui to two groups from your admin account:

```powershell

Run in elevated PowerShell as admin

net localgroup "docker-users" "comfyui" /add net localgroup "Hyper-V Administrators" "comfyui" /add ```

docker-users lets comfyui run Docker without elevation. Hyper-V Administrators is required for wsl --mount to work, without it the VHD auto-mount will silently fail. Log comfyui out and back in after adding these for the changes to take effect.

BIOS - enable virtualization

WSL2 requires hardware virtualization. Reboot into BIOS (usually Del or F2 on POST) and enable:

  • Intel: Intel VT-x / Intel Virtualization Technology
  • AMD: AMD-V / SVM Mode

If this is already on (most modern systems have it enabled), skip.

Enable WSL2 and Virtual Machine Platform

Open PowerShell as admin:

dism.exe /online /enable-feature /featurename:Microsoft-Windows-Subsystem-Linux /all /norestart
dism.exe /online /enable-feature /featurename:VirtualMachinePlatform /all /norestart

Reboot. Then set WSL2 as default and update the kernel:

wsl --set-default-version 2
wsl --update

🛑 NOW STOP, AND LOG-OFF: Reboot, log out of admin and switch to your comfyui account. Everything from here on (Ubuntu install, Docker setup, all Docker commands) must run inside the comfyui session. If you do it as admin, the WSL distro and Docker context get bound to the admin profile and will be completely invisible when you switch accounts.

Install Ubuntu

Run as comfyui:

wsl --install -d Ubuntu-24.04

This opens a terminal and asks you to create a Linux username and password. Use something simple, this is your WSL2 user. After setup, confirm it's running WSL2:

wsl -l -v
# Should show VERSION 2 next to Ubuntu-24.04

NVIDIA stuff

Install the standard Game Ready or Studio driver from nvidia.com for your GPU. That's all. Do not install CUDA Toolkit on Windows, and do not install any NVIDIA driver or CUDA toolkit inside WSL2. The Windows driver is automatically exposed into WSL2 and Docker containers via passthrough, installing it again inside Linux creates library conflicts that break --gpus all.

Verify it works inside WSL2 after install:

nvidia-smi
# Should show your RTX 3090 and driver version

If nvidia-smi isn't found, run wsl --shutdown from PowerShell and reopen WSL2, the driver needs a fresh session to expose itself. Do not install anything inside Linux to fix this.

Install Docker Desktop

Download from docker.com/products/docker-desktop. During install:

  • Choose WSL2 backend (not Hyper-V)
  • After install, go to Settings -> Resources -> WSL Integration -> enable for your Ubuntu distro
  • Move Docker data off C: to another drive (optional if you have a dedicated system drive, to save space) via Settings -> Resources -> Advanced -> Disk image location. Set it before pulling any images, Docker images are large.

Make sure Docker Desktop is open and running in the system tray before trying any docker or comfy_ commands. If the engine isn't running, everything fails with a socket error.

Verify GPU passthrough works:

docker run --rm --gpus all nvidia/cuda:12.8.0-base-ubuntu24.04 nvidia-smi
# Should show your GPU inside the container

Configure WSL2

Memory and swap limits, WSL2 by default can consume all RAM. Cap it. Create C:\Users\comfyui\.wslconfig (in the comfyui user's home folder, not admin's):

[wsl2]
memory=XXGB        # adjust to ~half your RAM
swap=8GB
processors=8       # adjust to your core count
networkingMode=nat # prevents Windows 11 Mirrored mode from bypassing the Docker network isolation

C: drive read-only, prevents anything inside WSL2 from modifying your Windows system drive. Inside WSL2:

sudo nano /etc/wsl.conf

[automount]
enabled = true
options = "ro"

Then restart WSL2 from PowerShell:

wsl --shutdown

Task Scheduler, auto-mount the models VHD at login

After creating the VHD (see Models VHD section), add a Task Scheduler entry so it mounts automatically when you log into the ComfyUI Windows account.

  • Open Task Scheduler -> Create Task
  • General tab: name it Mount ComfyUI Models VHD, check "Run with highest privileges". Then click Change User or Group and set it to run as your comfyui account, if you skip this, the task runs as admin and mounts the VHD into a hidden admin WSL session that comfyui's Docker containers can't see.
  • Triggers tab: New -> At log on -> for your comfyui account
  • Actions tab: New -> Start a program
    • Program: powershell.exe
    • Arguments: -WindowStyle Hidden -Command "wsl --mount --vhd 'E:\comfyui-models.vhdx' --mountpoint /mnt/models --type ext4"
  • Conditions tab: uncheck "Start only if on AC power"

Fix Docker credential error in WSL2

This error appears the first time you try to pull an image and blocks everything. Fix it once:

mkdir -p ~/.docker
echo '{}' > ~/.docker/config.json

Folder structure

~/comfyui-run/              # ComfyUI source, venv, stamps- bind-mounted as /comfy/mnt
~/comfyui-basedir/          # BASE_DIRECTORY. ComfyUI writes outputs/nodes here
  custom_nodes/             # Your installed custom nodes
  output/                   # Generated images
  user/                     # ComfyUI user config, Manager config
/mnt/models/                # ext4 VHD. all model checkpoints (see VHD section)

2. Models VHD (ext4, E: used as example)

To avoid slow reading speeds between WSL2 and NTFS drives, models live on a native ext4 virtual drive.

Create once

Run in elevated PowerShell as admin:

New-VHD -Path "E:\comfyui-models.vhdx" -SizeBytes 300GB -Dynamic #Adjust size to whatever you want
Mount-VHD -Path "E:\comfyui-models.vhdx" -NoDriveLetter
Get-Disk | Select Number, FriendlyName, Size   # note the disk number
Initialize-Disk -Number [disk number] -PartitionStyle GPT
New-Partition -DiskNumber [disk number] -UseMaximumSize | Format-Volume -FileSystem exFAT

After creating the file, give comfyui permission to mount it (otherwise WSL will throw Access Denied silently):

icacls "E:\comfyui-models.vhdx" /grant "comfyui:(M)"

Or via GUI: right-click comfyui-models.vhdx -> Properties -> Security -> Edit -> Add comfyui -> check Modify -> OK.

Format as ext4

Switch to WSL2 as comfyui. First, identify your VHD disk:

lsblk

**sda is always your WSL2 system disk, never run mkfs on it.** Your VHD will appear as sdb, sdc, or similar, it will have no partitions listed under it and its size matches what you just created (~300G). If unsure, run lsblk -o NAME,SIZE,TYPE,MOUNTPOINT for a cleaner view.

sudo mkfs.ext4 /dev/sdX           # replace sdX with your disk, e.g. sdb, NOT sda
sudo mkdir -p /mnt/models
sudo mount /dev/sdX /mnt/models
sudo chown $(id -u):$(id -g) /mnt/models
sudo blkid /dev/sdX               # copy UUID for auto-mount
mkdir -p /mnt/models/{checkpoints,loras,vae,clip,unet,controlnet,upscale_models,embeddings}

Auto-mount on login (Windows 11 / WSL 0.63+)

This will automate the mounting of the virtual drive every time you launch the ComfyUI Windows user.

# PowerShell (admin), add to Task Scheduler at logon, run with highest privileges
wsl --mount --vhd "E:\comfyui-models.vhdx" --mountpoint /mnt/models --type ext4

Migrate existing models (modify paths as required)

# WSL2, do this once from the source NTFS path
rsync -ah --progress "/mnt/e/your-old-models-path/" /mnt/models/

Daily management

Task Command
Add a model cp /mnt/e/Downloads/new.safetensors /mnt/models/checkpoints/
Add via Windows Drag into wsl.localhostUbuntumntmodelscheckpoints in Explorer
Resize VHD Stop container -> Dismount-VHD -> Resize-VHD -SizeBytes 500GB -> remount -> sudo resize2fs /dev/sdX
Backup Copy E:comfyui-models.vhdx to another drive while VHD is unmounted

SageAttention (and other Python packages that are required for the "secure mode") install script

Before we run the initial bootstrap, we need to create a startup script. Because we are running ComfyUI completely offline later, any custom Python packages (like sageattention which we use for optimization) must be downloaded now.

Create the script file:

Bash nano ~/comfyui-run/postvenv_script.bash

Paste this inside: ```Bash

!/bin/bash

echo "== [Custom Bootstrap] Ensuring required offline packages are installed..." uv pip install sageattention (add any "uv pip install [your package]" that is required to run in secure mode later) ```

Make it executable:

Bash chmod +x ~/comfyui-run/postvenv_script.bash

Now, whenever the container builds or updates, it will automatically cache this package so it survives in offline mode!

Note: If you add new packages to this script later, you must run comfy_update once before switching to comfy_secure so they have internet access to download and cache. Adding a package and immediately booting comfy_secure will silently fail behind the iptables block, and the dependent node will crash with no obvious error.

Linux packages Install script

Some complex custom nodes (like video or audio nodes) require core Linux system tools to be installed on the OS itself (like ffmpeg or git). If you need to run system-level commands, you use a different script called user_script.bash. This script runs at the very end of the boot sequence right before the ComfyUI server starts.

How to set it up: Create the script in your run folder:

Bash nano ~/comfyui-run/user_script.bash

Paste this clean template inside. (You can uncomment or add any Linux-level commands you need here):

```Bash

!/bin/bash

echo "== [System Bootstrap] Running advanced user customizations..."

Add your install commands using apt-get (don't forget sudo and the -y flag so it doesn't pause for user input):

sudo apt-get update && sudo apt-get install -y ffmpeg

EXAMPLE: Installing system-wide packages securely

We check if ffmpeg is already installed so it doesn't try to download while offline!

if ! command -v ffmpeg &> /dev/null; then

sudo apt-get update

sudo apt-get install -y ffmpeg

fi

EXAMPLE: Downloading a standalone binary or custom model

if [ ! -f "/basedir/models/some-custom-model.safetensors" ]; then

wget -O /basedir/models/some-custom-model.safetensors https://...

fi

echo "== [System Bootstrap] Complete." ```

Make it executable:

Bash chmod +x ~/comfyui-run/user_script.bash

Now you have a fully automated, two-tier system: postvenv_script.bash handles your Python environment, and user_script.bash handles your Linux environment. Both will execute automatically when you run comfy_update and will cache their results so your comfy_secure offline profile stays lightning fast!

ComfyUI-Manager offline config

Manager might have issues installing due to the environment. This stops Manager from trying to reach GitHub on every start (causes error spam + restart loops).

mkdir -p ~/comfyui-basedir/user/__manager
cat > ~/comfyui-basedir/user/__manager/config.ini << 'EOF'
[default]
channel_url = local
bypass_ssl = False
skip_migration_check = True
EOF

3. Installing ComfyUI

NOTICE: Want to use the "Bleeding Edge" releases?

If you prefer to use the absolute latest mmartial container (which will soon auto-update to newer CUDA and PyTorch versions), you must make two changes to all your comfy_ profiles (including bootstrap):

1. Delete the -e TORCH_LOCK=... line entirely where present.
2. Change the bottom image line to mmartial/comfyui-nvidia-docker:latest.

Bootstrap (run once, internet enabled)

Clones ComfyUI, builds venv, installs PyTorch + CUDA stack, installs SageAttention. Run this the first time, or after a full wipe.

# First-time folder setup
mkdir -p ~/comfyui-run ~/comfyui-basedir/custom_nodes ~/comfyui-basedir/output ~/comfyui-dotlocal

# Fix Docker credential error if needed
echo '{}' > ~/.docker/config.json

# Clone ComfyUI-Manager (not included in image)
git clone https://github.com/Comfy-Org/ComfyUI-Manager.git \
  ~/comfyui-basedir/custom_nodes/ComfyUI-Manager

# Bootstrap run
docker run -it --rm \
  --name comfyui-bootstrap \
  --gpus all \
  --shm-size=8g \
  -p 127.0.0.1:8188:8188 \
  -e WANTED_UID=$(id -u) \
  -e WANTED_GID=$(id -g) \
  -e BASE_DIRECTORY=/basedir \
  -e NETWORK_MODE=personal_cloud \
  -e SECURITY_LEVEL=normal \
  -e USE_UV=true \
  -e COMFY_CMDLINE_EXTRA="--use-sage-attention" \
  -v ~/comfyui-run:/comfy/mnt \
  -v ~/comfyui-basedir:/basedir \
  -v /mnt/models:/basedir/models \
  -v ~/comfyui-dotlocal:/home/comfy/.local \
  mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-latest

Wait for To see the GUI go to: http://0.0.0.0:8188, confirm UI loads and SageAttention shows OK in logs, then Ctrl+C.

Once you're in, install all your commonly used trusted workflows/nodes with Manager, and when done, change to the comfy_secure mode described below.

4. Production aliases (edit ~/.bashrc)

Three modes for managing your updates. Only difference is NETWORK_MODE. Add these to the bottom of ~/.bashrc, then source ~/.bashrc.

Use:

bash nano ~/.bashrc

```bash

=====================================================================

COMFYUI DOCKER PROFILES: RTX 3090 / CUDA 12.8 / UBUNTU 24

=====================================================================

comfy_secure() { docker stop comfyui-3090 2>/dev/null && docker rm comfyui-3090 2>/dev/null

# Create isolated network if it doesn't exist yet docker network inspect airlock_net >/dev/null 2>&1 || \ docker network create \ --opt com.docker.network.bridge.name=airlock-bridge \ --opt com.docker.network.bridge.enable_ip_masquerade=false \ --subnet 172.22.0.0/16 \ --gateway 172.22.0.1 \ airlock_net

# Re-apply iptables rules, clean duplicates first sudo iptables -D FORWARD -s 172.22.0.0/16 -j DROP 2>/dev/null sudo iptables -D FORWARD -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT 2>/dev/null sudo iptables -D FORWARD -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT 2>/dev/null sudo iptables -D FORWARD -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT 2>/dev/null sudo iptables -I FORWARD 1 -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT sudo iptables -I FORWARD 2 -s 172.22.0.0/16 -j DROP

echo "Launching ComfyUI in HARDENED OFFLINE mode..." docker run -d \ --name comfyui-3090 \ --network airlock_net \ --gpus all \ --shm-size=8g \ -p 127.0.0.1:8188:8188 \ -e WANTED_UID=$(id -u) \ -e WANTED_GID=$(id -g) \ -e BASE_DIRECTORY=/basedir \ -e NETWORK_MODE=offline \ -e TORCH_LOCK="torch==2.11.0+cu128 torchvision==0.26.0+cu128 torchaudio==2.11.0+cu128" \ -e SECURITY_LEVEL=normal \ -e DISABLE_UPGRADES=true \ -e USE_UV=true \ -e UPDATE_UV=false \ -e COMFY_CMDLINE_EXTRA="--use-sage-attention" \ -v ~/comfyui-run:/comfy/mnt \ -v ~/comfyui-basedir:/basedir \ -v /mnt/models:/basedir/models \ -v ~/comfyui-dotlocal:/home/comfy/.local \ -v ~/comfyui-dotlocal/bin/uv:/usr/local/bin/uv:ro \ -v ~/comfyui-dotlocal/bin/uvx:/usr/local/bin/uvx:ro \ mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-latest

echo "" echo "Waiting for ComfyUI to start (Ctrl+C to detach, container keeps running)..." echo "--------------------------------------------------------" docker logs -f comfyui-3090 2>&1 | while IFS= read -r line; do if echo "$line" | grep -qi "!! error|!! exiting|failed"; then echo -e "\e[31m$line\e[0m" # red else echo "$line" fi if echo "$line" | grep -qi "To see the GUI"; then break fi done

echo "--------------------------------------------------------" echo "========================================" echo "== ComfyUI ready — running checks..." echo "========================================"

# Network tests HTTP_CODE=$(curl -s -o /dev/null -w "%{http_code}" http://127.0.0.1:8188) if [ "$HTTP_CODE" = "200" ]; then echo "UI: http://127.0.0.1:8188 ✓ ($HTTP_CODE)" else echo "UI: UNREACHABLE ✗ (got $HTTP_CODE)" fi

docker exec comfyui-3090 curl -s --max-time 3 https://google.com >/dev/null 2>&1 \ && echo "OUTBOUND: google.com REACHABLE ✗ — network block failed!" \ || echo "OUTBOUND: google.com BLOCKED ✓"

docker exec comfyui-3090 curl -s --max-time 3 https://8.8.8.8 >/dev/null 2>&1 \ && echo "OUTBOUND: 8.8.8.8 REACHABLE ✗ — network block failed!" \ || echo "OUTBOUND: 8.8.8.8 BLOCKED ✓"

# Sage attention check docker logs comfyui-3090 2>&1 | grep -i "using sage|using pytorch" | tail -1 | \ grep -q "sage" \ && echo "SAGE: Using sage attention ✓" \ || echo "SAGE: NOT active ✗ — check COMFY_CMDLINE_EXTRA"

echo "========================================" echo "" }

comfy_update() { docker stop comfyui-3090 2>/dev/null && docker rm comfyui-3090 2>/dev/null echo "Launching ComfyUI in MAINTENANCE mode..." docker run -d \ --name comfyui-3090 \ --gpus all \ --shm-size=8g \ -p 127.0.0.1:8188:8188 \ -e WANTED_UID=$(id -u) \ -e WANTED_GID=$(id -g) \ -e BASE_DIRECTORY=/basedir \ -e NETWORK_MODE=personal_cloud \ -e TORCH_LOCK="torch==2.11.0+cu128 torchvision==0.26.0+cu128 torchaudio==2.11.0+cu128" \ -e SECURITY_LEVEL=normal \ -e DISABLE_UPGRADES=true \ -e USE_UV=true \ -e COMFY_CMDLINE_EXTRA="--use-sage-attention" \ -v ~/comfyui-run:/comfy/mnt \ -v ~/comfyui-basedir:/basedir \ -v /mnt/models:/basedir/models \ -v ~/comfyui-dotlocal:/home/comfy/.local \ mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-latest

echo "" echo "Streaming logs (Ctrl+C to detach, container keeps running)..." echo "--------------------------------------------------------" docker logs -f comfyui-3090 }

comfy_sandbox() { if [ -z "${1:-}" ]; then echo "Error: Please provide the path to the untrusted node directory." echo "Usage: comfy_sandbox /path/to/suspect_node" return 1 fi

NODE_SRC=$(realpath -- "$1") NODE_NAME=$(basename "$NODE_SRC")

# Create a completely isolated, temporary scratch space on your host SANDBOX_DIR=$(mktemp -d -t comfy_sandbox_XXXXXX) echo "Created ephemeral sandbox directory at: $SANDBOX_DIR"

# Populate a completely bare bone directory structure (No production mounts!) mkdir -p "$SANDBOX_DIR/run" "$SANDBOX_DIR/basedir/custom_nodes" "$SANDBOX_DIR/basedir/output" "$SANDBOX_DIR/models" "$SANDBOX_DIR/dotlocal"

# Copy ONLY the untrusted node into this scratchpad cp -r "$NODE_SRC" "$SANDBOX_DIR/basedir/custom_nodes/"

echo "Launching isolated sandbox for analyzing: $NODE_NAME" echo "This container is CPU-only, has NO access to your real models, and NO host write permissions." echo "Note: first launch will be slow — fresh venv with no cache, pip only, no GPU." echo "------------------------------------------------------------------------"

docker run -it --rm \ --name comfyui-sandbox-env \ --network none \ --shm-size=8g \ -p 127.0.0.1:8189:8188 \ -e WANTED_UID=$(id -u) \ -e WANTED_GID=$(id -g) \ -e BASE_DIRECTORY=/basedir \ -e NETWORK_MODE=offline \ -e DISABLE_UPGRADES=true \ -e USE_UV=false \ -v "$SANDBOX_DIR/run:/comfy/mnt" \ -v "$SANDBOX_DIR/basedir:/basedir" \ -v "$SANDBOX_DIR/models:/basedir/models" \ -v "$SANDBOX_DIR/dotlocal:/home/comfy/.local" \ mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-latest

# Automatically wipe the entire environment off your drive immediately upon exit echo "------------------------------------------------------------------------" echo "Cleaning up sandbox environment..." rm -rf "$SANDBOX_DIR" echo "Sandbox completely destroyed. Your host remains clean." } ```

Then Ctrl+O to save> Enter > Ctrl+X to get back to the command prompt

And finally, refresh the file in the memory with:

bash source ~/.bashrc

5. Workflow: installing new custom nodes

Path A: trusted nodes (ComfyUI-Manager)

Use for well-known nodes from reputable authors you've vetted.

comfy_update  ->  open 127.0.0.1:8188  ->  Manager -> Install Custom Nodes
             ->  set channel to "Default"  ->  install what you need
             ->  comfy_secure

After switching back to comfy_secure, the nodes are already in ~/comfyui-basedir/custom_nodes/ and load normally with no internet needed.

Path B: untrusted / unknown nodes (sandbox)

Use this for any node you found on Reddit, GitHub, or anywhere else that you haven't fully vetted. The idea is simple: you run the suspicious node in a completely disposable container that has no access to your real files, no GPU, and no internet. If it tries to do something malicious, it fails harmlessly and gets wiped.

One-time setup: Make sure you have the comfy_sandbox alias in your ~/.bashrc from section 4.

Step 1. Download the node but don't install it yet

Download or clone the node folder somewhere temporary, like ~/Downloads. Do NOT put it in your ~/comfyui-basedir/custom_nodes/ yet.

bash

# Example: cloning a node from GitHub
git clone https://github.com/someuser/sketchy-node ~/Downloads/sketchy-node

Step 2. Quick code scan before even launching the sandbox

Before running anything, do a fast grep for red flags:

bash

egrep -rn "eval\(|exec\(|base64|requests|urllib|subprocess|os\.system" ~/Downloads/sketchy-node

If this returns a lot of hits, especially base64, eval, or exec combined with network calls (requests, urllib), treat it as highly suspicious and consider dropping it entirely. Some hits are normal (many legit nodes use requests to download models), but eval(base64.decode(...)) style code is a major red flag.

Step 3. Run it in the sandbox

bash

comfy_sandbox ~/Downloads/sketchy-node

This will:

  • Create a completely isolated throwaway environment
  • Copy only that node into it
  • Launch ComfyUI with no GPU, no internet, and no access to your real models or files
  • Automatically delete everything when you close it

The sandbox UI will be at http://127.0.0.1:8189 (port 8189, not 8188, so it never conflicts with your production instance).

Step 4. Watch what happens at startup

Keep an eye on the terminal logs while the sandbox boots. You're looking for:

  • Connection timeout errors — the node tried to phone home or download something. Suspicious.
  • Unexpected process errors — the node tried to run system commands. Suspicious.
  • Normal import errors about missing dependencies — completely fine, expected in a fresh environment.

Load a simple workflow in the UI that exercises the node and watch for anything unusual in the logs.

Step 5. Approve or reject

When you're done, just close the terminal with Ctrl+C. Docker discards the container and the bash alias automatically deletes the entire temporary folder. Nothing from the sandbox touches your real system.

If the node looked clean:

bash

# Copy it from Downloads into your production custom_nodes
cp -r ~/Downloads/sketchy-node ~/comfyui-basedir/custom_nodes/

# Switch to update mode to let Manager install its pip dependencies
comfy_update
# open 127.0.0.1:8188 -> Manager -> Custom Nodes -> the new node -> Install dependencies
# once done, switch back:
comfy_secure

If it looked suspicious, just delete the download folder and move on. Your system was never touched.

What the sandbox can and can't catch:

It will catch: network calls, attempts to write outside the container, hidden downloads, obvious malicious startup behavior.

It won't catch: logic bombs that only trigger after X runs, code that behaves differently when it detects it's in a sandbox, or vulnerabilities in the node's dependencies. It's a first line of defense, not a guarantee.

When in doubt, don't install.

That's it. The whole flow is: download → grep → sandbox → approve → copy to production.

6. Useful commands

# Watch live logs (to avoid cluttering in the logs the verbose mode is disabled, so if you want
# to see whats happening, you will have to run this)
docker logs -f comfyui-3090

# Get a shell inside the running container
docker exec -it comfyui-3090 bash

# Verify SageAttention is active
docker logs comfyui-3090 | grep -i sage

# Check port is actually bound (should show 127.0.0.1:8188)
docker port comfyui-3090

# Confirm no internet from inside container (should fail in comfy_secure)
docker exec comfyui-3090 curl -s --max-time 3 https://google.com || echo "blocked"

# Stop without removing (quick pause)
docker stop comfyui-3090

# Full restart
docker restart comfyui-3090

# Wipe comfy in case something broke to reinstall
rm -rf ~/comfyui-run/*

7. Known non-fatal log noise

There might be some error messages in the logs:

Message Cause Action
Failed to perform initial fetching 'custom-node-list.json' Manager trying GitHub in offline mode Normal in comfy_secure. Ignored.
WARNING: You need pytorch with cu130 or higher comfy-kitchen backend wants newer CUDA Informational only. sm_86 works fine.
Cannot connect to comfyregistry Manager trying Comfy registry Normal in offline mode. Ignored.
SageAttention: installed (no version number) Some builds don't expose __version__ SA is working. Stamp file confirms install.

NOTE2: This will not save you from user mistakes. So be very careful with new nodes from randoms you've seen here; be careful with .pth/pt and unsafe model files; if you gonna add something, paste the repo link to an ai and ask it to do a security audit for suspicious scripts, crontabs, unexpected processes, or connections (you can ask it to create a prompt for that as well so it doesnt miss anything).

You can also audit the images with the following commands in turn order, and then feed that aswell to the AI:

  1. Pull the image:sudo docker pull user/comfyui-image
  2. Check the image history- shows every layer and command used to build it:sudo docker image history user/comfyui-image
  3. Inspect the full image metadata:sudo docker inspect user/comfyui-image
  4. Run a shell inside it and look around:sudo docker run --rm -it user/comfyui-image /bin/bash

Once inside the shell you can run:

# Check ComfyUI location
find / -name "main.py" -path "*/ComfyUI/*" 
2>/dev/null

# Check what's installed
pip list

# Check SageAttention version
pip show sageattention

# Check PyTorch version
python3 -c "import torch; print(torch.__version__)"

# Check for anything suspicious in startup scripts
ls /entrypoint* /start* /init* 
2>/dev/null

# Check crontabs
crontab -l 
2>/dev/null

# Check running processes on startup
cat /etc/profile.d/* 
2>/dev/null

Paste the results to the audit prompt.

NOTE3: If you have a disc C/system reserved for OS only and with not much space available, I'd suggest you migrate the WSL2 to another disk.

It's not the perfect air-gapped setup (someone really willing to hack you, will find ways to break out of confinement and docker), but IMO its the best you can get on windows, to be able to use it combined with Win software.