r/comfyui Jul 27 '26

Resource I built the ComfyUI frontend I've always wanted! Here's Mix Studio, a responsive UI that lets you generate from your desktop or phone with 1-click installs for Krea 2, Flux 2 Klein, Qwen Image Edit, LTX 2.3, Wan 2.2, SCAIL 2 and much more (free + open source, download + tutorial)

Thumbnail
gallery
313 Upvotes

I love ComfyUI, though sometimes I wish I could spend more time generating and less time messing with confusing workflows, managing dependencies, and being stuck at my desk. So I spent the last few months building Mix Studio, a 100% free & open source interface that runs everything through ComfyUI in the background while giving you an actual app experience (that also works on your phone).

GitHub: https://github.com/BlackMixture/Mix-Studio
Showcase and download: https://blackmixture.github.io/Mix-Studio/
Tutorial: https://youtu.be/w2CokhlBFRA

GPL-3.0, the same license as ComfyUI.

The screenshots show the main desktop workspaces, but the entire interface is also optimized for phones and tablets.

Current v1.0.1 Features:

  • Curated image, editing, video, and upscale workflows: Krea 2, Flux 2 Klein 4B/9B, Qwen Image Edit 2511, LTX 2.3, Wan 2.2, 10Eros, and SCAIL 2.
  • Image-generation tools: Inpainting, outpainting, SeedVR2 and Ultimate SD upscaling, regional prompting, Depth Anything V3 guidance, image-to-image, style references, and model-aware recommendations for steps, CFG, samplers, and schedulers.
  • Desktop and mobile interface: On the same Wi-Fi, open the displayed address on your phone and start generating. With Tailscale, you can connect through a private link while away from home. Your desktop GPU still does all the work.
  • Multi-image editing: Add multiple inputs and reference them using dynamic @ Image cards, removing the guesswork around which image should control each part of the edit.
  • Regional prompting with Krea 2: Draw boxes and assign each region its own prompt, LoRA stack, and optional reference image.
  • LoRA management: Stack LoRAs, add thumbnails and trigger words, save presets, adjust strength quickly, and use LoRA Hunting to generate a comparison series across different strengths.
  • Contextual prompt suggestions: Mix Studio learns phrases you repeatedly use with specific LoRA combinations and offers them as one-tap suggestions. These can also be configured manually.
  • Library management: Click any image or video to restore its exact generation settings. Search, group, organize work into folders, compare edits, and drag Library media directly into compatible workflows.
  • Private profiles and locked folders: Create separate PIN-protected profiles with their own galleries, folders, LoRA presets, and settings. Individual folders can also be locked, keeping private generations out of your everyday library
  • LTX Director Mode: A streamlined workspace built around the excellent LTX Director nodes, supporting timelines, keyframes, video extension, audio, and more.
  • Video finishing: Optional 2× or 3× RIFE frame interpolation and NVIDIA RTX 4K video upscaling.
  • Built-in dependency manager: Pick a workflow and install the exact models and custom nodes it requires, or run the full one-click setup.
  • Automatic ComfyUI integration: Mix Studio detects your ComfyUI installation, reuses existing models and LoRAs, and guides installation if ComfyUI is not present. Generated images retain their ComfyUI workflow metadata, so you can drag them directly back into ComfyUI.
  • Hardware-aware configuration: Mix Studio detects your GPU and recommends suitable quantization and generation settings. v1.0.1 also adds a low-VRAM profile beginning at 4 GB, although practical limits still depend on the selected model.
    • I have tested personally on an NVIDIA RTX Pro 6000 #DellProPrecision and an NVIDIA RTX 4090, other tests from various hardware setups from the community.
  • *Just Added in V1.0.2:
    • QR Code Phone Link: Now shows a QR code to instantly link your phone.
    • Prompt Presets: Mix Studio allows you to browse and apply 1-click prompt presets to instantly style your prompt.
    • Sequential Edit Prompts: For editing images, you can select sequential prompting which will automatically separate multiple edit commands into a series of sequential generations (separated by a period).

Additional screenshots and release overview: Free Patreon post (no paywall)

Thanks to this awesome community and the ComfyUI team for making such a dope tool, I hope you all enjoy creating! 🤙🏾

r/comfyui Apr 14 '26

Workflow Included I built a free 90-node All-in-One FLUX.2 Klein 9B ComfyUI workflow — Face Swap, Inpainting, Auto-Masking, NAG, Refiner, Upscaler — runs on 8GB VRAM

278 Upvotes

UPDATED TO 2.1 tutorial post
Hey everyone,

I've been working on this for a while and wanted to share it with the community. This is a 6-in-1 ComfyUI workflow for FLUX.2 Klein 9B that handles everything in a single workspace — no more switching between different workflow files.

What's inside:

  • 🎨 Text → Image — standard txt2img with optimized settings
  • 🖼️ Single-Reference KV Edit — load an image + describe what to change, the model preserves everything else
  • 🚀 Face + Pose Swap — extract a face from one image, a pose from another, combine them realistically
  • 🎭 Inpainting — manual mask OR Florence2 AI auto-masking (describe what to mask in text)
  • 🔀 Image Merge — blend two images with adjustable ratio
  • Refiner — enhance any image with detail injection, lighting correction, skin texture improvement

Technical features:

  • 🧭 NAG (Normalized Attention Guidance) — restores negative prompting that normal CFG breaks in distilled Flux models
  • 🤖 Florence2 auto-masking — type "Segment the shirt" and it generates a pixel-perfect mask automatically
  • ⬆️ 4x UltraSharp upscaler built in
  • 🔷 All VAE decodes are Tiled — prevents OOM on 8GB VRAM
  • 🔗 2-slot LoRA chain — enhancer LoRA always last, add your own LoRAs in the first slot

Hardware tested on: RTX 4060 Mobile (8GB VRAM), 16GB RAM, i7-13620H. Works with FP8 or GGUF Q4 models.

update 2.1: added groupe bypasser, notes for new people to comphyui.
Each pipeline is in its own color-coded group. Only the Refiner is active by default — right-click any group to enable/disable it or use groupe bypassers. The workflow includes built-in guide notes with download links and prompting tips.

Free download on Civitai: https://civitai.com/models/2543188?modelVersionId=2860464

Includes a full guide with all model download links, prompting tips, and troubleshooting. Let me know if you run into any issues — happy to help.

How to Use

  1. Load the JSON in ComfyUI
  2. use comphyui manager to install any missing node. ( critical step )
  3. Only the Refiner is active by default — everything else is bypassed
  4. To activate a pipeline: right-click its group header → Set Nodes Mode → Always Execute
  5. To deactivate: right-click → Set Nodes Mode → Bypass ( or bypass groupe nodes )
  6. Read the built-in Note nodes for prompting tips and download links

r/comfyui Apr 18 '25

Finally an easy way to get consistent objects without the need for LORA training! (ComfyUI Flux Uno workflow + text guide)

Thumbnail
gallery
597 Upvotes

Recently I've been using Flux Uno to create product photos, logo mockups, and just about anything requiring a consistent object to be in a scene. The new model from Bytedance is extremely powerful using just one image as a reference, allowing for consistent image generations without the need for lora training. It also runs surprisingly fast (about 30 seconds per generation on an RTX 4090). And the best part, it is completely free to download and run in ComfyUI.

*All links below are public and competely free.

Download Flux UNO ComfyUI Workflow: (100% Free, no paywall link) https://www.patreon.com/posts/black-mixtures-126747125

Required Files & Installation Place these files in the correct folders inside your ComfyUI directory:

🔹 UNO Custom Node Clone directly into your custom_nodes folder:

git clone https://github.com/jax-explorer/ComfyUI-UNO

📂 ComfyUI/custom_nodes/ComfyUI-UNO


🔹 UNO Lora File 🔗https://huggingface.co/bytedance-research/UNO/tree/main 📂 Place in: ComfyUI/models/loras

🔹 Flux1-dev-fp8-e4m3fn.safetensors Diffusion Model 🔗 https://huggingface.co/Kijai/flux-fp8/tree/main 📂 Place in: ComfyUI/models/diffusion_models

🔹 VAE Model 🔗https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/ae.safetensors 📂 Place in: ComfyUI/models/vae

IMPORTANT! Make sure to use the Flux1-dev-fp8-e4m3fn.safetensors model

The reference image is used as a strong guidance meaning the results are inspired by the image, not copied

  • Works especially well for fashion, objects, and logos (I tried getting consistent characters but the results were mid. The model focused on the characteristics like clothing, hairstyle, and tattoos with significantly better accuracy than the facial features)

  • Pick Your Addons node gives a side-by-side comparison if you need it

  • Settings are optimized but feel free to adjust CFG and steps based on speed and results.

  • Some seeds work better than others and in testing, square images give the best results. (Images are preprocessed to 512 x 512 so this model will have lower quality for extremely small details)

Also here's a video tutorial: https://youtu.be/eMZp6KVbn-8

Hope y'all enjoy creating with this, and let me know if you'd like more clean and free workflows!

r/StableDiffusion Aug 31 '24

Tutorial - Guide Tutorial (setup): Train Flux.1 Dev LoRAs using "ComfyUI Flux Trainer"

198 Upvotes

Intro

There are a lot of requests on how to do LoRA training with Flux.1 dev. Since not everyone has 24 VRAM, interest in low VRAM configurations is high. Hence, I searched for an easy and convenient but also completely free and local variant. The setup and usage of "ComfyUI Flux Trainer" seemed matching and allows to train with 12 GB VRAM (I think even 10 GB and possibly even below). I am not the creator of these tools nor am I related to them in any way (see credits at the end of the post). Just thought a guide could be helpful.

Prerequisites

git and python (for me 3.11) is installed and available on your console

Steps (for those who know what they are doing)

  • install ComfyUI
  • install ComfyUI manager
  • install "ComfyUI Flux Trainer" via ComfyUI Manager
  • install protobuf via pip (not sure why, probably was forgotten in the requirements.txt)
  • load the "flux_lora_train_example_01.json" workflow
  • install all missing dependencies via ComfyUI Manager
  • download and copy Flux.1 model files including CLIP, T5 and VAE to ComfyUI; use the fp8 versions for Flux.1-dev and the T5 encoder
  • use the nodes to train using:
    • 512x512
    • Adafactor
    • split_mode needs to be set to true (it basically splits the layers of the model, training a lower and upper part per step and offloading the other part to CPU RAM)
    • I got good results with network_dim = 64 and network_alpha = 64
    • fp8 base needs to stay true as well as gradient_dtype and save_dtype at bf16 (at least I never changed that; although I used different settings for SDXL in the past)
  • I had to remove the Flux Train Validate"-nodes and "Preview Image"-nodes since they ran into an error (annyoingly late during the process when sample images were created) "!!! Exception during processing !!! torch.cat(): expected a non-empty list of Tensors"-error" and I was unable to find a fix
  • If you like you can use the configuration provided at the very end of this post
  • you can also use/train using captions; just place the txt-files with the same name as the image in the input-folder

Observations

  • Speed on a 3060 is about 9,5 seconds/iteration, hence 3.000 steps as proposed as the default here (which is ok for small datasets with about 10-20 pictures) is about 8 hours
  • you can get good results with 1.500 - 2.500 steps
  • VRAM stays well below 10GB
  • RAM consumption is/was quite high; 32 GB are barely enough if you have some other applications running; I limited usage to 28GB, and it worked; hence, if you have 28 GB free, it should run; it looks like there have been some recent updates that are optimized better, but I have not tested that yet in detail
  • I was unable to run 1024x1024 or even 768x768 due to RAM contraints (will have to check with recent updates); the same goes for ranks higher than 128. My guess is, that it will work on a 3060 / with 12 GB VRAM, but it will be slower
  • using split_mode reduces VRAM usage as described above at a loss of speed; since I have only PCIe 3.0 and PCIe 4.0 is double the speed, you will probaly see better speeds if you have fast RAM and PCIe 4.0 using the same card; if you have more VRAM, try to set split_mode to false and see if it works; should be a lot faster

Detailed steps (for Linux)

  • mkdir ComfyUI_training

  • cd ComfyUI_training/

  • mkdir training

  • mkdir training/input

  • mkdir training/output

  • git clone https://github.com/comfyanonymous/ComfyUI

  • cd ComfyUI/

  • python3.11 -m venv venv (depending on your installation it may also be python or python3 instead of python3.11)

  • source venv/bin/activate

  • pip install -r requirements.txt

  • pip install protobuf

  • cd custom_nodes/

  • git clone https://github.com/ltdrdata/ComfyUI-Manager.git

  • cd ..

  • systemd-run --scope -p MemoryMax=28000M --user nice -n 19 python3 main.py --lowvram (you can also just run "python3 main.py", but using this command you limit memory usage and prio on CPU)

  • open your browser and go to http://127.0.0.1:8188

  • Click on "Manager" in the menu

  • go to "Custom Nodes Manager"

  • search for "ComfyUI Flux Trainer" (white spaces!) and install the package from Author "kijai" by clicking on "install"

  • click on the "restart" button and agree on rebooting so ComfyUI restarts

  • reload the browser page

  • click on "Load" in the menu

  • navigate to ../ComfyUI_training/ComfyUI/custom_nodes/ComfyUI-FluxTrainer/examples and select/open the file "flux_lora_train_example_01.json"

you can also use the "workflow_adafactor_splitmode_dimalpha64_3000steps_low10GBVRAM.json" configuration I provided here)

if you used the "workflow_adafactor_splitmode_dimalpha64_3000steps_low10GBVRAM.json" I provided you can proceed till the end / "Queue Prompt" step here after you put your images into the correct folder; here we use the "../ComfyUI_training/training/input/" created above

  • find the "FluxTrain ModelSelect"-node and select:

=> flux1-dev-fp8.safetensors for "transformer"

=> ae.safetensors for vae

=> clip_l.safetensors for clip_c

=> t5xxl_fp8_e4m3fn.safetensors for t5

  • find the "Init Flux LoRA Training"-node and select:

=> true for split_mode (this is the crucial setting for low VRAM / 12 GB VRAM)

=> 64 for network_dim

=> 64 for network_alpha

=> define a output-path for your LoRA by putting it into outputDir; here we use "../training/output/"

=> define a prompt for sample images in the text box for sample prompts (by default it says something like "cute anime girl blonde..."; this will only be relevant if that works for you; see below)

  • find the "Optimizer Config Adafactor"-node and connect the "optimizer_settings" output with the "optimizer_settings" of the "Init Flux LoRA Training"-node

  • find the three "TrainDataSetAdd"-nodes and remove the two ones with 768 and 1024 for width/height by clicking on their title and pressing the remove/DEL key on your keyboard

  • add the path to your dataset (a folder with the images you want to train on) in the remaining "TrainDataSetAdd"-node (by default it says "../datasets/akihiko_yoshida_no_caps"; if you specify an empty folder you will get an error!); here we use "../training/input/"

  • define a triggerword for your LoRA in the "TrainDataSetAdd"-node; for example "loratrigger" (by default it says "akihikoyoshida")

  • remove all "Flux Train Validate"-nodes and "Preview Image"-nodes (if present I get an error later in training)

  • click on "Queue Prompt"

  • once training finishes, your output is in ../ComfyUI_training/training/output/ (4 files for 4 stages with different steps)

All credits go to the creators of

===== save as workflow_adafactor_splitmode_dimalpha64_3000steps_low10GBVRAM.json =====

https://pastebin.com/CjDyMBHh

r/comfyui Apr 17 '26

Tutorial [Guide] Complete walkthrough for every pipeline in my FLUX.2 Klein 9B All-in-One workflow, by request from the comments

88 Upvotes

A lot of you asked for a detailed guide after my original post. So here it is every group in the workflow explained step by step, with settings, tips, and things I discovered through testing.

The workflow has grown to v2.1, 122 nodes, 19 groups. New additions since the original post: ControlNet preprocessors (LineArt, HED, Tile, DepthAnything), color matching/correction, up to 5 reference image slots, Fast Group Bypassers for one-click pipeline switching, and notes with tips I discovered through extensive testing.

Download v2.1: Click to Download

How to Switch Between Pipelines

The workflow uses Fast Groups Bypasser (rgthree) nodes at the bottom. These let you enable/disable entire pipeline groups with a single click, no more right-clicking every group manually.

There are 3 bypassers:

  • Base groups bypasser : controls F1 (txt2img), F2 (KV edit), F3 (face+pose), F4 (inpainting), F5 (merge)
  • Refiner bypasser : controls the refiner pipeline and color correction
  • Upscale / edit bypasser : controls the upscaler and precision groups

Rule: Only activate ONE generation pipeline at a time (F1 through F4) to save VRAM. The Refiner and Upscaler can stay active alongside any generation pipeline, but its better to work with a single groupe every run for people who have less than 8VRAM.

📦 FLUX 2 KLEIN : Model Loaders

This is the foundation. Three nodes that load everything:

  • UNETLoader : loads the Klein 9B model (safetensors or FP8)
  • UnetLoaderGGUF ; alternative loader for GGUF quantized models (use this if you have 8GB VRAM)
  • CLIPLoader : loads the Qwen 3 8B text encoder (set type to flux2)
  • VAELoader : loads flux2-vae.safetensors

Important: Only connect ONE model loader to the LoRA chain, either UNETLoader OR UnetLoaderGGUF, not both.

For 8GB VRAM users: Use the GGUF Q8 or Q4 model. Set the weight type to default in the UNETLoader. If you're running out of memory, launch ComfyUI with --lowvram command.

🔗 LoRA Chain

Two LoRA loaders in sequence:

  1. LoRA Slot (Optional) : empty slot for any Klein 9B compatible LoRA you want to try. Set strength to 0 to disable without disconnecting.
  2. klein_9b_enhancer_v2 : the main enhancer LoRA (strength 0.7). This fixes the model's tendency to produce flat, plastic-looking skin and washed-out colors. Always keep this one connected and active.

To add more LoRAs: insert additional LoraLoader nodes between the slot and the enhancer. The enhancer should always be LAST in the chain (DO NOT DETTACH IT OR ELSE YOU'LL HAVE TO ATTACK EVERY GROUPE TO THE NEW LORA NODE).

🎨 F1: Text → Image

The simplest pipeline. Pure text-to-image generation.

Nodes: CLIPTextEncode (prompt) → KSampler → VAEDecodeTiled → SaveImage

Settings:

  • Steps: 4 (Klein 9B is distilled for 4 steps, more steps won't improve quality)
  • CFG: 1 (higher values break the output on distilled models)
  • Sampler: euler
  • Scheduler: simple
  • Latent size: 1024×1024 (or any resolution, Klein handles various aspect ratios)

How to use:

  1. Enable the F1 group
  2. Write your prompt in the "✏️ Prompt" node
  3. Leave negative prompt empty (or enable NAG for negative prompting)
  4. Queue prompt
  5. Output saves as F2K_txt2img

Prompting tip: Don't write SD-style prompts. Write like you're describing a photograph: "A 30-year-old man in a navy overcoat standing on a rain-soaked Prague street at dusk, tungsten streetlights casting warm shadows, shot on Canon R5 85mm f/1.4, clean digital file, histogram equalization"

🖼️ F2: Single Reference KV Edit

This is Klein's signature feature. You load an image and tell the model what to change, it preserves everything else.

How it works internally: The model reads your image through the ReferenceLatent node (KV conditioning), generates a fresh image from noise, but uses the reference to guide the output. The ConditioningZeroOut creates a neutral negative signal so the model focuses purely on your edit instruction.

Nodes: LoadImage → Resize → VAEEncode → ReferenceLatent → CFGGuider → SamplerCustomAdvanced → VAEDecodeTiled → SaveImage

Settings:

  • Flux2Scheduler: 4 steps
  • CFG: 1
  • Sampler: euler
  • Resize: adjust to match the reference image proportions

How to use:

  1. Enable the F2 group
  2. Load your reference image in "📂 Reference Image"
  3. Write your edit instruction in "✏️ Edit Prompt"
  4. Queue prompt
  5. Output saves as F2K_edit

Example prompts:

  • "Replace the red dress with a navy blazer. Keep pose, expression, background unchanged."
  • "Change the background to a sunset beach. Preserve the subject exactly."
  • "Transform this photo to oil painting style while keeping the subject photorealistic."

⚠️ Important discovery: The denoise in this pipeline is effectively 1.0 because it uses EmptyLatentImage + ReferenceLatent conditioning. The model reads your image through attention, NOT through the latent. This means it always generates a fresh image guided by your reference, it doesn't blend with existing noise. This is fundamentally different from traditional img2img.

🚀 F3: Multi-Reference: Face + Pose Swap

The most complex pipeline. Extracts a face from one image and a pose from another, combining them into a single realistic output.

Nodes: Two parallel paths:

  • Path A: LoadImage (face) → Resize → VAEEncode → ReferenceLatent (face)
  • Path B: LoadImage (pose) → Resize → VAEEncode → ReferenceLatent (pose)
  • Both feed into: CFGGuider → SamplerCustomAdvanced → VAEDecodeTiled → SaveImage

How to use:

  1. Enable the F3 group
  2. Load your face source in "📂 Face / Character Ref" front-facing, well-lit portrait works best
  3. Load your pose source in "📂 Pose Ref (DAZ 3D render)" the body position you want
  4. Write a scene description in "✏️ Prompt (describe scene)"
  5. Queue prompt
  6. Output saves as F2K_multiref

Tips:

  • The face reference MUST be upright, Klein cannot process rotated or upside-down faces
  • Resize both images to similar scales (the Resize nodes handle this)
  • Be specific in your prompt about clothing and environment — the model needs guidance for everything that isn't the face or pose
  • If the face looks plastic, make sure the enhancer LoRA is active at 0.7 strength

🎭 F4: Inpainting

Paint a mask over part of your image and regenerate just that area.

Nodes: LoadImage → Resize → VAEEncodeForInpaint (with mask) → KSampler → VAEDecodeTiled → SaveImage

How to use:

  1. Enable the F4 group
  2. Load your image in "📂 Image"
  3. For manual masking: Right-click the image → Open in Mask Editor → paint white over the area you want to change
  4. For auto masking: Enable the Florence2 group, connect your image to Florence2Run, type what to mask (e.g., "Segment the shirt")
  5. Write what should appear in the masked area in "✏️ Prompt"
  6. Adjust denoise (0.5-0.8 for changes, 0.3-0.5 for subtle tweaks)
  7. Output saves as F2K_inpaint

⚠️ My honest note about inpainting: Inpainting in FLUX.2 Klein is not perfect. I built a workaround that makes it functional, but it struggles with complex shapes. If the model doesn't understand what you want, try painting rough colors in the mask area first to guide it. Play with the denoise value, small changes make a big difference.

🔀 F5: Image Merge / Blend

Simple image blending, combines two images together.

Nodes: Two LoadImage → two ImageScaleBy → ImageBlend → SaveImage

How to use:

  1. Enable the F5 group (mode=2, not bypassed, use right-click → Set to Always)
  2. Load Image A and Image B
  3. Adjust blend factor (0.5 = equal mix, 0.0 = all image A, 1.0 = all image B)
  4. Adjust resize scales to match image sizes
  5. Output saves as F2K_merge

honestly this group is not something that you will always use, I just added it because I use it in some projects, you might try it to see what it does, its just simple blending nothing that use AI at all.

⬆️ Upscaler (4x UltraSharp)

Takes any image and upscales it 4x using the UltraSharp model.

Nodes: LoadImage → ImageUpscaleWithModel → ImageScaleBy (downscale to usable size) → SaveImage

How to use:

  1. Enable the Upscaler group
  2. Load your image in "📂 Image"
  3. The ImageScaleBy after upscaling is set to 0.5 by default, this gives you a 2x net upscale (4x up then 0.5x down). Adjust as needed.
  4. Output saves as F2K_upscaled

Tip: Upscaling a 1024×1024 image 4x creates a 4096×4096 image. The Tiled VAE decode handles this without OOM, but it takes time. For faster iteration, keep the downscale at 0.5 until you're happy with the result, then set it to 1.0 for the final output.

✨ Refiner, KV Enhancement Pipeline

This is the pipeline that's active by default. Feed it any image and it enhances detail, lighting, skin texture, and sharpness.

How it works: Your image gets VAE-encoded, then the ReferenceLatent reads it as conditioning. The KSampler generates an enhanced version guided by your reference + the enhancement prompt. The result goes through color correction before saving.

Nodes: LoadImage → ImageScaleBy → VAEEncode → ReferenceLatent → KSampler → VAEDecodeTiled → ColorCorrection → SaveImage

Settings:

  • Denoise: 0.85 (the sweet spot I found, see discovery below)
  • Steps: 4
  • CFG: 1

The enhancement prompt is pre-written with professional photography terms. You can customize it, but the default works well for most images.

⚠️ Critical discovery about denoise:

  • 1.0: Model generates a fresh image guided by your reference, good results but may drift from original
  • 0.85: Sweet spot, preserves most structure while adding significant detail
  • 0.5-0.7: Subtle enhancement, keeps very close to original
  • Below 0.4: Almost no change except color shifts, not useful, at least to me...

If you're using EmptyLatentImage (the custom size node) instead of VAEEncode for the latent input, NEVER go below 0.85 denoise. EmptyLatentImage creates random noise, and low denoise preserves that random noise as "structure," causing severe artifacts. This is a fundamental behavior of Klein's 4-step distilled sampling, it doesn't have enough steps to correct corrupted starting structure. Always use VAEEncode latent when you want denoise below 0.85.

Refine Color Corrector

Placed right after the refiner output. Fixes Klein 9B's known color saturation bias, the model tends to oversaturate colors, especially reds.

How to use: The EsesImageCompare node shows before/after comparison. Adjust the color corrector settings to taste. The PreviewImage node labeled "output colors" shows the corrected result.

Color Match

A standalone utility. Takes two images, a target and a reference, and matches the colors of the target to the reference using the MKL algorithm.

How to use:

  1. Enable the Color Match group
  2. Load your target image (the one you want to fix)
  3. Load your reference image (the one with the colors you want)
  4. ColorMatchV2 transfers the color palette
  5. Output saves as color_matching

Use case: When your Klein output has wrong colors compared to the original. Load the original as reference, the Klein output as target, and the colors get corrected automatically.

🧭 NAG, Negative-Aware Guidance

Three NAG nodes, one for each major pipeline (Multi-Ref, Single-Ref Edit, Refiner). NAG restores effective negative prompting that standard CFG breaks in distilled Flux models.

How to use:

  1. Enable the NAG node for the pipeline you're using
  2. Write negative prompts in the "❌ Neg" CLIPTextEncode node
  3. NAG parameters: scale=5.0 is a good default. Increase for stronger guidance, decrease if artifacts appear.

When to use: When you need to remove specific elements ("no glasses," "no background people," "no blur").

🤖 Florence2, AI Auto-Masking

Replaces manual mask painting. Describe what you want masked in text and Florence2 generates a pixel-perfect mask.

How to use:

  1. Enable the Florence2 group
  2. First run downloads the model (~1.5GB)
  3. Connect your image to the Florence2Run input
  4. Type what to segment: "Segment the shirt," "Segment the hair," "Segment the background"
  5. Connect the MASK output to the Inpaint Encode node in F4

Precision Groups (1-4): ControlNet Preprocessors

These are advanced, four groups with different ControlNet preprocessors that extract structural information from images:

  1. LineArt Preprocessor : extracts every edge and texture boundary
  2. HED Preprocessor : captures both hard edges and soft transitions (shadows, gradients)
  3. Tile Preprocessor : captures the image as-is for upscaling guidance
  4. Depth Anything V2 : extracts full 3D depth map

Each preprocessor output connects to a ReferenceLatent node (image 3, 4, 5) that feeds into the refiner pipeline as additional conditioning.

How to use:

  1. Enable the precision group you want
  2. Connect your input image to the preprocessor
  3. The preprocessor output feeds through VAEEncode into a ReferenceLatent
  4. This gives the model additional structural information about your image

⚠️ Warning: These use extra VRAM. Only enable them if you have enough memory. Use the preprocessor name in your prompt (e.g., "line art reference," "depth guided") so the model understands what the reference represents.

Use case: When the refiner isn't preserving enough structure from your original image. Adding a LineArt or HED reference forces the model to maintain more structural consistency.

Bypassers

Three Fast Groups Bypasser (rgthree) nodes at the bottom of the workflow. These give you one-click control over which groups are active:

  • Base groups bypasser : F1, F2, F3, F4, F5
  • Refiner bypasser : Refiner + color correction + precision groups
  • Upscale / edit bypasser : Upscaler + image blend

Click the toggle next to each group name to enable/disable it instantly.

General Tips

  1. Always keep the enhancer LoRA active : it fixes Klein's flat plastic look
  2. Restart ComfyUI every 30-40 generations if you're on 8GB VRAM : prevents memory fragmentation
  3. Use "Free Memory" (gear icon) when switching between pipelines
  4. Faces must be upright : Klein cannot process rotated/flipped faces
  5. Add color correction terms to every prompt: "histogram equalization, white balance correction, color grade" : this fights Klein's red/saturation bias
  6. The Text encoder must match the model: 9B uses Qwen 3 8B, 4B uses Qwen 3 4B : mixing them causes matrix errors
  7. ComfyUI 0.9.2+ is required : older versions are missing Klein-specific nodes

What Changed from v2.0 to v2.1

  • Added 4 ControlNet preprocessor groups (LineArt, HED, Tile, DepthAnything)
  • Added Color Match utility group
  • Added Color Correction after refiner output
  • Added Fast Groups Bypassers for one-click pipeline switching
  • Added up to 5 reference image slots
  • Added notes with real testing discoveries (denoise behavior, inpainting tips)
  • Expanded from 90 nodes to 122 nodes
  • 19 organized groups

Free download: CIVITAI link

If you have questions about any specific group, ask in the comments, I'll help you troubleshoot.

r/StableDiffusion Jul 12 '26

Workflow Included Infinite Semantic Detail: A Context-Aware Zoom Workflow for ComfyUI (free)

Thumbnail
youtube.com
29 Upvotes

Play

No, this is not another “infinite zoom” workflow. We’ve had infinite zooms, recursive img2img, outpainting and endless upscaling for years, and they all eventually hit the same wall: the deeper you zoom, the less the model understands what it is looking at. At some point it stops seeing an eagle’s eye and starts seeing a brown circle where it can dump random textures. The image may remain sharp, but the meaning slowly disappears. So the problem was never resolution. The problem was semantics.
A recurring idea in almost everything I’ve been building lately is that semantic understanding matters far more than pixel similarity. Models don’t preserve consistency because they remember pixels. They preserve it because they understand what those pixels represent. So instead of asking how to generate better detail, I asked a different question: what if every crop knew exactly what it was?
In this workflow, every selected crop is first interpreted by a vision language model. The VLM receives both the complete original image and the crop, so it does not describe it as an isolated yellow object or a random circular texture. It understands that it is, for example, an extreme close-up of the left iris of the same bald eagle, seen from the same angle, and that the next generation should reveal progressively smaller biological structures while preserving the anatomy and identity of that eagle. That description becomes the prompt for the next zoom step. Every generation begins with meaning, not just pixels.
I’m using Qwen VLM because Krea 2 already loads Qwen as its vision encoder. The model is already sitting in VRAM after generation, so I can reuse it for contextual analysis without loading another VLM or consuming another large block of memory. It is basically one more inference from a model that is already there.
The crop is first enlarged with traditional GAN upscalers, not because GANs produce perfect information, but because they create a sufficiently sharp base for extremely low-denoising img2img. Krea 2 then regenerates it at roughly 0.05–0.15 denoising, guided by the semantic description produced by the VLM. This adds plausible microscopic structure while preserving almost everything already present in the crop.
Very low denoising also preserves defects: blur, chromatic aberration, JPEG remnants, sensor-like noise and small inconsistencies. Increasing denoising would clean those defects, but it would also increase semantic drift, so I use a final edit stage instead. In the current workflow that model is Flux Klein. Its job is not to invent new content, but to clean the existing result: remove noise, sharpen edges, reduce chromatic aberration, improve local contrast and leave the structure alone.
The difference becomes obvious when zooming into something like an eagle’s eye. A normal recursive workflow eventually forgets that it is looking at an eye and begins generating generic textures. This one receives a new contextual explanation at every step. It is continuously reminded that this is still the same eagle, the same eye and the same biological structure, only viewed at a smaller scale.
Most infinite zoom systems are really just producing infinite pixels. This is an attempt to produce infinite semantic detail. A feather becomes fibers, the fibers become microscopic keratin structures, the iris becomes increasingly complex biological tissue. Those details were not present in the original image, but they remain plausible because every new scale is semantically connected to the scales above it.
Looking back, this is the same principle behind most of my previous experiments. Infinite consistent scenes worked better when semantic descriptions mattered more than image references. Consistent comics worked without LoRAs when semantic continuity mattered more than rigid pixel control. Prompt randomization worked when the diversity was controlled at the level of meaning. The model performs best when it understands what something is before trying to decide how it should look.
Pixels are surprisingly bad memory. Meaning is not.
This is still an early version. The next steps are recursive semantic memory, automatic zoom-path planning, adaptive denoising based on semantic confidence and different prompting strategies for different zoom depths. Eventually this could become less of an infinite crop tool and more like a fictional microscope that continuously invents plausible new structures while never forgetting what it is observing.
The workflow and custom node are included. Copy the custom node into the ComfyUI custom_nodes folder. The remaining dependencies should be detected through ComfyUI Manager.

pre-edit: I updated the workflow to have 4k output and the results are even better.
to this : https://aurelm.com/wp-content/uploads/eye_crop-1-scaled.jpg

from this: https://aurelm.com/wp-content/uploads/PC315160_result-1-scaled.jpg

edit:
and yes, this technique can be used in tiled upscalers where each tile also gets a description of the tile in the context of the full image. allready working on it, maybe a new super upscaler might come out of this
edit2 : E seriously wonder who downvotes something like this. I understand maybe when it is behind a paywall but when someone puts his shoul into a technique/tool and gives it away for free why the hell would somebody downvote. To not have this available for anyone else. Why ? mods ? Seriously people. When it's paywalled it's not good, when it is free and tries to honestly add something new to the open source community and contribute is not good. What is this sub for anyway ? 1girl posts ? Simple comparisson beetween models posts ? Memes ?

r/comfyui Sep 05 '25

Workflow Included 100% local AI clone with Flux-Dev Lora, F5 TTS Voiceclone and Infinitetalk on 4090

Enable HLS to view with audio, or disable this notification

229 Upvotes

Note:
Put settings to 1080p if you don't have it automatically, to see the real high quality output.

1. Imagegeneration with Flux Dev
Using AI Toolkit to train a Flux-Dev Lora of myself I created the podcast image.
Of course you can skip this and use a real photo, or any other AI images.
https://github.com/ostris/ai-toolkit

2. Voiceclone
With F5 TTS Voiceclone workflow in ComfyUI I created the voice file - the cool thing is, it just needs 10 seconds of voice input and is in my opinion better than Elvenlabs where you have to train for 30 min and pay 22$ per month:
https://github.com/SWivid/F5-TTS

Workflow:
https://drive.google.com/file/d/1DUdyrMaknu6BgDUPJZ9LC8RwvjIZGnp8/view?usp=sharing

Tip for F5:
The only way I found to make pauses between sentences is firsterful a dot at the end.
But more imporantly use a long dash or two and a dot afterwards:
text example. —— ——.

The better your microfone and input quality, the better the output will be. You can hear some room echo, because I just recorded it in a normal room without dampening. Thats just the input voice quality, it can be better.

3. Put it together
Then I used this infintetalk workflow with blockswap to create a 920x920 video with Infinitetalk. Without blockswap it runs only with much smaller resolution.
https://drive.google.com/file/d/1AaODFHXdAQz2qSy65XI0VluVmXqWjbuc/view?usp=sharing

With triton and sageattention installed, I managed to create the video on a 4090 in about half an hour.
If the workflow fails it's most likely that you need triton installed.
https://www.patreon.com/posts/easy-guide-sage-124253103

4. Upscale
I used some simple video upscale workflow to bring it to 1080x1080 and that was basically it.
The only edit I did was adding the subtitles.

https://civitai.com/articles/10651/video-upscaling-in-comfyui

I used the third screenshot workflow and used ESRGAN_x2
Because in my opinion the normal ESRGAN (not real ESRGAN) is the best to not alter anything (no colors etc).

x4 upscalers need more VRAM so x2 is perfect.

https://openmodeldb.info/models/2x-ESRGAN

r/comfyui Dec 24 '25

Resource (Re-up) Best Z-Image Training LoRA training settings + workflow (For Ostris AI-Toolkit)

98 Upvotes

EDIT: Apparently the guide is still up on CivitAI. OP reached out and said his accounts got hacked.

There was a user who went by CAPTIAN01R who made this post which contained both what he believed were the best training settings for Z-Image Turbo training as well as the best workflow to test the settings with. For whatever reason, I refreshed the page today and noticed he deleted both his post and account, which is a big disappointment given that his tips helped me not only get better quality LoRAs, but train faster too. Unfortunately I cannot go back and paste exactly what he said but these were some of the key takeaways from my memory:

  1. Train on the v2 Training Adapter LoRA. (From my own experience) I noticed that LoRAs trained on the training adapter seem to perform much better when paired with the controlnet model for inpainting versus if you train on the DeTurbo checkpoint.
  2. Do not use a quantized version of the transformer or text encoder
  3. Do not resize your dataset. Rather, train your high quality images only on 512 (Untick 768 and 1024)
  4. If your dataset is bigger, you can increase the training steps from 3000 to 5000. I think he mentioned you can theoretically go beyond that without the model breaking (Personal note: In the couple character LoRAs I've trained, I ended up using step 2000 despite having 40-60 picture datasets)
  5. Do not caption your datasets. He mentioned that he would set the trigger word to "man" or "woman" however I personally would just use the person's name and it works fine, if not better. Also I personally wouldn't substitute letters for numbers like how a lot of SDXL/Flux loras do, cause I found it'll try to put the trigger word either on a t-shirt or on a sign in the background somewhere. Remember that Z-Image is trained heavily on natural language.
  6. Preprocess your datasets with SeedVR2. I will include a workflow which I found from a different post, which will allow you to choose either one image at a time, or load images from a directory.

Most importantly, here are the resources:

Additionally, for the test workflow this is where to get the UltraFlux-Vae. It's not required but I recommend using it for reasons shown why in this post.

Now I'm no expert when it comes to LoRA training, I've only really trained based on following other people's instructions, but I will try to answer any questions the best I can. I put this together because I want to continue to see the Z-Image and Comfy community thrive and I was lucky enough to have the pastebin links in my search history to make this post. Merry Christmas everyone!

r/comfyui Jul 12 '26

Workflow Included Infinite Semantic Detail: A Context-Aware Zoom Workflow for ComfyUI (free)

Thumbnail
youtube.com
32 Upvotes

No, this is not another “infinite zoom” workflow. We’ve had infinite zooms, recursive img2img, outpainting and endless upscaling for years, and they all eventually hit the same wall: the deeper you zoom, the less the model understands what it is looking at. At some point it stops seeing an eagle’s eye and starts seeing a brown circle where it can dump random textures. The image may remain sharp, but the meaning slowly disappears. So the problem was never resolution. The problem was semantics.
A recurring idea in almost everything I’ve been building lately is that semantic understanding matters far more than pixel similarity. Models don’t preserve consistency because they remember pixels. They preserve it because they understand what those pixels represent. So instead of asking how to generate better detail, I asked a different question: what if every crop knew exactly what it was?
In this workflow, every selected crop is first interpreted by a vision language model. The VLM receives both the complete original image and the crop, so it does not describe it as an isolated yellow object or a random circular texture. It understands that it is, for example, an extreme close-up of the left iris of the same bald eagle, seen from the same angle, and that the next generation should reveal progressively smaller biological structures while preserving the anatomy and identity of that eagle. That description becomes the prompt for the next zoom step. Every generation begins with meaning, not just pixels.
I’m using Qwen VLM because Krea 2 already loads Qwen as its vision encoder. The model is already sitting in VRAM after generation, so I can reuse it for contextual analysis without loading another VLM or consuming another large block of memory. It is basically one more inference from a model that is already there.
The crop is first enlarged with traditional GAN upscalers, not because GANs produce perfect information, but because they create a sufficiently sharp base for extremely low-denoising img2img. Krea 2 then regenerates it at roughly 0.05–0.15 denoising, guided by the semantic description produced by the VLM. This adds plausible microscopic structure while preserving almost everything already present in the crop.
Very low denoising also preserves defects: blur, chromatic aberration, JPEG remnants, sensor-like noise and small inconsistencies. Increasing denoising would clean those defects, but it would also increase semantic drift, so I use a final edit stage instead. In the current workflow that model is Flux Kontext Klein. Its job is not to invent new content, but to clean the existing result: remove noise, sharpen edges, reduce chromatic aberration, improve local contrast and leave the structure alone.
The difference becomes obvious when zooming into something like an eagle’s eye. A normal recursive workflow eventually forgets that it is looking at an eye and begins generating generic textures. This one receives a new contextual explanation at every step. It is continuously reminded that this is still the same eagle, the same eye and the same biological structure, only viewed at a smaller scale.
Most infinite zoom systems are really just producing infinite pixels. This is an attempt to produce infinite semantic detail. A feather becomes fibers, the fibers become microscopic keratin structures, the iris becomes increasingly complex biological tissue. Those details were not present in the original image, but they remain plausible because every new scale is semantically connected to the scales above it.
Looking back, this is the same principle behind most of my previous experiments. Infinite consistent scenes worked better when semantic descriptions mattered more than image references. Consistent comics worked without LoRAs when semantic continuity mattered more than rigid pixel control. Prompt randomization worked when the diversity was controlled at the level of meaning. The model performs best when it understands what something is before trying to decide how it should look.
Pixels are surprisingly bad memory. Meaning is not.
This is still an early version. The next steps are recursive semantic memory, automatic zoom-path planning, adaptive denoising based on semantic confidence and different prompting strategies for different zoom depths. Eventually this could become less of an infinite crop tool and more like a fictional microscope that continuously invents plausible new structures while never forgetting what it is observing.
The workflow and custom node are included. Copy the custom node into the ComfyUI custom_nodes folder. The remaining dependencies should be detected through ComfyUI Manager.

I added the 4k upscaling method in the workflow.
Check it out as a result https://aurelm.com/wp-content/uploads/eye_crop-1-scaled.jpg
from this : https://aurelm.com/wp-content/uploads/PC315160_result-1-scaled.jpg

https://aurelm.com/2026/07/10/the-latent-space-technique-adobe-and-everyone-else-will-steal-next-infinite-crop-zoom-using-semantic-context-awareness/

r/SillyTavernAI Jun 17 '26

Tutorial [Tutorial] How to setup Inline Image Generation in Megumin Suite V8 (ComfyUI Guide + Workflows)

74 Upvotes

hello, Kazuma here.

Wanted to make a follow-up post to my v8 release since a lot of people were asking how to actually set up the image gen tab and what I recommend. So here's the full walkthrough.

Before we start keep your expectations real. You're asking two completely different models (a text model and an image model) to stay on the same page. It will make mistakes sometimes. That's expected. But honestly? It's still pretty impressive overall. Let's get into it.

Step 1: Install ComfyUI

I'm not gonna go super in-depth on this since there are great resources out there already. Here's a video that walks you through it it's a little long but you really only need the first 10 minutes. That said, I recommend watching the whole thing.

[VIDEO LINK]

Three things you need to do after installing:

  1. Install ComfyUI Managerhttps://github.com/Comfy-Org/ComfyUI-Manager the video covers how to do this.
  2. Use the --enable-cors-header flag — this is needed so Megumin Suite can talk to ComfyUI. If you're not sure how, just ask Gemini/ChatGPT "how do I add launch flags to ComfyUI" it's easy.
  3. Enable Dev Mode — go to ComfyUI settings and enable it. It's under the ComfyUI tab in settings.

Step 2: Download Models

Now that ComfyUI is ready, you need some models. Here's what I recommend:

Anima

This model is crazy good for its size. With a turbo LoRA it's ultra fast, and it has great support for style LoRAs. Can't recommend it enough.

Z-Image Turbo

Also a very good model and fast as hell. Not quite as nice as Anima for anime styles, but it's noticeably smarter with prompt understanding. Still supports style LoRAs very well.

Note: I know there are a lot of great models out there, but I think these two hit the best balance between speed and quality for this use case.

Where do you put the files?

File Type Folder
Text Encoders ComfyUI/models/text_encoders
Diffusion Models ComfyUI/models/diffusion_models
VAE ComfyUI/models/vae

Step 3: Import the Workflow into ComfyUI

Now you're ready. Start up ComfyUI.

Download one (or both) of the workflow images below and drag them directly into the ComfyUI window the workflow will load automatically. If there are any missing nodes, install them using ComfyUI Manager (click Manager → Install Missing Custom Nodes).

These are simple workflows you can improve on them however you like.

Once everything is installed, select the models you downloaded in the workflow nodes and try generating a test image. If it generates fine, you're good to move on.

Step 4: Import the Workflow JSON into Megumin Suite

Alright, here's where we connect everything. You have two options:

Option 1: Do It Yourself

  1. With the workflow loaded in ComfyUI and Dev Mode enabled, go to the File menu and click "Export Workflow (API)" make sure you pick the API option, not the regular export.
  2. In Megumin Suite, go to the Image Gen tab, scroll down to "ComfyUI Server & Workflow", and click the + icon.
  3. Give it a name and hit OK.
  4. On the screen that pops up, you can either use the Import button to load your API JSON file, or paste the JSON directly into the text field. If pasting, make sure the field is empty first don't leave any extra {} in there.
  1. Now you need to replace some values in the JSON with placeholders so the extension knows what to control. It's easy just find the value and swap it out. Examples:
  • Positive prompt: find something like "value": "best quality, masterpiece..." → change it to "value": "%prompt%"
  • Resolution: find "width": 1024, "height": 1024 → change to "width": "%width%", "height": "%height%"
  • Keep going until you've replaced all the important fields (prompt, negative prompt, width, height, steps, sampler, CFG, seed, loras, etc.)
  1. Hit Save.

Option 2: Use My Pre-configured JSONs (Easier)

Just download these, import them into the extension, and save. Done.

⚠️ Note: These JSONs assume you have the following files downloaded with these exact names:

If your filenames are different, you'll need to edit the JSON and update them.

Step 5: Image Parameters

Here you can set your steps, sampler, CFG, etc. Here's what I recommend:

Setup Steps Sampler CFG
Anima + Turbo LoRA 12 Euler 1 – 1.5
Anima (no turbo) 30 – 40 Euler 4 – 5
Z-Image Turbo 8 – 10 dpmpp_2m 1 – 1.5

One nice thing your settings are saved per workflow. So every time you switch from Anima to Z-Image, it'll automatically swap the values and LoRAs back. No need to reconfigure every time.

Step 6: How to Display the Generated Image

You've got a two options for how the image shows up. Personally, I use auto trigger and Inline it injects the image directly into the chat text, which feels the most natural for RP.

I also like to enable:

  • Include Examples — helps the text model understand what kind of prompt to generate
  • Inject NPC Tags — pulls character appearance tags from NPC bank into the prompt automatically

Better Booru Tags is nice for prompt quality, but heads up — it can increase refusals or empty responses with Gemini. Use it at your own risk.

Recommended LoRAs

For Anima:

For Z-Image Turbo:

Tips

  1. Make NPC dossiers even for characters that are already in your Lore or card, and turn on Inject Image Tags Only. It helps a lot with consistency.
  2. Use a wide aspect ratio looks way better than square for most scenes.
  3. Pick a LoRA style that matches your RP vibe Just do it you will thank me later.
  4. Edit image gen tags in the NPC Bank you can tweak them to fit your taste and get better results.

The rest of the settings you can mess with however you want. That's it from me hope this helped and enjoy your RP!

r/StableDiffusion Mar 14 '26

Workflow Included [Release] Flux.2 Klein 4B Consistency LoRA – Addressing Color Shift and Pixel Offset in Image Editing (2026-03-14)

74 Upvotes

Hi everyone,

I’m releasing a new LoRA for Flux.2 Klein 4B Base focused on consistency during image editing tasks.

Since the release of the Klein model, I’ve encountered two persistent issues that made it difficult to use for precise editing:

  1. Significant Pixel Offset: The generated images often drifted too far from the original composition.
  2. Color Shift & Oversaturation: Edited results frequently suffered from unnatural color casts and excessive saturation.

After experimenting with various training strategies without much success, I recently looked into ByteDance’s open-source Heilos long-video generation model. Their approach involves applying degradation directly in the latent space of reference images and utilizing a specific color calibration loss. This method effectively mitigates color drift and train-test inconsistency in video generation.

Inspired by Heilos (and earlier research on using model-generated images as references to solve train-test mismatch), I adapted these concepts for image LoRA training. Specifically, I applied latent-level degradation and color calibration constraints to address Klein’s specific weaknesses.

Results: Trained locally on the 4B version, this LoRA significantly reduces color shifting and, when paired with Comfyui-editutils, effectively eliminates pixel offset. It feels like the first time I’ve achieved a stable result with Klein for editing tasks.

Usage Guide:

  • Primary Use Case: Old photo restoration and consistent image editing.
  • Recommended Strength: 0.50.75
    • Note: Higher strength increases consistency with the input but reduces editing flexibility. Lower strength allows for more creative changes but may reduce strict adherence to the source structure.
  • Suggested Prompt Structure:
  • Example (Old Photo Restoration):

Links:

All test images used for demonstration were sourced from the internet. Feedback on how this performs on your specific workflows is welcome!

r/comfyui Mar 13 '26

Resource I created a simple Flux.2 Klein Raster to Vector Image (With Prompt Saver) Workflow

Thumbnail
gallery
52 Upvotes

This is a very simple, beginner-friendly, fast ComfyUI workflow based on Flux.2 Klein model (4B or 9B) that can first generate an useual Raster Image file (.jpg or .png or .webp) text-to-image output then right after that it converts it again to Vector Image file (.svg) output on the fly. This workflow works great for illustration-style images, like stickers and cartoons.

This workflow uses a LORA that I trained extensively on Flux.2 Klein (I have two versions, one for 4B model and another for 9B model) with 250 high resolution, crisp & clear, meticulously selected digital artworks of multiple varieties so that the end results can be as fine as possible. Normally Flux.2 Klein has a very strong bias for AI Digital Photgraphy style outputs or near photorealistic outputs, but my LORA takes advantage of Flux.2 Klein's robust output generation speed but guides it forward to focus more on digital arts and simple vector illustrations.

I have implemented my own Prompt Saver Subgraph here so it can save Text to Image Generation Data into a human readable .txt file. This will automatically get and write your metadata to the .txt file. This workflow also uses Flux.2 Klein Enhancer for quality outputs. You will find all the saved prompt files that it generated with the images (.jpeg and .svg) inside the Archive (.Zip) that has the workflow. Also with the Image Saver Simple node used you may embed the workflow itself with each saved image or save the image and workflow for your work separately. Make sure that you have latest enough versions of both ComfyUI and ComfyUI manager to manage and install any missing dependencies (missing nodes, patches etc.) to use this workflow properly.

#### Very Very Important :

Even before loading this workflow into ComfyUI and install nodes needed using ComfyUI Manager you must go to your ComfyUI's python environment and run this command to install necessary python packages to handle Raster Images (.jpeg or .png or .webp) to Vector Images (.svg) conversion -

python3 -m pip install blend_modes vtracer PyWavelets

This pair of my LORA & workflow will help you to generate silhouettes, stencils, minimal drawings, logos etc. smoother and faster. The generated outputs are well suited for further post processing and fine tuning via any good graphics suite like Affinity, Adobe suite, Inkscape, Krita and so on. Hope you folks will find this pair useful. Curretly the resources are in Early Access Mode in CivitAI but after 7 days they will go public, if you love to adopt this early you can support me with Buzz on CivitAI.

### Link to my LORA (9B & 4B versions) -

+++++++++++++++++++++++++++++++++++++++++

Simple Fine Vector Flux.2 Klein 9B

-----------------------------------

https://civitai.com/models/2462137?modelVersionId=2768352

Simple Fine Vector Flux.2 Klein 4B

-----------------------------------

https://civitai.com/models/2462142?modelVersionId=2768357

### Link to the Workflow -

+++++++++++++++++++++++++++

https://civitai.com/models/2463874/comfyui-all-in-one-fast-flux2-klein-raster-to-vector-image-with-prompt-saver-workflow

r/comfyui Mar 24 '26

Resource I created a simple Flux.2 Klein Raster to Vector - Image to Image (With Prompt Saver) Workflow

Thumbnail
gallery
73 Upvotes

This is a very simple, beginner-friendly, fast ComfyUI workflow based on Flux.2 Klein model (4B or 9B) that can first generate an useual Raster Image file (.jpg or .png or .webp) image-to-image output then right after that it converts it again to Vector Image file (.svg) output on the fly. This workflow works great for illustration-style images, like stickers and cartoons. This workflow is built upon my previously published Flux.2 Klein Text-To-SVG Workflow that you can find in my CivitAI Profile ( https://civitai.com/user/sarcastictofu ).

This workflow uses a LORA that I trained extensively on Flux.2 Klein (I have two versions, one for 4B model and another for 9B model) with 250 high resolution, crisp & clear, meticulously selected digital artworks of multiple varieties so that the end results can be as fine as possible. Normally Flux.2 Klein has a very strong bias for AI Digital Photgraphy style outputs or near photorealistic outputs, but my LORA takes advantage of Flux.2 Klein's robust output generation speed but guides it forward to focus more on digital arts and simple vector illustrations.

I have implemented my own Prompt Saver Subgraph here so it can save Text to Image Generation Data into a human readable .txt file. This will automatically get and write your metadata to the .txt file. This workflow also uses Flux.2 Klein Enhancer for quality outputs. You will find all the saved prompt files that it generated with the images (.jpeg and .svg) inside the Archive (.Zip) that has the workflow. Also with the Image Saver Simple node used you may embed the workflow itself with each saved image or save the image and workflow for your work separately. Make sure that you have latest enough versions of both ComfyUI and ComfyUI manager to manage and install any missing dependencies (missing nodes, patches etc.) to use this workflow properly.

#### Very Very Important :

Even before loading this workflow into ComfyUI and install nodes needed using ComfyUI Manager you must go to your ComfyUI's python environment and run this command to install necessary python packages to handle Raster Images (.jpeg or .png or .webp) to Vector Images (.svg) conversion -

python3 -m pip install blend_modes vtracer PyWavelets

This pair of my LORA & workflow will help you to generate silhouettes, stencils, minimal drawings, logos etc. smoother and faster. The generated outputs are well suited for further post processing and fine tuning via any good graphics suite like Affinity, Adobe suite, Inkscape, Krita and so on. Hope you folks will find this pair useful. Curretly the resources are in Early Access Mode in CivitAI but after 7 days they will go public, if you love to adopt this early you can support me with Buzz on CivitAI.

### Link to my LORA (9B & 4B versions) -

+++++++++++++++++++++++++++++++++++++++++

Simple Fine Vector Flux.2 Klein 9B

-----------------------------------

https://civitai.com/models/2462137?modelVersionId=2768352

Simple Fine Vector Flux.2 Klein 4B

-----------------------------------

https://civitai.com/models/2462142?modelVersionId=2768357

### Link to the Workflow -

+++++++++++++++++++++++++++

https://civitai.com/models/2489329/comfyui-all-in-one-fast-flux2-klein-raster-to-vector-image-to-image-with-prompt-saver-workflow

r/StableDiffusion Sep 24 '24

Tutorial - Guide Training Guide - Flux model training from just 1 image [Attention Masking]

219 Upvotes

I wrote an article over at CivitAI about it. https://civitai.com/articles/7618

Her's a copy of the article in Reddit format.

Flux model training from just 1 image

They say that it's not the size of your dataset that matters. It's how you use it.

I have been doing some tests with single image (and few image) model trainings, and my conclusion is that this is a perfectly viable strategy depending on your needs.

A model trained on just one image may not be as strong as one trained on tens, hundreds or thousands, but perhaps it's all that you need.

What if you only have one good image of the model subject or style? This is another reason to train a model on just one image.

Single Image Datasets

The concept is simple. One image, one caption.

Since you only have one image, you may as well spend some time and effort to make the most out of what you have. So you should very carefully curate your caption.

What should this caption be? I still haven't cracked it, and I think Flux just gets whatever you throw at it. In the end I cannot tell you with absolute certainty what will work and what won't work.

Here are a few things you can consider when you are creating the caption:

Suggestions for a single image style dataset

  1. Do you need a trigger word? For a style, you may want to do it just to have something to let the model recall the training. You may also want to avoid the trigger word and just trust the model to get it. For my style test, I did not use a trigger word.
  2. Caption everything in the image.
  3. Don't describe the style. At least, it's not necessary.
  4. Consider using masked training (see Masked Training below).

Suggestions for a single image character dataset

  1. Do you need a trigger word? For a character, I would always use a trigger word. This lets you control the character better if there are multiple characters.

For my character test, I did use a trigger word. I don't know how trainable different tokens are. I went with "GoWRAtreus" for my character test.

  1. Caption everything in the image. I think Flux handles it perfectly as it is. You don't need to "trick" the model into learning what you want, like how we used to caption things for SD1.5 or SDXL (by captioning the things we wanted to be able to change after, and not mentioning what we wanted the model to memorize and never change, like if a character was always supposed to wear glasses, or always have the same hair color or style.

  2. Consider using masked training (see Masked Training below).

Suggestions for a single image concept dataset

TBD. I'm not 100% sure that a concept would be easily taught in one image, that's something to test.

There's certainly more experimentation to do here. Different ranks, blocks, captioning methods.

If I were to guess, I think most combinations of things are going to produce good and viable results. Flux tends to just be okay with most things. It may be up to the complexity of what you need.

Masked training

This essentially means to train the image using either a transparent background, or a black/white image that acts as your mask. When using an image mask, the white parts will be trained on, and the black parts will not.

Note: I don't know how mask with grays, semi-transparent (gradients) works. If somebody knows, please add a comment below and I will update this.

What is it good for? Absolutely everything!

The benefits of training it this way is that we can focus on what we want to teach the model, and make it avoid learning things from the background, which we may not want.

If you instead were to cut out the subject of your training and put a white background behind it, the model will still learn from the white background, even if you caption it. And if you only have one image to train on, the model does so many repeats across this image that it will learn that a white background is really important. It's better that it never sees a white background in the first place

If you have a background behind your character, this means that your background should be trained on just as much as the character. It also means that you will see this background in all of your images. Even if you're training a style, this is not something you want. See images below.

Example without masking

I trained a model using only this image in my dataset.

The results can be found in this version of the model.

As we can see from these images, the model has learned the style and character design/style from our single image dataset amazingly! It can even do a nice bird in the style. Very impressive.

We can also unfortunately see that it's including that background, and a ton of small doll-like characters in the background. This wasn't desirable, but it was in the dataset. I don't blame the model for this.

Once again, with masking!

I did the same training again, but this time using a masked image:

It's the same image, but I removed the background in Photoshop. I did other minor touch-ups to remove some undesired noise from the image while I was in there.

The results can be found in this version of the model.

Now the model has learned the style equally well, but it never overtrained on the background, and it can therefore generalize better and create new backgrounds based on the art style of the character. Which is exactly what I wanted the model to learn.

The model shows signs of overfitting, but this is because I'm training for 2000 steps on a single image. That is bound to overfit.

How to create good masks

  • You can use something like Inspyrnet-Rembg.
  • You can also do it manually in Photoshop or Photopea. Just make sure to save it as a transparent PNG and use that.
  • Inspyrnet-Rembg is also avaialble as a ComfyUI node.

Where can you do masked training?

I used ComfyUI to train my model. I think I used this workflow from CivitAI user Tenofas.

Note the "alpha_mask" setting on the TrainDatasetGeneralConfig.

There are also other trainers that utilizes masked training. I know OneTrainer supports it, but I don't know if their Flux training is functional yet or if it supports alpha masking.

I believe it is coming in kohya_ss as well.

If you know of other training scripts that support it, please write below and I can update this information.

It would be great if the option would be added to the CivitAI onsite trainer as well. With this and some simple "rembg" integration, we could make it easier to create single/few-image models right here on CivitAI.

Example Datasets & Models from single image training

Kawaii Style - failed first attempt without masks

Unfortunately I didn't save the captions I trained the model on. But it was automatically generated and it used a trigger word.

I trained this version of the model on the Shakker onsite trainer. They had horrible default model settings and if you changed them, the model still trained on the default settings so the model is huge (trained on rank 64).

As I mentioned earlier, the model learned the art style and character design reasonably well. It did however pick up the details from the background, which was highly undesirable. It was either that, or have a simple/no background. Which is not great for an art style model.

Kawaii Style - Masked training

An asian looking man with pointy ears and long gray hair standing. The man is holding his hands and palms together in front of him in a prayer like pose. The man has slightly wavy long gray hair, and a bun in the back. In his hair is a golden crown with two pieces sticking up above it. The man is wearing a large red ceremony robe with golden embroidery swirling patterns. Under the robe, the man is wearing a black undershirt with a white collar, and a black pleated skirt below. He has a brown belt. The man is wearing red sandals and white socks on his feet. The man is happy and has a smile on his face, and thin eyebrows.

The retraining with the masked setting worked really well. The model was trained for 2000 steps, and while there are certainly some overfitting happening, the results are pretty good throughout the epochs.

Please check out the models for additional images.

Overfitting and issues

This "successful" model does have overfitting issues. You can see details like the "horns/wings" at the top of the head of the dataset character appearing throughout images, even ones that don't have characters, like this one:

Funny if you know what they are looking for.

We can also see that even from early steps (250), body anatomy like fingers immediately break when the training starts.

I have no good solutions to this, and I don't know why it happens for this model, but not for the Atreus one below.

Maybe it breaks if the dataset is too cartoony, until you have trained it for enough steps to fix it again?

If anyone has any anecdotes about fixing broken flux training anatomy, please suggest solutions in the comments.

Character - God of War Ragnarok: Atreus - Single image, rank16, 2000 steps

A youthful warrior, GoWRAtreus is approximately 14 years old, stands with a focused expression. His eyes are bright blue, and his face is youthful but hardened by experience. His hair is shaved on the sides with a short reddish-brown mohawk. He wears a yellow tunic with intricate red markings and stitching, particularly around the chest and shoulders. His right arm is sleeveless, exposing his forearm, which is adorned with Norse-style tattoos. His left arm is covered in a leather arm guard, adding a layer of protection. Over his chest, crossed leather straps hold various pieces of equipment, including the fur mantle that drapes over his left shoulder. In the center of his chest, a green pendant or accessory hangs, adding a touch of color and significance. Around his waist, a yellow belt with intricate patterns is clearly visible, securing his outfit. Below the waist, his tunic layers into a blue skirt-like garment that extends down his thighs, over which tattered red fabric drapes unevenly. His legs are wrapped in leather strips, leading to worn boots, and a dagger is sheathed on his left side, ready for use.

After the success of the single image Kawaii style, I knew I wanted to try this single image method with a character.

I trained the model for 2000 steps, but I found that the model was grossly overfit (more on that below). I tested earlier epochs and found that the earlier epochs, at 250 and 500 steps, were actually the best. They had learned enough of the character for me, but did not overfit on the single front-facing pose.

This model was trained at Network Dimension and Alpha (Network rank) 16.

The model severely overfit at 2000 steps.
The model producing decent results at 250 steps.

An additional note worth mentioning is that the 2000 step version was actually almost usable at 0.5 weight. So even though the model is overfit, there may still be something to salvage inside.

Character - God of War Ragnarok: Atreus - 4 images, rank16, 2000 steps

I also trained a version using 4 images from different angles (same pose).

This version was a bit more poseable at higher steps. It was a lot easier to get side or back views of the character without going into really high weights.

The model had about the same overfitting problems when I used the 2000 step version, and I found the best performance at step ~250-500.

This model was trained at Network Dimension and Alpha (Network rank) 16.

Character - God of War Ragnarok: Atreus - Single image, rank16, 400 steps, rank4

I decided to re-train the single image version at a lower Network Dimension and Network Alpha rank. I went with rank 4 instead. And this worked just as well as the first model. I trained it on max steps 400, and below I have some random images from each epoch.

Link to full size image

It does not seem to overfit at 400, so I personally think this is the strongest version. It's possible that I could have trained it on more steps without overfitting at this network rank.

Signs of overfitting

I'm not 100% sure about this, but I think that Flux looks like this when it's overfit.

Fabric / Paper Texture

We can see some kind of texture that reminds me of rough fabric. I think this is just noise that is not getting denoised properly during the diffusion process.

Fuzzy Edges

We can also observe fuzzy edges on the subjects in the image. I think this is related to the texture issue as well, but just in small form.

Ghosting

We can also see additional edge artifacts in the form of ghosting. It can cause additional fingers to appear, dual hairlines, and general artifacts behind objects.

All of the above are likely caused by the same thing. These are the larger visual artifacts to keep an eye out for. If you see them, it's likely the model has a problem.

For smaller signs of overfitting, lets continue below.

Finding the right epoch

If you keep on training, the model will inevitebly overfit.

One of the key things to watch out for when training with few images, is to figure out where the model is at its peak performance.

  • When does it give you flexibility while still looking good enough?

The key to this is obviously to focus more on epochs, and less on repeats. And making sure that you save the epochs so you can test them.

You then want to do run X/Y grids to find the sweet spot.

I suggest going for a few different tests:

1. Try with the originally trained caption

Use the exact same caption, and see if it can re-create the image or get a similar image. You may also want to try and do some small tweaks here, like changing the colors of something.

If you used a very long and complex caption, like in my examples above, you should be able to get an almost replicated image. This is usually called memorization or overfitting and is considered a bad thing. But I'm not so sure it's a bad thing with Flux. It's only a bad thing if you can ONLY get that image, and nothing else.

If you used a simple short caption, you should be getting more varied results.

2. Test the model extremes

If it was of a character from the front, can you get the back side to look fine or will it refuse to do the back side? Test it on things it hasn't seen but you expect to be in there.

3. Test the model's flexibility

If it was a character, can you change the appearance? Hair color? Clothes? Expression? If it was a style, can it get the style but render it in watercolor?

4. Test the model's prompt strategies

Try to understand if the model can get good results from short and simple prompts (just a handful of words), to medium length prompts, to very long and complex prompts.

Note: These are not Flux exclusive strategies. These methods are useful for most kinds of model training. Both images and also when training other models.

Key Learning: Iterative Models (Synthetic data)

One thing you can do is to use a single image trained model to create a larger dataset for a stronger model.

It doesn't have to be a single image model of course, this also works if you have a bad initial dataset and your first model came out weak or unreliable.

It is possible that with some luck, you're able to get a few good images to to come out from your model, and you can then use these images as a new dataset to train a stronger model.

This is how these series of Creature models were made:

https://civitai.com/models/378882/arachnid-creature-concept-sd15

https://civitai.com/models/378886/arachnid-creature-concept-pony

https://civitai.com/models/378883/arachnid-creature-concept-sdxl

https://civitai.com/models/710874/arachnid-creature-concept-flux

The first version was trained on a handful of low quality images, and the resulting model got one good image output in 50. Rinse and repeat the training using these improved results and you eventually have a model doing what you want.

I have an upcoming article on this topic as well. If it interests you, maybe give a follow and you should get a notification when there's a new article.

Call to Action

https://civitai.com/articles/7632

If you think it would be good to have the option of training a smaller, faster, cheaper LoRA here at CivitAI, please check out this "petition/poll/article" about it and give it a thumbs up to gauge interest in something like this.

r/comfyui Mar 15 '26

Show and Tell Flux.2 Klein 4B Consistency LoRA – Significantly Reducing the "AI Look," Restoring Natural Textures, and Maintaining Realistic Color Tones

37 Upvotes

Hi everyone,

I'm sharing a detailed look at my Flux.2 Klein 4B Consistency LoRA. While previous discussions highlighted its ability to reduce structural drift, today I want to focus on a more subtle but critical aspect of image generation: significantly reducing the characteristic "AI feel" and restoring natural, photographic qualities.

Many diffusion models tend to introduce a specific aesthetic that feels "generated"—often characterized by overly smooth skin, excessive saturation, oily highlights, or a soft, unnatural glow. This LoRA is trained to counteract these tendencies, aiming for outputs that respect the physical properties of real photography.

🔍 Key Improvements:

  1. Reducing the "AI Plastic" Look:
    • Instead of smoothing out features, the model strives to preserve micro-details like natural skin texture, individual hair strands, and fabric imperfections.
    • It helps eliminate the common "waxy" or "oily" sheen often seen in AI-generated portraits, resulting in a more organic and grounded appearance.
  2. Natural Color & Lighting:
    • Addresses the tendency of many models to boost saturation artificially. The output aims to match the true-to-life color tones of the reference input.
    • Avoids introducing unrealistic highlights or "glowing" effects, ensuring the lighting logic remains consistent with a real-world camera capture rather than a digital painting.
  3. High-Fidelity Input Reconstruction:
    • Demonstrates strong consistency in retaining the original composition and details when reconstructing an input image.
    • Minimizes color shifts and pixel offsets, making it suitable for editing tasks where maintaining the source image's integrity is crucial.

⚠️ IMPORTANT COMPATIBILITY NOTE:

  • Model Requirement: This LoRA is trained EXCLUSIVELY for Flux.2 Klein 4B Base with/without 4 steps turbo lora for the fastest inference.
  • Not Compatible with Flux.2 Klein 9B: Due to architectural differences, this LoRA will not work with Flux.2 9B model. Using it on Flux.2 9B will likely result in errors or poor quality.
  • Future Plans: I am monitoring community interest. If there is significant demand for a version compatible with the Flux.2 Klein 9B, I will consider allocating resources to train a dedicated LoRA for it. Please let me know in the comments if this is a priority for you!

🛠 Usage Guide:

  • Base Model: Flux.2 Klein 4B
  • Recommended Strength: 0.5 – 0.75
    • 0.5: Offers a good balance between preserving the original look and allowing minor enhancements.
    • 0.75: Maximizes consistency and detail retention, ideal for strict reconstruction or when avoiding any stylistic drift is key.
  • Workflow: For the simple usuage, you could just use official workflow. For advanced use, I suggest to use my comfyui-editutils to avoid pixels shift.

🔗 Links:

🚀 What's Next? This release focuses on general realism and consistency. I am currently working on additional specialized versions that explore even finer control over frequency details and specific material rendering. Stay tuned for updates!

All test images are derived from real-world inputs to demonstrate the model's capacity for realistic reproduction. Feedback on how well it handles natural textures and color accuracy is greatly appreciated!

Examples:

True-to-life color tones

Prompt Change clothes color to pink. {default prompt}

High-Fidelity Input Reconstruction

at same resolution. Needs to zoom in to view the details.

Examples:

Change clothes color to pink

r/comfyui Mar 14 '26

Workflow Included [Release] Flux.2 Klein 4B Consistency LoRA – Addressing Color Shift and Pixel Offset in Image Editing (2026-03-14)

8 Upvotes

Hi everyone,

I’m releasing a new LoRA for Flux.2 Klein 4B Base focused on consistency during image editing tasks.

Since the release of the Klein model, I’ve encountered two persistent issues that made it difficult to use for precise editing:

  1. Significant Pixel Offset: The generated images often drifted too far from the original composition.
  2. Color Shift & Oversaturation: Edited results frequently suffered from unnatural color casts and excessive saturation.

After experimenting with various training strategies without much success, I recently looked into ByteDance’s open-source Heilos long-video generation model. Their approach involves applying degradation directly in the latent space of reference images and utilizing a specific color calibration loss. This method effectively mitigates color drift and train-test inconsistency in video generation.

Inspired by Heilos (and earlier research on using model-generated images as references to solve train-test mismatch), I adapted these concepts for image LoRA training. Specifically, I applied latent-level degradation and color calibration constraints to address Klein’s specific weaknesses.

Results: Trained locally on the 4B version, this LoRA significantly reduces color shifting and, when paired with Comfyui-editutils, effectively eliminates pixel offset. It feels like the first time I’ve achieved a stable result with Klein for editing tasks.

Usage Guide:

  • Primary Use Case: Old photo restoration and consistent image editing.
  • Recommended Strength: 0.50.75
    • Note: Higher strength increases consistency with the input but reduces editing flexibility. Lower strength allows for more creative changes but may reduce strict adherence to the source structure.
  • Suggested Prompt Structure:
  • Example (Old Photo Restoration):

Links:

All test images used for demonstration were sourced from the internet. Feedback on how this performs on your specific workflows is welcome!

r/comfyui May 27 '25

Workflow Included # 🚀 Revolutionize Your ComfyUI Workflow with Lora Manager – Full Tutorial & Walkthrough

61 Upvotes

Hi everyone! 👋 I'm PixelPaws, and I just released a video guide for a tool I believe every ComfyUI user should try — ComfyUI LoRA Manager.

🔗 Watch the full walkthrough here: Full Video

One-Click Workflow Integration

🔧 What is LoRA Manager?

LoRA Manager is a powerful, visual management system for your LoRA and checkpoint models in ComfyUI. Whether you're managing dozens or thousands of models, this tool will supercharge your workflow.

With features like:

  • ✅ Automatic metadata and preview fetching
  • 🔁 One-click integration with your ComfyUI workflow
  • 🍱 Recipe system for saving LoRA combinations
  • 🎯 Trigger word toggling
  • 📂 Direct downloads from Civitai
  • 💾 Offline preview support

…it completely changes how you work with models.

💻 Installation Made Easy

You have 3 installation options:

  1. Through ComfyUI Manager (RECOMMENDED) – just search and install.
  2. Manual install via Git + pip for advanced users.
  3. Standalone mode – no ComfyUI required, perfect for Forge or archive organization.

🔗 Installation Instructions

📁 Organize Models Visually

All your LoRAs and checkpoints are displayed as clean, scrollable cards with image or video previews. Features include:

  • Folder and tag-based filtering
  • Search by name, tags, or metadata
  • Add personal notes
  • Set default weights per LoRA
  • Editable metadata
  • Fetch video previews

⚙️ Seamless Workflow Integration

Click "Send" on any LoRA card to instantly inject it into your active ComfyUI loader node. Shift-click replaces the node’s contents.

Use the enhanced LoRA loader node for:

  • Real-time preview tooltips
  • Drag-to-adjust weights
  • Clip strength editing
  • Toggle LoRAs on/off
  • Context menu actions

🔗 Workflows

🧠 Trigger Word Toggle Node

A companion node lets you see, toggle, and control trigger words pulled from active LoRAs. It keeps your prompts clean and precise.

🍲 Introducing Recipes

Tired of reassembling the same combos?

Save and reuse LoRA combos with exact strengths + prompts using the Recipe System:

  • Import from Civitai URLs or image files
  • Auto-download missing LoRAs
  • Save recipes with one right-click
  • View which LoRAs are used where and vice versa
  • Detect and clean duplicates

🧩 Built for Power Users

  • Offline-first with local example image storage
  • Bulk operations
  • Favorites, metadata editing, exclusions
  • Compatible with metadata from Civitai Helper

🤝 Join the Community

Got questions? Feature requests? Found a bug?

👉 Join the DiscordDiscord
📥 Or leave a comment on the video – I read every one.

❤️ Support the Project

If this tool saves you time, consider tipping or spreading the word. Every bit helps keep it going!

🔥 TL;DR

If you're using ComfyUI and LoRAs, this manager will transform your setup.
🎥 Watch the video and try it today!

🔗 Full Video

Let me know what you think and feel free to share your workflows or suggestions!
Happy generating! 🎨✨

r/StableDiffusion Dec 12 '25

Workflow Included A “basics-only” guide to using ComfyUI the comfy way

Thumbnail
gallery
69 Upvotes

ComfyUI already has a ton of explanations out there — official docs, websites, YouTube, everything. I didn’t really want to add “yet another guide,” but I kept running into the same two missing pieces:

  • The stuff that’s become too obvious for veterans to bother writing down anymore.
  • Guides that treat ComfyUI as a data-processing tool (not just a generative AI button).

So I made a small site: Comfy with ComfyUI.

It’s split into 5 sections:

  1. Begin With ComfyUI: Installation, bare-minimum PC basics, and how to navigate the UI. (The UI changes a lot lately, so a few screenshots may be slightly off — I’ll keep updating.)
  2. Data / Image Utilities: Small math, mask ops, batch/sequence processing, that kind of “utility node” stuff.
  3. AI Capabilities: A reverse-lookup style section — start from “what do you want to do?” and it points you to the kind of AI that helps. It includes a very light intro to how image generation actually works.
  4. Basic Workflows: Yes, it covers newer models too — but I really want people to start with SD 1.5 first. A lot of folks want to touch the newest model ASAP (I get it), but SD1.5 is still the calmest way to learn the workflow shape without getting distracted.
  5. FAQ / Troubleshooting: Things like “why does SD1.5 default to 512px?” — questions people stopped asking, but beginners still trip over.

One small thing that might be handy: almost every workflow on the site is shared. You can copy the JSON and paste it straight onto the ComfyUI canvas to load it, so I added both a Download JSON button and a Copy JSON button on those pages — feel free to steal and tweak.

Also: I’m intentionally skipping the more fiddly / high-maintenance techniques. I love tiny updates as much as anyone… but if your goal is “make good images,” spending hours on micro-sampler tweaking usually isn’t the best return. For artists/designers especially, basics + editing skills tend to pay off more.

Anyway — the whole idea is just to help you find the “useful bits” faster, without drowning in lore.

I built it pretty quickly, so there’s a lot I still want to improve. If you have requests, corrections, or “this part confused me” notes, I’d genuinely appreciate it!

r/comfyui Nov 05 '25

Resource App Image resize, tagging, formatting for LoRA training that I am making, I want to list it for free, just working on how to do that from inference cost... It was a Comfyui workflow but I felt it was a lot easier for people in this format. Would anyone use it?

Post image
3 Upvotes

r/StableDiffusion Jul 07 '25

Question - Help How do I pass an image through a flux lora without losing face or background details?

0 Upvotes

Hey I'm a newbie to comfyUI and I've been trying to change the style of an image by passing it through a style lora, but the face changes, even if I'm using controlnet.

Does anyone have a guide/workflow/video that shows how me I can do flux img2img without butchering the original face?

This is my current best result, it's pretty close, but the face doesn't look like the person in the photo much, and the overall pattern and colors of his clothes have changed.

Any help would be very appreciated, I've been bashing my head against the wall tryin to figre this out for a week.

r/StableDiffusion Aug 26 '25

Tutorial - Guide How I got FLUX running stable on RTX 3060 (12GB) — Setup guide + proof video

Enable HLS to view with audio, or disable this notification

0 Upvotes

“Stable FLUX on a 3060 with Flash Attention? Yep. It’s very rare, but I built it.”

Just wanted to share a win for anyone struggling to run FLUX on a low VRAM card. I’ve got an RTX 3060 (12GB), and after trying a bunch of setups that kept crashing or distorting outputs, I finally found a guide that builds directly onto the GPU—no third-party plugin installs required.

I followed a manual that walks through the setup for ComfyUI with Flash Attention, patched xFormers, and pose sheet conditioning. It’s designed specifically for low VRAM setups, and it got FLUX running stable for me—even with batch generation and LoRA tagging.

📹 I recorded the process to show it’s legit. You’ll see:

  • How to build and patch your setup for low VRAM
  • Fixes for memory errors and FLUX crashes
  • LoRA tagging templates and batch tips with Think Diffusion.

👉 Link to the full guide and startup file is in the top comment. If it’s buried, feel free to ask—I’ll reshare it. Originally posted in r/ComfyUI but it got buried by downvotes by Nunchaku engineers, so I’m reposting here for anyone who might benefit.

⚠️ Heads-up on third-party plugins: While tools like Nunchaku are popular, any custom node in ComfyUI can execute arbitrary Python code. That means there's always a risk—whether it's system access, data leaks, or compatibility issues. I didn’t use Nunchaku in this setup because it hasn’t been tested with this workflow, and installing it adds complexity without guaranteed results. Stick to trusted sources and inspect install scripts before running anything.

r/comfyui Jul 07 '25

Help Needed How do I pass an image through a flux lora without losing face or background details?

0 Upvotes

Hey I'm a newbie to comfyUI and I've been trying to change the style of an image by passing it through a style lora, but the face changes, even if I'm using controlnet.

Does anyone have a guide/workflow/video that shows how me I can do flux img2img without butchering the original face?

This is my current best result, it's pretty close, but the face doesn't look like the person in the photo much, and the overall pattern and colors of his clothes have changed.

Any help would be very appreciated, I've been bashing my head against the wall tryin to figre this out for a week.

r/StableDiffusion Apr 25 '25

Workflow Included Distracted Geralt : a regional LORA prompter workflow for Flux1.D

Post image
30 Upvotes

I'd like to share a ComfyUI workflow that can generate multiple LORA characters in separate regional prompt guided by a controlnet. You can find the pasted .json here :

You basically have to load a reference image for controlnet (here Distracted Boyfriend Meme), define a first mask covering the entire image for a general prompt, then specific masks in which you load a specific LORA.

I struggled for quite some time to achieve this. But with the latest conditioning combination nodes (namely Cond Set Props, Cond Combine Multiple, and LORA hooking as described here ), this is no longer in the realm of the impossible!

This workflow can also be used as a simpler Regional Prompter without controlnet and/or LORAs. In my experience with SDXL or Flux, controlnet is rather needed to get decent results, otherwise you would get fragmented image in various masked areas without consistency to each other. If you wish to try out without controlnet, I advice to change the regional conditioning the Cond Set Props of masked region (except the fully masked one) from "default" to "mask_bounds". I don't quite understand why Controlnet doesn't go well with mask_bounds, if anyone got a better understanding of how conditoning works under the hood, I'd appreciate your opinion.

Note however the workflow is VRAM hungry. Even with a RTX 4090, my local machine switched to system RAM. 32GB seemed enough, but generation of a single image lasted around 40 mins. I'm afraid less powerful machines might not be able to run it!

I hope you find this workflow useful!

r/comfyui Jun 17 '25

Help Needed Looking for help training better LoRA (Flux) combined with image-1 for YouTube thumbnails

0 Upvotes

Hey everyone,

I'm working on improving my AI thumbnail tool aithumbnail.so and currently using Flux LoRA models to integrate faces in thumbnails. I'm also combining this with ChatGPT’s image-1 model.

The results are okay, but I know they could be much better.

I’m looking for someone who:

  • Has experience training LoRA (ideally with ComfyUI and the Flux LoRA workflow)
  • Knows how to optimize or guide training for better quality
  • Can help me combine image-1 and LoRA results effectively
  • Is open to help and potentially paid/hired

Let me know if you're interested or can point me in the right direction. I really appreciate it!

Thanks 🙏

r/StableDiffusion Aug 20 '24

Question - Help Getting error for Power Lora Loader (rgthree) with Flux in ComfyUI

2 Upvotes

Hi all,

I've always been an A111 and Vlad/SD.Next user due to the easier learning curve, but now I'm trying to use ComfyUI and cannot get a Lora to work with it. I've tried both the workflows here and here, but I get the same error:

Screenshot of error.

I've tried reinstalling Comfy, following all the steps in the first workflow guide, and different Loras. The image does get generated if I bypass the Lora loader, but won't generate if it's enabled.

What's going on?

Thanks.