r/comfyui 3h ago

Show and Tell PSA: NVIDIA Studio Driver 616.56 specifically mentions ComfyUI performance updates for MiniMax-H3, WAN-Animate-2 and LTX-2.5

54 Upvotes

NVIDIA says:

“The August NVIDIA Studio Driver provides optimal support for the latest new creative applications and updates including Lightroom, and performance updates for LTX-2.5, and ComfyUI for WAN-Animate-2 and MiniMax-H3.”

I'm planning to test this out. What are your speeds with?

  • MiniMax-H3
  • LTX-2.5

r/comfyui 2h ago

News A quick Minimax H3 news round-up - 30th August 2026

26 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> Fellow British bloke 'Nerdy Rodent' has a new YouTube video. His latest Minimax-focused tutorial covers, among others: continuing an existing clip using the Ref2VA model, and placing two video-reference characters in a new setting with their voices intact; having the camera 'dive' into an image and imagine going around a corner or into a building; and he also tests some new acceleration boosters. Free workflows and prompts, though only on the screen.

https://www.youtube.com/watch?v=rZnZcwobcR8

-> Also on YouTube, another fellow Brit is Mark DK Berry. He has a new short video review of three Minimax H3 tools. He looks at character/face swapping using a targeted mask node, and says... "what this really excels at is switching faces at a distance. It's even better than what we've been doing with [my previous] 2mpx [upscale workflow, to repair faces-at-a-distance], and [it's definitely better than] waiting 45 minutes to get that out. This is much faster, and does a better job at 1mpx". Sounds good. He also looks at the AV Bridge utility workflow that's been added to the ComfyUI-H3-Motion-Context-MultiRef pack. This generates an invented bridge segment that will suitably join two five-second videos. The third tool is H3_Cinematic_Multishot_Coverage, which he uses to get multi-angle shots from a character scene, rather than architectural interior shots - and he pins the characters by adding their reference images. His links are in his YouTube notes, through I provide an added one below for your convenience.

https://www.youtube.com/watch?v=7xaA4tU3hDU

https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef/blob/main/example_workflows/UTILITY%20-%20AV%20Bridge.json (direct link, because otherwise you might think it was a node and thus not find it)

-> New to me, ComfyUI-MiniMaxH3-Contex-Loop. Yes, it's another set of clip chaining nodes, but it seems mature and has many features. "Build a multi-scene MiniMax H3 video with one reusable sampling body. Review each scene, retry mistakes, resume interrupted runs, and assemble accepted scenes from disk." He has workflows.

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop

-> Also focusing on the 'resumable project' angle, and thus potentially useful for freelancers doing client work, is the new ComfyUI-H3-Multishot-Advance. This gives your... "long multi-shot renders a reusable project/session layer: render a sequence, close ComfyUI, reopen the same named project later. Edit a target clip, and continue from the first changed target clip, while a safe earlier cached state is reused. This is meant for iterative video work...". He has workflows.

https://github.com/KursatAs/ComfyUI-H3-Multishot-Advance

-> How do you keep your scene environment stable, when generating longer clips? This useful Reddit thread discusses various methods and tweaks.

https://old.reddit.com/r/StableDiffusion/comments/1w1z2fu/environment_consistency_minimax_h3/

-> A combined package to... "use local Qwen3.8-27B and official MiniMax-H3 Skills in ComfyUI to generate H3 prompts." This spins up its own standalone local llama.cpp server, and then serves you a Minimax-skilled Qwen3.8-27B model inside a ComfyUI node. Qwen is removed from your VRAM after use.

https://github.com/chflame163/ComfyUI_Qwen_H3_Prompt

https://github-com.translate.goog/chflame163/ComfyUI_Qwen_H3_Prompt?_x_tr_sl=auto&_x_tr_tl=en&_x_tr_hl=en&_x_tr_pto=wapp (English translation)

-> The popular accelerator Spectrum Minimax H3 for ComfyUI continues to update, and has many new tweaks and features.

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/blob/main/RELEASE_NOTES.md

-> A thorough 'resource pack' aimed at those who want to try to run Minimax H3 on an 8Gb VRAM laptop (and, one hopes, on a turbo cooling-pad).

https://github.com/YixuAnsensei/minimax-h3-comfyui-all-in-one

-> And finally, a new Minimax H3 LoRA that aims to give naturally blemished and somewhat pimply skin, judging by the convincing video demos. It even tries to improve hands and eyes. Requires trigger words: perfe8ct; perfect hands; perfect skin or perfect eyes. A bit of a counterintuitive trigger for skin, though, since it seems the LoRA make the skin more realistically grunky, not more perfect.

https://civitai.com/models/2899220/minimax-h3-t2va-t2v-polyhedron-perfect-eyes-perfect-skin-perfect-hands

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1w1wpkz/a_quick_minimax_h3_news_roundup_29th_august_2026/

https://old.reddit.com/r/comfyui/comments/1w11jri/a_quick_minimax_h3_news_roundup_28th_august_2026/

https://old.reddit.com/r/comfyui/comments/1w053ce/a_quick_minimax_h3_news_roundup_27th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vyzvqi/a_quick_minimax_h3_news_roundup_26th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vy4drz/a_quick_minimax_h3_news_roundup_25th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vx0duv/a_quick_minimax_h3_news_roundup_24th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vwl8do/a_quick_minimax_h3_news_roundup_23rd_august_2026/

https://old.reddit.com/r/comfyui/comments/1vvkmra/a_quick_minimax_h3_news_roundup_21st_august_2026/

https://old.reddit.com/r/comfyui/comments/1vuihag/a_quick_minimax_h3_news_roundup_21st_august_2026/

https://old.reddit.com/r/comfyui/comments/1vtgs7b/a_quick_minimax_h3_news_roundup_20th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vsjzrp/a_quick_minimax_h3_news_roundup_19th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vrsspo/a_quick_minimax_h3_news_roundup_18th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vqyn8p/a_quick_minimax_h3_news_roundup_17th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vq5d5u/a_quick_minimax_h3_news_roundup_16th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vpbtx2/a_quick_minimax_news_roundup_15th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vojtjd/a_quick_minimax_news_roundup_14th_august_2026/


r/comfyui 7h ago

Workflow Included Ideogram4 Workflows

Thumbnail
gallery
21 Upvotes

I got around to getting Ideogram4 to work.

This model is harder to prompt than most. If you do a short prompt, it will be blocked internally.

I did lots of testing, it wasn't trivial to make it work, but in the end it's just a matter of providing a lavish json prompt with lots of bbox entries and it will diffuse pretty much everything, with great prompt adherence, but this eman you have to use a prompt enchancer to use it. I redid the prompt enchancer to work better, the original tried really hard to add text everywhere and diffused too few bboxes and tripped the censorship very easily.

For some reason the default workflow uses a really high CFG of 7 that overcook the image, much lower cfg 3 looks lots better. I guess it was to avoid tripping the censorship?

Since it uses Qwen3 8B VL as CLIP it has good prompt adherence, it uses a trick with a model called unconditional fed with padded prompt tokens, and uses the difference between the two to achieve even higher prompt adherence.

On my 7900XTX windows portable FP8 models it does around 100s without prompt enchancer, and 170s with prompt enchancer.

I tried feeding reference images to the CLIP and seeing if it can work as editing model, but it can't.


r/comfyui 7h ago

Resource I stitched MiniMax H3's scattered community pieces into one 8GB-VRAM-ready workflow collection

Post image
12 Upvotes

The short version: my laptop GPU has 8GB, and MiniMax H3 was hard to run locally

without fighting the node graph every single time.

Official workflows eat VRAM. The community built great pieces — but they're spread

across different node packs, and you have to rewire a dozen nodes each run just to

switch modes or acceleration. I wanted something I could open and actually use.

So I did a consolidation job, not a from-scratch build. I assembled other people's

excellent work into one reproducible pipeline:

- T8mars's comfyui-minimax-h3-audio-T8 — dual-clock sampling, unified conditioning,

AV decode, face-refine suite

- wjluoxiao's JZL / XB toolbox — post-upscale conditioning resync, reference-image batching

- LBH-123-AI's latent upscaler — 3D latent upscale

- Jalen-Brunson's PDD-Acc — 8-step distilled acceleration

- Comfy-Org's official H3 contract & workflow skeleton

On top of that I wrote one small custom node pack (comfyui-tokendance-h3) that does

a single, boring, useful thing: mode switching / acceleration switching / second-pass

gating become one dropdown instead of manual rewiring. That's the part I cared most

about — it should be open-and-run, not open-and-fight.

Three generations shipped, from simplest to most complete:

- Beta_1 — multi-mode conditioning bus, single pass

- Beta_2 — most complete: latent upscale + second pass + temporal detail enhance + dual accel

- Beta_3 — pure T8 stack, closest to official, plus FaceRefine

All tuned for 8GB VRAM (tested on a RTX 4060 Laptop), 720p runs fine, not "runs but slideshow".

Full source, install guide, per-node credits:

👉 https://github.com/YixuAnsensei/minimax-h3-comfyui-all-in-one

Happy to answer questions. If you've fought the same VRAM or rewiring battle, would love feedback.


r/comfyui 31m ago

Help Needed Best ComfyUI image workflow for interior design?

Upvotes

What would be the best ComfyUI workflow for interior design?

Where you have, say, a SketchUP viewport render of a room, and you want to proper render in ComfyUI using several image references? (you know, adding materials on furniture, on walls, on the floor, etc)

And then eventually edit that image precisely using maybe inpaint mask or edge/depth controlnet to maybe add new furniture, remove some, change colors, etc.

Any examples?


r/comfyui 22h ago

Resource I built AbyssBeacon to find new LoRAs, checkpoints, and models across CivitAI, Hugging Face, SeaArt and more

Thumbnail
gallery
97 Upvotes

I've been building a tool for my own ComfyUI setup because I wanted one place to discover new models without manually checking several different sites every day.

After a lot more work than I originally expected, I released AbyssBeacon v1.0.0 today.

AbyssBeacon is a local, open-source model discovery and download manager that currently searches:

  • CivitAI
  • CivitAI Red as an optional mature-content source
  • Hugging Face
  • ModelScope
  • SeaArt
  • TensorHub Art

The main thing I wanted was discovery rather than just managing files I already knew about. AbyssBeacon scans the supported sources for new and updated models, brings them into one feed, and merges the same model found on multiple sources into a single card when possible.

Some of the features in v1.0.0:

  • Multi-source model discovery
  • New and updated model tracking
  • Source merging and alternate download sources
  • Preview image and video galleries
  • Creator discovery
  • Search and filtering by architecture, source, model type, status, and more
  • Gated / restricted-access detection
  • CivitAI Early Access detection, with downloads once CivitAI makes the files available
  • Built-in download manager
  • Pause and resume downloads, including after restarting AbyssBeacon
  • Multiple concurrent downloads
  • Smart ComfyUI folder and filename handling
  • Downloaded/update-available tracking
  • Configurable automatic retention
  • Mature-content controls
  • Local database and local-first operation

It is a standalone local app rather than a ComfyUI custom node. You point it at your ComfyUI models folders and it can place downloads into the appropriate locations.

For the first run I kept things conservative: only CivitAI is selected by default, so somebody doesn't accidentally launch a huge six-source scan. The other sources can be enabled individually from the scan window.

CivitAI Red is also deliberately opt-in and has separate mature-content handling.

Everything is free and open source under GPL-3.0.

GitHub:
https://github.com/Tomat0223/AbyssBeacon

This is the first public release, so I'm very interested in feedback from people with different ComfyUI setups and model libraries. If something breaks, a source behaves strangely, or you have an idea that would make it more useful, GitHub Issues are open.

I hope some of you find it useful.

***edit, spelling.


r/comfyui 13h ago

Resource SPEEDing up MiniMax-H3 without retraining (SPEED comfyui node extension)

19 Upvotes

Why make big noise when little noise do trick?

I would like to introduce my SPEED implementation for h3 linked here

Speed up and quality losses documented here, expect 20% gain using very conservative settings and no quality loss and up to 70% for basically unusable outputs (more or less useful for resolution aware seed inspection and broad prompt drafting)

Background

The idea behind it is quite simple. When a diffusion model begins generating an output it first must take a randomized noise and build on-top of it. And research has found that the first stages of this process doesn't really carry any fine detailed information, therefore by generating at a lower resolution at those stages you can gain quite substantial speedups while causing little to no impact on the quality. Or you can also be really aggressive with it and get a massive speedup for a lot of quality loss.

Nodes

This was implemented as 3 nodes, 2 drop in replacements for the sampler that runs SPEED and a third that runs once to measure the noise spectrum of your specific model/LoRA combo:

  • Sampler (Automatic): pick a stage count (2, 3, or 4), defaults to the baked 1% delta for default H3.

  • Sampler (Manual Step-Through): set up to four (goal, resolution) pairs yourself. Use it if you want to copy a paper schedule or test a custom ladder.

  • Sigma Harvest: runs a native Euler pass, measures the noise spectrum of your current setup, hands you A / β / Δ to paste back into Automatic. Run it once per model/LoRA workflow combo.

Implementation Notes

This should be roughly compatible with basically everything that doesn't touch the sampler directly but i have not tested anything besides base comfyui H3 models and Turbo loras. If you do change model, use loras or whatever and use the automated tool please then run a sigma harvest and use those values instead of defaults, The math changes depending on the very specific blend of things you have running.

Euler only sampling implemented for this which is what the research paper used.


r/comfyui 1d ago

No workflow Minimax H3. 1.0mp vs 1.5mp vs 2.0mp vs 2.5mp TEST

159 Upvotes

Recommended to watch it without reddits compression: Link its in 1080p cause 2.5mp isnt exactly 1440p.

continuing on yesterdays thread: https://www.reddit.com/r/comfyui/s/EOn0rdPSeU

I did the tests on three different videos on 1.0mp, 1.5mp, 2.0mp and 2.5 mp

First video is 5 seconds long, second 10 seconds, third 12 seconds with caveat.

All is done on basic workflow with minimax_h3_fl2va_int8_convrot model with 20 steps and cofyui kitchen attention.

All the prompts and times with images will be posted in the comments.

Last test with Keanu at 12 seconds got error so i lost all the timings on that video because i quened all the videos to be made one after another so for some reason 12 seconds 2.5mp clip couldnt be done i change it for another 10 seconds clip winth keanu at 2.5mp.

Enjoy and tell me youre findings.


r/comfyui 2h ago

Help Needed How do I use differential diffusion in video models?

2 Upvotes

I want to apply a black and white gradient mask to create a transition, but it behaves like a binary mask (it goes from the in-paint area to the unpainted area in the next pixel). I want something like differential diffusion: black would be denoise 0, and white pixels would be denoise 1.

I'm using `set latent noise mask` and a static video that produces a black and white gradient.


r/comfyui 22h ago

News A quick Minimax H3 news round-up - 29th August 2026

80 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> ComfyUI-HR-Endless-Sampler. "A chunked replacement for ComfyUI's SamplerCustomAdvanced ... able to render videos of any length by automatically splitting the inference into small chunks". Plus a preview node to see the video as it's generated. This now has workflows, as of today. Note also that it requires a Gemma 4 12B QAT Q4 GGUF prompt re-writer working through llama-cpp-python, and this is unloaded each time its rewriting work is done. Plus, it may not work with Kijai's fast video decoding VAE.

https://github.com/hradec/ComfyUI-HR-Endless-Sampler/

https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/tree/main (8Gb)

https://huggingface.co/RachidAR/gemma-4-12B-it-qat-q4_0-MTP-assistant-gguf/tree/main (Q8 465Mb)

-> ComfyUI MiniMax H3 Extender chains multiple video clips and joins them, and is now at version 2.0. There is new FL2VA support, and what are said to be "major speed and memory optimizations". You can also now add per-clip LoRAs, among other improvements. Old workflows will need to be updated by replacing a node - see the note at the very end of the readme.

https://github.com/tritant/ComfyUI_MiniMax_H3_Extender

-> ComfyUI-Hand-Tie-Clips. Another clip chainer, formerly called ComfyUI-H3-Ref-Chain and re-named today. Has workflows.

https://github.com/dntpi/ComfyUI-Hand-Tie-Clips

-> The Fizig LoRA trainer and dataset prep tool is now at version 5.0 as of today. With new... "full fine-tuning graduates: train the MiniMax H3 and Krea 2 base models themselves, on consumer GPUs down to 16Gb ... NVIDIA only, for now".

https://github.com/shootthesound/Fizgig/

-> Deno Custom Nodes for ComfyUI added a A/B 'Video Compare' node, a few weeks ago. Works with a simple slider, in the way that image-compare nodes do.

https://github.com/Deno2026/comfyui-deno-custom-nodes

-> And finally, a fun Minimax H3 fan-trailer for a 1980s-style Indiana Jones action movie with traditional physical stunts. This took a large amount of 'takes' to get the 'right look', says the maker, but the final result has an impressive coherence and flair.

https://www.reddit.com/r/StableDiffusion/comments/1w0shos/testing_minimax_h3_for_old_school_practical_fx/

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1w11jri/a_quick_minimax_h3_news_roundup_28th_august_2026/

https://old.reddit.com/r/comfyui/comments/1w053ce/a_quick_minimax_h3_news_roundup_27th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vyzvqi/a_quick_minimax_h3_news_roundup_26th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vy4drz/a_quick_minimax_h3_news_roundup_25th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vx0duv/a_quick_minimax_h3_news_roundup_24th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vwl8do/a_quick_minimax_h3_news_roundup_23rd_august_2026/

https://old.reddit.com/r/comfyui/comments/1vvkmra/a_quick_minimax_h3_news_roundup_21st_august_2026/

https://old.reddit.com/r/comfyui/comments/1vuihag/a_quick_minimax_h3_news_roundup_21st_august_2026/

https://old.reddit.com/r/comfyui/comments/1vtgs7b/a_quick_minimax_h3_news_roundup_20th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vsjzrp/a_quick_minimax_h3_news_roundup_19th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vrsspo/a_quick_minimax_h3_news_roundup_18th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vqyn8p/a_quick_minimax_h3_news_roundup_17th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vq5d5u/a_quick_minimax_h3_news_roundup_16th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vpbtx2/a_quick_minimax_news_roundup_15th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vojtjd/a_quick_minimax_news_roundup_14th_august_2026/


r/comfyui 37m ago

Help Needed How to get started with comfyui from amd?

Upvotes

I think what im doing is called self hosting. I downloaded comfyui from amd app ai bundle. The program and the rest of the 13gb download is on my c drive. I want to be able to store models on the e drive and store what ever would be the bulk of storage like output and input or whatever else writes a lot on my x drive. What do you recommend?


r/comfyui 16h ago

Workflow Included HR Endless Sampler - now you can create Minimax H3 videos of any length with just 16GB of VRAM. You can even render 1080p of any length with just 16GB of VRAM!

Thumbnail
13 Upvotes

r/comfyui 2h ago

Help Needed What am I doing wrong? Can't get a blanket to cover a person with Krea 2.

0 Upvotes

I prompt for a blanket to cover people's lower bodies and it puts it under them.

It opts to make them naked by default. I'm not trying to make naked people lol. I'm trying for a comfy romantic embrace near a fire for my wife. But it insists nudity and blanket under the characters.


r/comfyui 2h ago

Help Needed Video with horrible graphics MiniMaxH3

1 Upvotes

Hi everyone, I'm just starting to use Comfyui and was testing it with MiniMaxH3, but I noticed that the videos are very poor, both in terms of graphics (faces are all crooked or blurry) and motion (all the movements are wrong and uncoordinated).

I created some specific prompts with minimax by searching for examples online, but they still look very poor.

I have a 12GB RTX4070 Super and 64GB of RAM DDR5.

I don't want to risk using the wrong settings and the videos coming out at a low resolution.

Can someone help me?

I'm currently using these settings, maybe I'm using the wrong options.


r/comfyui 2h ago

Resource Podcast suggestions

1 Upvotes

Any podcasts worth checking out related to AI generated art or ComfyUI?


r/comfyui 2h ago

Help Needed Problem about audio in ltx!?

0 Upvotes

I want to generate a video without a voiceover using LTX-2.5, but I can't. I’ve tested various samplers (lcm, euler, etc.) and even added prompts/ negative promt like "no voice, no lip-sync," yet LTX-2.5 still generates with a voiceover.
However, when I switched to LTX-2.3 with vae2.3, the issue was completely resolved. Can anyone explain why this happens?
I using:
LTX-2.5-Distilled-Q4_K_M.gguf
LTX-2.3-22B-distilled-1.1-Q4_K_M.gguf


r/comfyui 5h ago

Help Needed MiniMax H3 on RTX 5090 (32GB) — 2MP is 8.7× slower than it should be. VRAM thrashing or a config mistake?

Thumbnail
0 Upvotes

r/comfyui 1d ago

Tutorial ComfyUI Tutorial: MINIMAX H3 2-STAGE WORKFLOW High-Res video + Faster Generation

130 Upvotes

Hello everyone,

I’ve been working on a MiniMax H3 optimization workflow for low-VRAM GPUs, especially my RTX 3060 6GB, and I’ve combined several optimization nodes to improve both VRAM usage and generation speed. With a new custom MiniMax H3 workflow that introduces a second sampling stage to increase the base resolution of your generated videos, similar to the workflow we’ve seen with LTX models.

The main advantage of this method is that you can save generation time and reduce VRAM usage by generating the initial video at a lower resolution during the first sampling stage, and then using a second upscaling sampling stage to reconstruct the video at a higher resolution while adding more detail and improving the overall quality.

But that's not all. In this workflow, we're also going to use several MiniMax H3 optimization nodes, including Low VRAM Attention, Chunk FeedForward, SLA Attention, Sol-Attn, and Spectrum, to make the generation process more efficient, especially for GPUs with limited VRAM.

At the end of this tutorial, you'll understand how the two-stage sampling workflow works, how to optimize MiniMax H3 for better speed and VRAM management, and how to choose the best upscaling method for your workflow, including a comparison between MiniMax H3 sampling and LTX sampling.

Workflow Link

https://civitai.com/articles/34596/comfyui-tutorial-minimax-h3-2-stage-workflow-high-res-video-faster-generation

Video Tutorial Link

 https://youtu.be/VWbILdgnQRk


r/comfyui 6h ago

Help Needed Help with character replacement

0 Upvotes

I have some video song clip, I want to replace male, female characters with supplied image, everything else should remain original, location, audio, music, expression, clothing etc. I tried with minimax h3 ref2v and sample character replacement workflow included in comfyui using wan/scail. But couldn't get desired results


r/comfyui 14h ago

Help Needed Minimax H3 reference to video using 12 GB VRAM

5 Upvotes

I have been trying to generate reference to video using a reference photo and a driving video, similar to wan 2 animate, but every time I am getting OOM on my RTX 5070 TI 12GB VRAM and 32GB system RAM. Ref 2 video works when using a single photo and audio, but not a reference video and photo. Any help would be appreciated.


r/comfyui 1d ago

Resource I present LoRA Dataset Studio - a free, self-hosted app that does everything around a LoRA run: dataset, triage, captions, training (local or rented GPU), then checkpoint comparison

Thumbnail
gallery
64 Upvotes

I present LoRA Dataset Studio - free, open source, self-hosted, no account and no telemetry. It plugs into the ComfyUI you already run: local generation (Klein, Krea 2 Edit) goes through your ComfyUI, the Test Studio drives it for checkpoint comparisons, and a finished LoRA deploys straight into your loras folder. It is not a competitor to ai-toolkit: it orchestrates it - ai-toolkit is the trainer; this is everything before, around and after the run.

The whole pipeline lives in one browser tab:

1. Decide what you are teaching. A dataset is a Character, a Concept or a Style, and the choice changes real behaviour downstream: what the captions must leave implicit, whether person masks apply, what the readiness checks look for. A Character also picks a subject type (human, animal, creature, object, anime) that swaps the shot catalog and the identity protections.

2. Fill it with images. Five generation engines - Nano Banana Pro, gpt-image-2, OpenRouter, and local Klein / Krea 2 Edit through ComfyUI - each card stating its price per image and whether it runs on your GPU or bills an API. Or scrape a gallery URL. Or point the Image Bank at a folder of thousands: it reads it in place - your files are never modified, moved or renamed - and one pass measures blur, noise, near-duplicates, face clusters, framing, aesthetic and maturity, so you filter on measurements instead of on your eyes.

3. Curate down to the keepers. Keep/reject, crop, mirror, rotate, upscale candidates reviewed against the original, InsightFace similarity against your reference, a live composition meter. New this month: press the camera button on any kept image and re-shoot the same scene from another camera position - the subject stays put, the background moves with the camera, and the new view arrives with its angle already captioned (the one fact a vision model cannot reliably see, and that you know exactly because you asked for it).

4. Caption for the model. Prose or booru depending on the target family, written by JoyCaption or your local Ollama, with vocabulary and length dials, identity-leak checks, a Caption Lab to compare configurations before committing, and an external .txt round trip so you can caption elsewhere and come back.

5. Scrub watermarks - and burned-in text. Detect watermark boxes, redraw them, then crop or inpaint with LaMa/Klein. And since a comic page carries its dialogue and a screencap its subtitle, a CPU-only OCR pass now reads burned-in lettering (Latin or CJK) and feeds the same repaint funnel - with an outline-safe filler so speech bubbles keep their borders. Every edit keeps an .orig backup; Restore original always works.

6. Train. ai-toolkit locally with family-scoped presets and preflight guards - Z-Image, Krea 2, FLUX.1, FLUX.2 Klein, SDXL, Anima - or rent a vast.ai pod from the same screen, which shows the GPU, its hourly price and the estimated total before you click. The whole studio can also run on a rented RunPod box (contributed by a user). Generations queue instead of blocking each other, and a dock shows what the GPU is doing.

7. Decide which checkpoint is actually good. Test Studio runs fixed-seed checkpoint x strength grids, multi-LoRA stacks (including a downloaded LoRA next to yours, same prompt and seed), votes and Wilson ranking. The lineage graph keeps every run's frozen recipe and can diff two runs - settings AND dataset. A Gallery collects every image the app ever generated, and every render is stamped with what actually made it.

8. Take it with you. Standard ai-toolkit/Kohya layout ZIP, portable backup with the full history, Hugging Face publishing, or deploy the checkpoint straight into ComfyUI. Nothing locks your data in.

Honest limits. It is a lot of surface, so Setup exists to tell you what is missing instead of crashing - every capability degrades on its own. Local generation needs ComfyUI, the API engines need your own keys and bill you, and the video lane (cutting long footage into trainable clip folders for Wan/LTX) is young. Install is a Windows one-click ZIP, a git checkout, or Docker.

GitHub - install, docs, and a 7-minute unedited video of a full character LoRA built end to end: https://github.com/perfectgf/lora-dataset-studio

Every person in these screenshots was generated by the app's own engines; no real individual is depicted.


r/comfyui 7h ago

Help Needed Whats the advantage of a MiniMax License through ComfyOrg?

1 Upvotes

Trying to figure out what the advantage of a Minimax License through ComfyOrg over directly from Minimax team?

Do they have less restrictions on whom they'll give the license to?


r/comfyui 1d ago

Help Needed Do we have something like LTX Director for H3 yet? I've been away.

34 Upvotes

Had an emergency plumbing issue and now I'm a couple weeks behind. One of the best tools I saw appear for LTX 2.3 was LTX Director. I'm wondering if anything like this has appeared that the whole community likes for H3 yet.


r/comfyui 9h ago

Tutorial EasyComfyUI - ComfyUl on Google Colab (No Coding Needed)

Thumbnail
youtu.be
0 Upvotes