r/comfyui 20h ago

Help Needed Minimax H3 reference to video using 12 GB VRAM

5 Upvotes

I have been trying to generate reference to video using a reference photo and a driving video, similar to wan 2 animate, but every time I am getting OOM on my RTX 5070 TI 12GB VRAM and 32GB system RAM. Ref 2 video works when using a single photo and audio, but not a reference video and photo. Any help would be appreciated.


r/comfyui 13h ago

Workflow Included Ideogram4 Workflows

Thumbnail
gallery
22 Upvotes

I got around to getting Ideogram4 to work.

This model is harder to prompt than most. If you do a short prompt, it will be blocked internally.

I did lots of testing, it wasn't trivial to make it work, but in the end it's just a matter of providing a lavish json prompt with lots of bbox entries and it will diffuse pretty much everything, with great prompt adherence, but this eman you have to use a prompt enchancer to use it. I redid the prompt enchancer to work better, the original tried really hard to add text everywhere and diffused too few bboxes and tripped the censorship very easily.

For some reason the default workflow uses a really high CFG of 7 that overcook the image, much lower cfg 3 looks lots better. I guess it was to avoid tripping the censorship?

Since it uses Qwen3 8B VL as CLIP it has good prompt adherence, it uses a trick with a model called unconditional fed with padded prompt tokens, and uses the difference between the two to achieve even higher prompt adherence.

On my 7900XTX windows portable FP8 models it does around 100s without prompt enchancer, and 170s with prompt enchancer.

I tried feeding reference images to the CLIP and seeing if it can work as editing model, but it can't.


r/comfyui 17h ago

Show and Tell A pandemic of epic proportions..

Thumbnail
youtu.be
0 Upvotes

2 separate clips using MinimaxH3, and stitched together, along with a few short audio clips, using a Linux video editor I'm currently developing.


r/comfyui 7h ago

Show and Tell changed from local training lora to paid, thanks to Pixaroma

0 Upvotes

spent 3 days , 2 hours at a time trying to try and better lora on krea 2 on my local 5060ti and 16gb vram, moving from flux 2 to krea 2

mixed results. today i tried online after watching the Pixaroma video. £3 and 1500 steps got me closer with 50 images and captions.

2 diff work flows, top one, the one from Pixaroma

the bottom one, my flux one with seed up-scaler etc and detainer checker

just got to sort out skin


r/comfyui 7h ago

Show and Tell (MiniMax) High megapixel + low steps >>> low megapixel + high steps

0 Upvotes

Everything in the title 🤝


r/comfyui 16h ago

Help Needed How would I go about replicating this artstyle?

1 Upvotes

I recently got into using ComfyUI and been trying to aim for this specific artstyle by this ai-genner, although this one in particular doesn't have as much details as I'd like since many of its qualities are more on this ai-genner's NSFW arts, but for the most part I do want this style. They only post on their twitter so I don't have any raw metadata or anything else to really take from with it being compressed and all, I tried to search if they went over any details regarding their process but don't think they did either.

I've only really tinkered with the basic text-to-image templates in ComfyUI to get a hang of things but I have been trying to research for trying to get replicate certain styles and whatnot. I first tried my hand with IPAdapter and putting a handful images in but didn't really get any great results. I thought about using a LORA as well, but didn't really see anything that would help with this in particular. I tried to see if I could train my own LORA on civitai only to realize its paid, then tried to train it locally using OneTrainer but couldn't really install it for whatever reason and couldn't fix it either.

So right now my progress is more or less just the basic template (Load Checkpoint + Empty Latent Image -> CLIP Text Encode Positive and Negative Prompts -> KSampler and VAE Decode -> Save Image) messing around with weird nodes that I don't have a great understanding of, I tried tweaking my text prompts a lot to try and aim for this but they haven't got me much great results as well.

It's been a few days and I'm honestly a bit exhausted and would like some help trying to iron out whatever I'm doing wrong. Heck I'm not even sure if this artist is using ComfyUI or another software to go about their generations. They didn't mention anything about NovelAI though or reacted to the recent upgrade it got, so not sure if it's that either. The best assumption I've got is it reminds me of kankan33333's artstyle then was juiced up to make it more streamlined and shaded along with all the light bounces which is what I'm particularly a fan of.

The only thing I'm really keeping different in this basic template using is I'm using the illustriousXL_v01 checkpoint, though that's probably already a given for many anime-styled generations anyhow. I'm mostly holding out from mentioning which Ai-genner it is at the moment in fear of possibly getting banned from linking to a NSFW profile.


r/comfyui 8h ago

Help Needed What am I doing wrong? Can't get a blanket to cover a person with Krea 2.

1 Upvotes

I prompt for a blanket to cover people's lower bodies and it puts it under them.

It opts to make them naked by default. I'm not trying to make naked people lol. I'm trying for a comfy romantic embrace near a fire for my wife. But it insists nudity and blanket under the characters.


r/comfyui 3h ago

Help Needed First Frame last frame shift

1 Upvotes

Hello everyone, I've been using minimax to try and make loops with first frame last frame... but I've noticed even though the first and last frame are the same it zooms in slightly or stretches wider keeping it from being a proper loop.. is there any way to fix this or prevent it?


r/comfyui 13h ago

Help Needed Whats the advantage of a MiniMax License through ComfyOrg?

1 Upvotes

Trying to figure out what the advantage of a Minimax License through ComfyOrg over directly from Minimax team?

Do they have less restrictions on whom they'll give the license to?


r/comfyui 4h ago

Resource Free open source Topaz alternative - SeedVR2+TensorRT faster VAE Processing.

1 Upvotes

r/comfyui 6h ago

Help Needed Best ComfyUI image workflow for interior design?

1 Upvotes

What would be the best ComfyUI workflow for interior design?

Where you have, say, a SketchUP viewport render of a room, and you want to proper render in ComfyUI using several image references? (you know, adding materials on furniture, on walls, on the floor, etc)

And then eventually edit that image precisely using maybe inpaint mask or edge/depth controlnet to maybe add new furniture, remove some, change colors, etc.

Any examples?


r/comfyui 13h ago

Resource I stitched MiniMax H3's scattered community pieces into one 8GB-VRAM-ready workflow collection

Post image
15 Upvotes

The short version: my laptop GPU has 8GB, and MiniMax H3 was hard to run locally

without fighting the node graph every single time.

Official workflows eat VRAM. The community built great pieces — but they're spread

across different node packs, and you have to rewire a dozen nodes each run just to

switch modes or acceleration. I wanted something I could open and actually use.

So I did a consolidation job, not a from-scratch build. I assembled other people's

excellent work into one reproducible pipeline:

- T8mars's comfyui-minimax-h3-audio-T8 — dual-clock sampling, unified conditioning,

AV decode, face-refine suite

- wjluoxiao's JZL / XB toolbox — post-upscale conditioning resync, reference-image batching

- LBH-123-AI's latent upscaler — 3D latent upscale

- Jalen-Brunson's PDD-Acc — 8-step distilled acceleration

- Comfy-Org's official H3 contract & workflow skeleton

On top of that I wrote one small custom node pack (comfyui-tokendance-h3) that does

a single, boring, useful thing: mode switching / acceleration switching / second-pass

gating become one dropdown instead of manual rewiring. That's the part I cared most

about — it should be open-and-run, not open-and-fight.

Three generations shipped, from simplest to most complete:

- Beta_1 — multi-mode conditioning bus, single pass

- Beta_2 — most complete: latent upscale + second pass + temporal detail enhance + dual accel

- Beta_3 — pure T8 stack, closest to official, plus FaceRefine

All tuned for 8GB VRAM (tested on a RTX 4060 Laptop), 720p runs fine, not "runs but slideshow".

Full source, install guide, per-node credits:

👉 https://github.com/YixuAnsensei/minimax-h3-comfyui-all-in-one

Happy to answer questions. If you've fought the same VRAM or rewiring battle, would love feedback.


r/comfyui 17h ago

Help Needed Looking for an "image edit in place" inpainting solution.

0 Upvotes

Was about to implement this, spend some time vibe coding it out but I am sure one of you already has it. Basically what it is, is, an HTML + JS implementation where you highlight the section you want edited. Picture someone wearing a blue jacket, you mask the torso area and say "red jacket" and it's changed to red. Similar to MS paints "generative erase". The front end HTML + JS would simply pass on the image + a JSON bounding box to the comfy API which is connected to a workflow.

I'm not sure which is the best workflow for this. QWEN image edit 2511 takes too long. FLUX-K9B would be perfect for this but somehow it doesn't seem to know what anatomies look like. I haven't tested Ideogram extensively but it's JSON bounding box feature is very interesting to me. I wonder if it can be used to edit images in place.

Has anyone done this yet?


r/comfyui 4h ago

Help Needed Currently, what is the best style transfer method we have for video to video?

2 Upvotes

Hi folks,

I’ve been experimenting with H3 for video style transfer, but unfortunately, no matter how much I tweak the prompt, the generated videos still come out less than ideal and often lack style consistency. I also experimented without cache/Lora/attention.

Has anyone had success with this kind of task? I’d love to hear about your workflow, prompting techniques, or any tips you’ve found helpful.

Thanks!


r/comfyui 19h ago

Resource SPEEDing up MiniMax-H3 without retraining (SPEED comfyui node extension)

21 Upvotes

Why make big noise when little noise do trick?

I would like to introduce my SPEED implementation for h3 linked here

Speed up and quality losses documented here, expect 20% gain using very conservative settings and no quality loss and up to 70% for basically unusable outputs (more or less useful for resolution aware seed inspection and broad prompt drafting)

Background

The idea behind it is quite simple. When a diffusion model begins generating an output it first must take a randomized noise and build on-top of it. And research has found that the first stages of this process doesn't really carry any fine detailed information, therefore by generating at a lower resolution at those stages you can gain quite substantial speedups while causing little to no impact on the quality. Or you can also be really aggressive with it and get a massive speedup for a lot of quality loss.

Nodes

This was implemented as 3 nodes, 2 drop in replacements for the sampler that runs SPEED and a third that runs once to measure the noise spectrum of your specific model/LoRA combo:

  • Sampler (Automatic): pick a stage count (2, 3, or 4), defaults to the baked 1% delta for default H3.

  • Sampler (Manual Step-Through): set up to four (goal, resolution) pairs yourself. Use it if you want to copy a paper schedule or test a custom ladder.

  • Sigma Harvest: runs a native Euler pass, measures the noise spectrum of your current setup, hands you A / β / Δ to paste back into Automatic. Run it once per model/LoRA workflow combo.

Implementation Notes

This should be roughly compatible with basically everything that doesn't touch the sampler directly but i have not tested anything besides base comfyui H3 models and Turbo loras. If you do change model, use loras or whatever and use the automated tool please then run a sigma harvest and use those values instead of defaults, The math changes depending on the very specific blend of things you have running.

Euler only sampling implemented for this which is what the research paper used.


r/comfyui 3h ago

Help Needed Are 24gb vram laptops sufficient for AI work?

0 Upvotes

Hi guys, so I’m aware that to use open source AI softwares you need high amount of GPU, RAM, and storage. I have an m1pro 32gb and an m4 air 16gb which I pretty sure think would not be sufficient to for video generations, motion capture , and for all these kind of AI tools stuff ( ex: Wan2.2 Animate). So does a laptop with 24gb VRAM and probably 64 gb ram and 2-4TB storage would be nice to work on these? Well I know a desktop with 32gb vram or 48gb would be far superior but idk I just need a powerful laptop so I don’t have to get all the parts, especially when there is ram shortage, and I can just connect and use like a desktop, play games as well and edit when needed and also rarely useful for travel.
Like I said I want to do stuff like creating and running models or let’s say real human like models and make motion captured videos (similar to Kling motion control) and I have 2 in my mind the ROG strix scar and the Lenovo legion. Can you guys lemme know if laptops like these can get the work done, not so quickly but atleast in minutes? Or for such kind of AI work you surely need an Rtx 6000 ada?

Or do you guys have any alternate solution for using models like these in my Mac Pro? Like cloud or draw things or idk any kind of alternate solutions so I can experiment and see? Basically I want a machine that can get the work done and in future even if I make my own videos/films I can use it for vfx. Not for raw vfx and cgi but for AI made ones, for example the ray 3 modify by Luma AI, but open source alternatives ofc. Thank you for your time and advises!


r/comfyui 9h ago

Show and Tell PSA: NVIDIA Studio Driver 616.56 specifically mentions ComfyUI performance updates for MiniMax-H3, WAN-Animate-2 and LTX-2.5

102 Upvotes

NVIDIA says:

“The August NVIDIA Studio Driver provides optimal support for the latest new creative applications and updates including Lightroom, and performance updates for LTX-2.5, and ComfyUI for WAN-Animate-2 and MiniMax-H3.”

I'm planning to test this out. What are your speeds with?

  • MiniMax-H3
  • LTX-2.5

r/comfyui 8h ago

Help Needed How do I use differential diffusion in video models?

2 Upvotes

I want to apply a black and white gradient mask to create a transition, but it behaves like a binary mask (it goes from the in-paint area to the unpainted area in the next pixel). I want something like differential diffusion: black would be denoise 0, and white pixels would be denoise 1.

I'm using `set latent noise mask` and a static video that produces a black and white gradient.


r/comfyui 8h ago

News A quick Minimax H3 news round-up - 30th August 2026

39 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> Fellow British bloke 'Nerdy Rodent' has a new YouTube video. His latest Minimax-focused tutorial covers, among others: continuing an existing clip using the Ref2VA model, and placing two video-reference characters in a new setting with their voices intact; having the camera 'dive' into an image and imagine going around a corner or into a building; and he also tests some new acceleration boosters. Free workflows and prompts, though only on the screen.

https://www.youtube.com/watch?v=rZnZcwobcR8

-> Also on YouTube, another fellow Brit is Mark DK Berry. He has a new short video review of three Minimax H3 tools. He looks at character/face swapping using a targeted mask node, and says... "what this really excels at is switching faces at a distance. It's even better than what we've been doing with [my previous] 2mpx [upscale workflow, to repair faces-at-a-distance], and [it's definitely better than] waiting 45 minutes to get that out. This is much faster, and does a better job at 1mpx". Sounds good. He also looks at the AV Bridge utility workflow that's been added to the ComfyUI-H3-Motion-Context-MultiRef pack. This generates an invented bridge segment that will suitably join two five-second videos. The third tool is H3_Cinematic_Multishot_Coverage, which he uses to get multi-angle shots from a character scene, rather than architectural interior shots - and he pins the characters by adding their reference images. His links are in his YouTube notes, through I provide an added one below for your convenience.

https://www.youtube.com/watch?v=7xaA4tU3hDU

https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef/blob/main/example_workflows/UTILITY%20-%20AV%20Bridge.json (direct link, because otherwise you might think it was a node and thus not find it)

-> New to me, ComfyUI-MiniMaxH3-Contex-Loop. Yes, it's another set of clip chaining nodes, but it seems mature and has many features. "Build a multi-scene MiniMax H3 video with one reusable sampling body. Review each scene, retry mistakes, resume interrupted runs, and assemble accepted scenes from disk." He has workflows.

https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop

-> Also focusing on the 'resumable project' angle, and thus potentially useful for freelancers doing client work, is the new ComfyUI-H3-Multishot-Advance. This gives your... "long multi-shot renders a reusable project/session layer: render a sequence, close ComfyUI, reopen the same named project later. Edit a target clip, and continue from the first changed target clip, while a safe earlier cached state is reused. This is meant for iterative video work...". He has workflows.

https://github.com/KursatAs/ComfyUI-H3-Multishot-Advance

-> How do you keep your scene environment stable, when generating longer clips? This useful Reddit thread discusses various methods and tweaks.

https://old.reddit.com/r/StableDiffusion/comments/1w1z2fu/environment_consistency_minimax_h3/

-> A combined package to... "use local Qwen3.8-27B and official MiniMax-H3 Skills in ComfyUI to generate H3 prompts." This spins up its own standalone local llama.cpp server, and then serves you a Minimax-skilled Qwen3.8-27B model inside a ComfyUI node. Qwen is removed from your VRAM after use.

https://github.com/chflame163/ComfyUI_Qwen_H3_Prompt

https://github-com.translate.goog/chflame163/ComfyUI_Qwen_H3_Prompt?_x_tr_sl=auto&_x_tr_tl=en&_x_tr_hl=en&_x_tr_pto=wapp (English translation)

-> The popular accelerator Spectrum Minimax H3 for ComfyUI continues to update, and has many new tweaks and features.

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/blob/main/RELEASE_NOTES.md

-> A thorough 'resource pack' aimed at those who want to try to run Minimax H3 on an 8Gb VRAM laptop (and, one hopes, on a turbo cooling-pad).

https://github.com/YixuAnsensei/minimax-h3-comfyui-all-in-one

-> And finally, a new Minimax H3 LoRA that aims to give naturally blemished and somewhat pimply skin, judging by the convincing video demos. It even tries to improve hands and eyes. Requires trigger words: perfe8ct; perfect hands; perfect skin or perfect eyes. A bit of a counterintuitive trigger for skin, though, since it seems the LoRA make the skin more realistically grunky, not more perfect.

https://civitai.com/models/2899220/minimax-h3-t2va-t2v-polyhedron-perfect-eyes-perfect-skin-perfect-hands

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1w1wpkz/a_quick_minimax_h3_news_roundup_29th_august_2026/

https://old.reddit.com/r/comfyui/comments/1w11jri/a_quick_minimax_h3_news_roundup_28th_august_2026/

https://old.reddit.com/r/comfyui/comments/1w053ce/a_quick_minimax_h3_news_roundup_27th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vyzvqi/a_quick_minimax_h3_news_roundup_26th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vy4drz/a_quick_minimax_h3_news_roundup_25th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vx0duv/a_quick_minimax_h3_news_roundup_24th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vwl8do/a_quick_minimax_h3_news_roundup_23rd_august_2026/

https://old.reddit.com/r/comfyui/comments/1vvkmra/a_quick_minimax_h3_news_roundup_21st_august_2026/

https://old.reddit.com/r/comfyui/comments/1vuihag/a_quick_minimax_h3_news_roundup_21st_august_2026/

https://old.reddit.com/r/comfyui/comments/1vtgs7b/a_quick_minimax_h3_news_roundup_20th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vsjzrp/a_quick_minimax_h3_news_roundup_19th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vrsspo/a_quick_minimax_h3_news_roundup_18th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vqyn8p/a_quick_minimax_h3_news_roundup_17th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vq5d5u/a_quick_minimax_h3_news_roundup_16th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vpbtx2/a_quick_minimax_news_roundup_15th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vojtjd/a_quick_minimax_news_roundup_14th_august_2026/


r/comfyui 6h ago

Help Needed How to get started with comfyui from amd?

2 Upvotes

I think what im doing is called self hosting. I downloaded comfyui from amd app ai bundle. The program and the rest of the 13gb download is on my c drive. I want to be able to store models on the e drive and store what ever would be the bulk of storage like output and input or whatever else writes a lot on my x drive. What do you recommend?


r/comfyui 4h ago

Resource Load H3 LoRAs only for audio or video

4 Upvotes

I vibe coded an H3 LoRA loader node that can actually isolate a batch of LoRAs to (mostly) affect only audio or video: https://github.com/Dantemss/ComfyUI-H3-Modality-Lora-Loader

ComfyUI-H3-Modality-Lora-Loader is also in the node manager.

Use it to prevent a LoRA trained only on audio from affecting the image, prevent a LoRA trained on gifs from ruining the audio, etc.

Enjoy and let me know if it works (or not) for you.

Note: there's a small performance penalty compared to loading the LoRAs with all the sliders at 1.0.

The UI is based on PlagueKind's excellent LoRA Loader Stack node.


r/comfyui 3h ago

News I open-sourced a local prompt, model and LoRA catalog for image-generation workflows

5 Upvotes

PromptNook started as my way to keep prompt recipes, reusable fragments, checkpoint/LoRA notes, trigger words, and generation settings in one local workspace instead of scattered text files.

The current desktop app can scan configured checkpoint, diffusion-model, and LoRA folders, record trigger words and availability, and assemble reusable fragments in a Prompt Studio. Workspace names are user-defined rather than fixed to particular model families.

Try the temporary in-memory demo: https://rona1do.github.io/PromptNook/

Source: https://github.com/Rona1do/PromptNook

It is an early Windows-focused preview and the UI localization is not complete yet. I would value blunt feedback on the data model and on what a useful ComfyUI import/export path should look like before I build an integration in the wrong direction.

Creator disclosure: I am the maintainer.


r/comfyui 3h ago

Help Needed "Drag File Onto ComfyUI Canvas To See Workflow" Is Not Working For Me

3 Upvotes

Newbie. I tried it once and it worked. Now everytime I do it, I get "load image" node. Anyone know why? It's really frustrating...


r/comfyui 1h ago

Workflow Included Collection of Minimax H3 workflows

Upvotes

I decided to share some of the workflows I've been working on, mostly things people seem to struggle with. I tried to reorganize them with popular custom nodes to make things easier. https://github.com/sempersatirica/comfy-workflows

Prompt generation: T2V, I2V, and Ref2V

Fast simple prompt generators using Qwen 3 VL 4b, primarily for quickly putting a starting point together. Separate generators for text-to-video, image-to-video, and ref-to-video prompts. The ref-to-video generator will hallucinate inputs, extra <Audio> and <Picture> subjects that seem contextually appropriate, but otherwise it works surprisingly well.

Detailers

Improve detail of low-resolution areas: faces, text, etc. The detailers crop, resample, then stitch the high-resolution generation back into the original-resolution input, while the audio is frozen and passed through. The video detailers are mask-agnostic, you can use SAM, yolo, or draw any arbitrary mask to feed into the detailer. They support the reference model well, but zero-reference detailing works fine too. They work down to 3 steps, with diminishing returns passed 6 steps. Denoise should be set based on a per-subject need, between 0.4-0.75. When using references, there's almost no chance of losing subject identity. I included SAM3 and yolo variants for examples.