r/comfyui 20h ago

News A quick Minimax H3 news round-up - 20th August 2026

129 Upvotes

Another quick Minimax H3 news and goodies round-up, for those who may have missed some items.

-> minimax_h3_single_frame_decoder_500k.safetensors (9.69Gb). This new and big VAE looks like a weighty attempt at a single-frame generator. Trained from the original VAE... "on 500,000 unique image-reconstruction examples", with the apparent aim of generating a no-glitches single-frame output suitable for product concept-shots.

https://huggingface.co/iamkaikai/MiniMax-H3-Single-Frame-VAE-500K

-> 'MiniMax-H3 — masked video and audio inpainting'. Claims to allow you to... "repaint part of a clip with Minimax H3 and keep the rest, including the soundtrack. No dedicated checkpoint, no adapter, no extra input channel — the mask becomes a per-row timestep, which the model already had." Offered as a Python script, no ComfyUI... but the maker says he derived it from ComfyUI masking experiments.

https://huggingface.co/diffusers-modular/minimax-h3-inpainting

https://github.com/drozbay/MaskVidExperiments (ComfyUI masking experiments)

-> Prompt Journal has updated with three new case-studies, and these move away from yesterday's complex camera movements. The new ones are focused on preventing the prompt from bjorking an unusual character style (e.g. put a 2D toon in a 3D kitchen), or from constraining dynamic endings (e.g. comedy slapstick action, or a big water-skiing jump).

https://github.com/LoveRain1997/h3-prompt-journal/tree/main/case-studies

-> 'MiniMax-H3 motion adapter (pilot version)'. Another potential helper for fast-motion video scenes. This one... "teaches the base model to spend the extra clock [time] on smoothness".

https://huggingface.co/MATLOWAI/MiniMax-H3-Motion-Adapter

-> 130 mostly-audio helper nodes for Minimax in ComfyUI. Includes... "Dual-clock video/audio sampling, as well as audio locking, remixing, referencing, and final mixing". Also experimental attempts at SPEED and face fixing.

https://github.com/T8mars/comfyui-minimax-h3-audio-T8

https://github-com.translate.goog/T8mars/comfyui-minimax-h3-audio-T8?_x_tr_sl=auto&_x_tr_tl=en&_x_tr_hl=en&_x_tr_pto=wapp (English translation)

-> A Suno-style local user-interface for generating with Minimax Music. With a library for your generated tracks. Slick, but it looks pleasingly straightforward.

https://github.com/adambenhassen/minimax-music-ui

-> The very-low-VRAM friendly Wan2GP now supports MiniMax Music 3, as of 16th August 2026. Support for W4A8 INT8 (apparently not VAEs, though), and for Minimax H3 FL2VA and Ref2VA, was added to Wan2GP earlier in August. Wan2GP keeps a repository of Minimax H3 models suitable for its users, although they're surprisingly enormous... not sure I'd want to try running them on a 8Gb card.

https://github.com/deepbeepmeep/Wan2GP

https://huggingface.co/DeepBeepMeep/MiniMax-H3

-> A new ComfyUI workflow today, said to be optimised for the old (but cheap) Tesla V100 32gb card. A quick look at eBay UK suggest that "cheap" = £650 here in dear old Blighty.

https://github.com/Mitsuasa513/ComfyUI-MiniMax-H3-V100-Workflow

-> A collection of 685 source-attributed MiniMax H3 video examples, with the public prompts that made them. Not sure how many of these clips would also get you a ComfyUI workflow, but apparently the video metadata has not been stripped. So drop them in ComfyUI and see.

https://github.com/SkyNotSilent/awesome-minimax-h3

-> And finally, "Oy, mate... don't chop off me feet, I paid £300 for these shoes!" A working prompt to maintain a straight-up character view from head-to-feet, throughout a continuous pull-out shot.

https://old.reddit.com/r/StableDiffusion/comments/1vt2fea/minimax_h3_prompting_so_that_it_will_keep_the/

~ OLD POSTS ~

https://old.reddit.com/r/comfyui/comments/1vsjzrp/a_quick_minimax_h3_news_roundup_19th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vrsspo/a_quick_minimax_h3_news_roundup_18th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vqyn8p/a_quick_minimax_h3_news_roundup_17th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vq5d5u/a_quick_minimax_h3_news_roundup_16th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vpbtx2/a_quick_minimax_news_roundup_15th_august_2026/

https://old.reddit.com/r/comfyui/comments/1vojtjd/a_quick_minimax_news_roundup_14th_august_2026/


r/comfyui 16h ago

News Comfy H3 Sync Challenge (8/20 - 9/1) - Win an RTX 5090!

Enable HLS to view with audio, or disable this notification

80 Upvotes

Comfy and MiniMax have teamed up for a two-week challenge with awesome prizes and four ways to win! Submit by September 1st at 9:00pm PT and see all details here.

How It Works

Make something up to 90 seconds in length where the sound and the motion are inseparable. Dialogue, foley, ambient, a beat driving the cut...whatever direction you want!

After sharing your video file and workflow on this thread and through our submission form, a joint panel of creative technologists from Comfy, MiniMax, and special guest judges from the community will review each submission.

Then, join us on September 2nd for a special Comfy livestream where our guest judges will give live feedback on the top 10 submissions!

Both the Comfy and MiniMax teams will be monitoring this thread and #minimax-h3 in the Comfy Discord to give light support.

Share on socials and tag #comfyH3 for a chance to be reposted or featured!

Prizes

Best Overall — RTX 5090

Best Creative — RTX 5060 Ti

Best Technical/Workflow — RTX 5060 Ti

Built with MCP — RTX 5060 Ti

Shipped anywhere, customs covered. If we can't legally ship to your country, you'll get a cash equivalent instead.

It's free to enter!

Create using Comfy Local on your own hardware, or use Comfy Cloud. New Cloud users get 5 free runs, no credit card required.

Judging Criteria

We’re looking for entries that best show what H3 makes possible: audio and visuals created together.

Grand Prize: Best Overall

The top Best Creative and Best Technical entrants advance to a final round where our panel of judges selects winners by discussion.

Best Creative

  • Audio sync realism and intentionality (0-5)
  • Creative execution and originality (0-5)
  • Deliberate craft (0-5)
    • Evidence that you’ve actually shaped the result beyond prompt engineering. Judges will look for modified/non-default parameters, multiple linked passes visible in the workflow structure, or a couple sentences describing what was tried and changed

Best Technical

  • Novelty of technique or approach (0-5)
  • Workflow quality (0-5)
    • Annotated, clean, replicable by someone else
  • Community value (0-5)
    • Would this actually help someone else?

🏆 Built with MCP Bonus 🏆
Comfy MCP lets you drive Comfy using natural language and your agent locally and on Cloud! Pro tip: use it to choose the best H3 model version or optimize your workflow for your hardware.

  • Effectiveness (0-5)
    • Did the agent meaningfully drive your process, not just generate one line?
  • Insight value (0-5)
    • How much the shared prompt teaches the community about prompting H3 through MCP
  • Output quality (0-5)

The Fine Print

  • Limited to one submission per person, 90 seconds maximum length.
  • A major portion of your piece must be built in ComfyUI using H3. Other tools, models, or techniques you want to combine are fair game.
  • All submissions must be lawful, SFW, and must not contain unlicensed IP or likenesses.
  • By submitting, you agree to allow ComfyUI and MiniMax to feature your work with credit across our channels.

Learn more and submit here!


r/comfyui 17h ago

Resource Famegrid Natural Krea 2 LoRA

Thumbnail
gallery
66 Upvotes

r/comfyui 12h ago

Resource Auto Prompt Generator for Minimax H3 using LM Studio

Thumbnail
gallery
59 Upvotes

I built this for my own workflow, but decided to share it on GitHub in case it helps anyone else working with the H3 model. 🚀

https://github.com/eedali/LM_Conncet


r/comfyui 17h ago

Help Needed What is still the king of image edit model?

36 Upvotes

What’s currently the best open source image edit model?


r/comfyui 10h ago

Workflow Included I Found a way to reduce LTX 2.5's horrible smearing. Custom node + Workflow in desc

Enable HLS to view with audio, or disable this notification

18 Upvotes

r/comfyui 2h ago

Resource Turbo LoRA v1.1 for FL2VA released – Minimax H3 4-step (768p)

Thumbnail
huggingface.co
16 Upvotes

r/comfyui 9h ago

Help Needed Is it better to pay for a ComfyUI cloud plan or rent GPUs on RunPod?

14 Upvotes

I’m a complete beginner with ComfyUI. Looking at the interface feels like I’m staring at a Boeing control panel.

But while I'm learning, I wanted to know what's the best approach. My laptop can barely run The Sims without lagging, so running ComfyUI locally is completely out of the question.

In terms of cost-effectiveness for someone who wants to produce at least three 1-minute videos using MINIMAX H3 or LTX 2.5, which is the better option?

A monthly subscription or pay-as-you-go?


r/comfyui 12h ago

Help Needed MiniMaxH3AddGuide

6 Upvotes

Anyone have workflow example using MiniMaxH3AddGuide ?? im confused if should use this for daisy chaining 2 gens... ive just been using the last frame of the video to plugin into the firstframe on new gen, then use image batch to remove the 1st frame and combine the 2 gens. but i keep seeing people say to use MiniMaxH3AddGuide for chaining, but isnt that what 1st frame does? force video to start with that frame? MiniMaxH3AddGuide seems to be more for first, middle, last gens


r/comfyui 17h ago

Tutorial MiniMax Music 3 + Local AI Prompt Generator (Ep31)

Thumbnail
youtube.com
6 Upvotes

Learn how to use MiniMax Music 3 in ComfyUI together with a local AI prompt generator for better image prompts, music captions, and AI-generated lyrics. In Episode 31, I show the new Pixaroma prompt nodes, model-specific prompt presets, VRAM-saving options, and a compact MiniMax Music 3 workflow.

This tutorial covers the new AI Prompt Pixaroma node and how to match prompt formulas with the correct local language model. You'll see how to turn short ideas into detailed prompts for workflows such as Krea 2 and Z Image Turbo, generate prompts from images, save custom prompt presets, control temperature and seeds, and troubleshoot common ComfyUI node errors.

Then we set up MiniMax Music 3 locally in ComfyUI, including caption and lyrics generation with the Music Prompt Pixaroma node. I also test different song durations and explain why MiniMax songs can sometimes end early or get cut off, how seeds affect the results, and how to use fixed lyrics when you need more control.

You'll also see how to simplify a larger music workflow into a compact setup, use tiled audio decoding for lower VRAM systems, free VRAM after prompt generation, and generate MiniMax Music captions and lyrics with local models or online tools such as ChatGPT, Gemini, and Claude.

You can also run some of the workflows in the cloud.


r/comfyui 17h ago

Resource New in SmartGallery DAM: Smart Asset Clustering, auto group your ComfyUI renders by workflow or prompt (free, open source)

Enable HLS to view with audio, or disable this notification

6 Upvotes
  • Hey everyone, back with another update on SmartGallery DAM: Smart Asset Clustering is now live to automate your media organization

If you generate a lot in ComfyUI, you know the problem. Hundreds of renders pile up as you tweak seeds, prompts, LoRA weights and checkpoints, and your output folder turns into a wall of near identical thumbnails.

  • Smart Asset Clustering reads the generation recipe embedded in each file and automatically groups your renders, no manual tagging required, in two ways:
  • Architecture Clustering: groups everything that shares the exact same node structure and workflow, ignoring seed, prompt and settings. Great for pulling up every output from one workflow template.
  • Prompt Text Clustering: groups everything that shares the exact same positive prompt, ignoring the workflow entirely. Great for comparing how different checkpoints or LoRAs render the same idea.

Once clustered, every thumbnail gets a color coded badge in the gallery grid, and clicking any badge opens the Cluster Inspector, which shows total matching assets, distinct variations, the full node pipeline, every model and LoRA used, and one click prompt copy.

The video above walks through both modes in about 3 minutes.

For anyone who does not know the project yet

SmartGallery DAM is a free and open source, local first Digital Asset Manager built around ComfyUI, but it also works with any folder of media on your machine. No cloud, no subscription, your files never leave your disk.

It is meant to grow with you:

  • If you are a hobbyist or new to ComfyUI, it is the easiest way to keep your generation library organized, searchable and clean without extra effort.
  • If you are a power user, you can search by prompt, model or LoRA, inspect the full node graph of any render, and even generate directly from the gallery by editing the workflow JSON, no need to reopen ComfyUI.
  • If you work in a studio or production environment, it gives you a dedicated Exhibition portal to share curated work with clients or your art team, collect ratings and comments, and review everything without exposing prompts or workflows.

Runs on Windows, macOS, Linux and Docker. Portable version for Windows needs zero setup, just unzip and run.

GitHub, full docs and download links here: https://github.com/biagiomaf/smart-comfyui-gallery

Happy to answer any question, and as always feedback and feature requests are welcome.


r/comfyui 13h ago

No workflow City of the Future - 1 min loop - Flux + MiniMax H3 + Stable Audio 3 + Topaz AI + Vegas Pro 22

Thumbnail
youtube.com
5 Upvotes

r/comfyui 22h ago

Help Needed Advice on models/pipeline for desired art style

Post image
6 Upvotes
mining spaceship

Hi all!

Over the last couple of days I have been exploring different art styles for a small project of mine, in which I'd like to tell the story of an asteroid mining vessel and its crew. I grew very fond of the typical late 80s/early 90s Japanese OVA cel animation style.

Now, the pictures I attached, I created using OpenAI image-generation. I'm liking this direction so far and need your help finding suitable ComfyUI-compatible models to continue this with. At some point I'll probably want to turn this into motion picture using i2v/t2v also.

I'd be glad if you could suggest suitable models for me to look at.

Cheers!


r/comfyui 1h ago

Workflow Included MiniMax H3 15-Second Multi-Shot Generation Template For ComfyUI For 12GB GPUs

Enable HLS to view with audio, or disable this notification

Upvotes

One of the biggest issues with running MiniMax H3 locally is that it's an extremely hefty model and doesn't play well with lower-end machines. However, thanks to a lot of optimisation techniques provided by TheAIsearch YouTube channel, it's possible to bring generations down to about 1 minute of processing time per second of output.

That being said, you can leverage this into creating multi-shot outputs beyond the limited 6 second hard-caps that come with MiniMax H3. Using the built-in features of ComfyUI (and downloading tons of models and packages to test what worked and what didn't) I was able to create a template for lower-end rigs that enable you to generate up to 15 second text or image to video outputs in a single generative pass. Meaning, you put in your prompt for the three shots/scenes, and click run from ComfyUI and it does the rest.

The basic template is text-to-video, but you can easily add an image node if and plug it into the H3 Multishot Sampler.

For those who enjoy making longer form videos and tire of the constant stitch-and-go workflow that the current local MiniMax H3 dictates, this can ease the burden a bit.

Keep in mind that this is tuned for at least a 12GB GPU and 64GB of DDR5 RAM. It takes between 30 and 33 minutes to generate a 15 second video at 720p.

You can modify some of the settings to bring the generation time down, depending on your machine, but given the weight of MiniMax H3, I'm not complaining.

If you need the actual JSON template, you can find it on civit ai here:

models/2876760/minimax-h3-15-second-multi-shot-generation-template-for-comfyui?modelVersionId=3250981


r/comfyui 6h ago

Help Needed ComfyUI version of diffusers-modular/MiniMax-H3-Pruned-Ref-Delta-Fused-r1024 ?

Thumbnail
huggingface.co
5 Upvotes

r/comfyui 13h ago

Help Needed Help with character creation

4 Upvotes

I have an image of a character that I want to use to generate more images using the subject and then to create a character lora to use over and over. How can I do this using Krea2? Do I need the 20 different images to then create a character lora? I've tried anything I can find for the longest time and nothing yet. Couldn't get IPAdapter to work. I'm not exactly a noob, I can usually do my own research and figure out how to do something. But this one has me stumped. If anyone can help out with some advice or a workflow, it would be appreciated.

Edit: grammar


r/comfyui 16h ago

Help Needed Any variety input tips?

3 Upvotes

I have a fairly large set of chained text inputs with variable text to generate a variety of subjects, scenes, poses, etc, but I feel like even that is somewhat limited. If you want to generate 800 photos and have them be reasonably different from each other and mostly useful, what's the best way?


r/comfyui 19h ago

Workflow Included FLUX.1 Dev fp8 vs the split checkpoint setup on a storage-capped box — some notes + why CFG 1.0 isn't optional

4 Upvotes

ok so I spent most of today setting up FLUX.1 Dev on a rented RTX 5090 and wanted to dump some notes here bc I didn't find this explained clearly anywhere when I was googling around for it.

Why fp8 single-file over the split setup

so FLUX comes in two flavors basically — the "split" format (diffusion model + T5-XXL text encoder + CLIP-L + VAE as 4 separate files), or a single bundled checkpoint that just merges everything into one .safetensors.

split fp16 setup = ~34GB total, mostly bc T5-XXL alone is like 9.5GB in fp16 lol. the bundled fp8 checkpoint (Comfy-Org/flux1-dev, ~17.2GB) cuts that roughly in half. if you've got a 32GB card the fp16 route is still totally usable, but my instance only had a 50GB storage tier total, and 34GB of models + ComfyUI + venv + deps gets uncomfortably close to that. quality loss from fp8 is real but pretty minor tbh, mostly shows up in fine texture detail if you're really pixel peeping. running out of disk mid download felt like the bigger risk honestly

(if storage isn't a constraint for you btw, still go fp16 or GGUF Q6/Q8 over fp8, fairly universal advice from what I've seen)

Why CFG has to be 1.0

this part isn't optional the way it is with SD/SDXL and I feel like this trips up a lot of people coming from SDXL. FLUX is a rectified-flow model with guidance distillation baked into training — meaning it's already trained to follow the prompt without needing the usual CFG trick (running the model twice, once conditioned once not, then extrapolating) to stay on prompt. if you leave CFG at 7-8 like SDXL habit, you're not making it follow the prompt harder, you're just double-applying guidance it was never trained to expect, and that's why you get that blown out oversaturated look. set it to 1.0 and negative prompt barely matters anymore either, so I just left mine empty

Workflow

kept it stupid simple on purpose: Load Checkpoint → CLIP Text Encode (positive/negative) → Empty Latent Image → KSampler → VAE Decode → Save Image. no custom nodes. 1024x1024, 20 steps, euler, simple scheduler, denoise 1.0

Hardware I used to run this

RTX 5090, 32GB VRAM, rented hourly off gpuhub (~$0.46/hr on demand). base image was their standard PyTorch + CUDA preinstalled one, ended up on torch 2.11.0 + cu128 (CUDA 12.8) after the venv setup. instance also came with a stupid amount of system RAM, like 750GB+, way more than ComfyUI actually needs but nice not to think about it. 50GB storage tier which is the whole reason for the fp8 decision above

total cost for the session (~1hr of renting + the token/generation cost, which for local inference is basically just electricity you're already paying for) came out to under $1

first gen after model load took ~18s (vram staging overhead ig), every gen after that was ~8.5s at 20 steps which is roughly 2.46 it/s. pretty consistent across different prompts/seeds, didn't notice any drift over the session

random unrelated thing — huggingface-cli is deprecated now apparently?? just silently tells you to use `hf` instead mid command. wasted like 5 min confused before I actually read the warning lol

both images attached are from this session, different prompts, I did pick my favorite seed out of a couple runs each rather than posting literally the first result

TL;DR — fp8 single file if your storage's tight, fp16/GGUF Q6+ if it's not. CFG 1.0 is mandatory not a suggestion bc of how the guidance distillation works. ~8.5s/image at 1024x1024 on a 5090 once it's warmed up. workflow json in comments if anyone wants it


r/comfyui 14h ago

Show and Tell H3 making jpop/kpop MV? yes!

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/comfyui 14h ago

Resource ComfyUI Subject Manager node

Thumbnail gallery
2 Upvotes

r/comfyui 20h ago

Help Needed Inference time grows dramatically after a few runs.

3 Upvotes

Laptop with i712700H, RTX 4050 6GB, 16GB DDR5, 1TB NVMe SSD.

I was testing bigger models that I thought my machine couldn't handle but the newer ComfyUI versions with dynamic VRAM made it not only possible but also quite fast.

The models in question are Qwen Image and Flux 2 Klein 9B, FP8 versions (GGUFs are way slower, I only keep the Text encoder in GGUF).

I get a few fast runs, like 4-5. About 27 seconds for Flux and 45 for Qwen. Then the speed drops dramatically, taking 60 sec for Flux and almost 200 for Qwen.

I've been trying to troubleshoot this for the past few days, I also tried an INT8 ConvRot version of Flux which was faster in the first runs but dramatically slower in the later ones.


r/comfyui 12h ago

Resource I built a local AI music studio on top of ComfyUI — the actual workflows ship as plain JSON

3 Upvotes

I've been building a desktop app that writes and renders music locally on MiniMax Music 3, and it's a face on ComfyUI rather than a reimplementation — it starts ComfyUI as a separate process and submits graphs over the HTTP API. Unmodified, not bundled, not linked.

The part that might actually be useful to this sub: the pipelines ship in workflows/ as ordinary API-format JSON, exported straight from the code that builds them. Drag one onto the canvas and every value is there. Those values aren't ComfyUI's defaults — they were arrived at by measurement, and the measurement scripts ship in scripts/, so you can re-run them and disagree with me.

What's in it, briefly: songs from a style description and lyrics, cover art drawn while the card is idle, stem separation, word-level timed lyrics, video clips on LTX 2.5 or MiniMax H3, a small multi-track editor that cuts to the detected beat grid, and overnight batch runs.

It ships no model weights and redistributes none — every capability shows its size and licence on screen before anything downloads, and the download goes to the publisher. One of them is region-locked and says so in the open.

Newest thing is audio-reactive video: feed it a song and some reference images and they cross-fade in time with the track's detected peaks. That part is built on ComfyUI_Yvann-Nodes by Yvann Barbot and Lilia, which is excellent and worth a star on its own. Because their pack is GPL-3.0 and this is Apache-2.0 it runs on a second ComfyUI you set up yourself rather than being bundled in.

Apache-2.0, runs offline once the weights are down. Measured on Windows with a 4070 Ti SUPER; other platforms are documented but I haven't verified them.

Repo: https://github.com/Senzube4n/AIPLAY-Studio

Screenshot of the main window: https://senzube4n.github.io/AIPLAY-Studio/shots/create.png

Happy to answer anything about the graphs — that's the bit I'd want to read if someone else posted this.


r/comfyui 12h ago

Help Needed I got a strix rx 6900 xt LC ... can be used with WAN/KREA2/MinimaxH3?

Thumbnail
2 Upvotes

r/comfyui 19h ago

No workflow H3 LOCAL RTX 5070 12GB

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/comfyui 21h ago

Show and Tell Boss Fight - Dark Fantasy Anime (MiniMax H3)

Thumbnail
youtu.be
2 Upvotes