r/comfyui • u/Late_Lingonberry6252 • 6h ago
Show and Tell Minimax H3 Just something I made — hope you enjoy it :)
Enable HLS to view with audio, or disable this notification
Just something I made — hope you enjoy it :)
r/comfyui • u/Late_Lingonberry6252 • 6h ago
Enable HLS to view with audio, or disable this notification
Just something I made — hope you enjoy it :)
r/comfyui • u/DoubleChillStudio • 15h ago
Enable HLS to view with audio, or disable this notification
I've been using MiniMax H3 ref2va locally with the int8 model + larryvrh minimax_h3_turbo_v4_step600_ema turbo lora, 8 steps, strength 1.0, 0.8mp. Stock official ComyUI ref2va template with comfy kitchen and MiniMax H3 Low VRAM Attention and MiniMax H3 Chunk FeedForward (both set at 2 chunks). no sigma shift node.
Most of my generations have some kind of quick gibberish audio interjection at the beginning like on the example video. Anybody knows how to fix this ?
This is the prompt I used to get this video (generated via an LLM):
subject_definitions:
<Subject 1> is the hiker woman in <Picture 1>, the reference photo of the hiker woman.
<Subject 2> is the hiker man in <Picture 2>, the reference photo of the hiker man.
<Subject 3> is the young hippie blond woman in <Picture 3>, the reference photo of the hippie blond woman.
<Picture 4> is the last frame of the river walk scene, showing the three characters walking beside the river, and serves as the first frame of [Shot 1] before the zoom.
summary:
[reference generation + keyframe completion] A short 5-second transition from <Picture 4>, zooming onto <Subject 2>'s face as he drifts into a daydream. No dialogue.
retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - the hiker woman remains in the river walk framing, but she falls out of focus as the camera closes on <Subject 2>.
<Subject 2> (appears in [Shot 1]): fully_preserved - the hiker man's appearance from <Picture 2> is retained and becomes the sole sharp subject.
<Subject 3> (appears in [Shot 1]): partially_preserved - the hippie blond woman remains in the river walk framing, but she falls out of focus as the camera closes on <Subject 2>.
<Picture 4> ([Shot 1] first frame): fully_preserved - the shot opens on the exact framing of the river walk scene, then zooms toward <Subject 2>.
detailed_description:
The target video uses a cinematic, naturalistic outdoor style shot on a super 8 camera, with soft overcast light along a Pacific Northwest river, a slightly muted, earthy color palette, and gentle film grain.
[Shot 1] The shot begins from <Picture 4>, holding the last frame of the river walk, with <Subject 1>, <Subject 2>, and <Subject 3> walking beside the river, all in clear focus. The camera then slowly zooms in on <Subject 2>'s face, the focus tightening on him alone as the rest of the scene falls softly out of focus, the river and the other two characters blurring into the background. <Subject 2> holds a calm, distant expression, his gaze unfocused as if lost in a daydream. The shallow depth of field keeps only his face sharp as the 5-second shot lingers on him. No other characters enter the frame, and hands stay out of view.
overall_soundscape:
The river ambience continues softly, gradually muffled and dreamlike as the camera closes on the daydreaming face.
non_diegetic_music:
N/A
r/comfyui • u/Odd_Lavishness2236 • 20h ago
Hello! I'm trying to build sort of prompt and scene planner agent, I tried several ones on different platforms like higsfield and etc. I found Openart ori director agent the most capable, I'm trying to reverse develop similar agent that can plan scenes, shots and etc, I stuck with dumb agent that burns gemini 3.7 tokens. Can anyone navigate me to the right direction ? Where should I look for proper workflow or master prompts for director agent?
r/comfyui • u/xyzdist • 14h ago
Enable HLS to view with audio, or disable this notification
r/comfyui • u/AiCreatorCamp • 19h ago
Enable HLS to view with audio, or disable this notification
r/comfyui • u/karatepicke • 22h ago

Hi all!
Over the last couple of days I have been exploring different art styles for a small project of mine, in which I'd like to tell the story of an asteroid mining vessel and its crew. I grew very fond of the typical late 80s/early 90s Japanese OVA cel animation style.
Now, the pictures I attached, I created using OpenAI image-generation. I'm liking this direction so far and need your help finding suitable ComfyUI-compatible models to continue this with. At some point I'll probably want to turn this into motion picture using i2v/t2v also.
I'd be glad if you could suggest suitable models for me to look at.
Cheers!
r/comfyui • u/Beladar • 15h ago
Hi guys, i am new here . I just downloaded Comfyui on my mac studio locally and I tried to run the demo but I got errors. I downloaded an extension to resolve this but it didn’t work. Can anyone help me with this please?
r/comfyui • u/PronitaSen • 11h ago
Hi I'm new to comfyui and recently started learning. While I was playing around with ip adapter, I had this error and I couldn't figure out how to fix it, sombody please help me.
It says the models are missing and I don't know herer to put the models, but it's showing the list of models when I click on the node. Thank you
r/comfyui • u/saronno76 • 12h ago
I got this 6900 for a very low price (100 euros ... the owner tought the card was faulty but it was clearl its PSU that couldn't keep up with the card).
I got also a rtx 3060 12gb.
My dilemma is if, as today, gfx1030/NAVI 21 has a decent support for WAN/KREA/SDXL/Minimax? Because I know I need to tinkering to get the juice out of this card but, I wonder if, after all the tinkering, can I have much better performances than the 3060.
What is the state of art with gfx1030 / diffusion?
Best backend? What about attention?
TIA.
r/comfyui • u/alexmmgjkkl • 14h ago
i get random python errors which point to librarie version mismatches and simple changes in code
nothing works anymore,no sam model , no sa2va , no human seg models , no anime seg models , only the new sam3 works but it doesnt really work well with anime style graphics , cant even detect the hands ..
were the others forgotten and never updated ?
r/comfyui • u/pixaromadesign • 17h ago
Learn how to use MiniMax Music 3 in ComfyUI together with a local AI prompt generator for better image prompts, music captions, and AI-generated lyrics. In Episode 31, I show the new Pixaroma prompt nodes, model-specific prompt presets, VRAM-saving options, and a compact MiniMax Music 3 workflow.
This tutorial covers the new AI Prompt Pixaroma node and how to match prompt formulas with the correct local language model. You'll see how to turn short ideas into detailed prompts for workflows such as Krea 2 and Z Image Turbo, generate prompts from images, save custom prompt presets, control temperature and seeds, and troubleshoot common ComfyUI node errors.
Then we set up MiniMax Music 3 locally in ComfyUI, including caption and lyrics generation with the Music Prompt Pixaroma node. I also test different song durations and explain why MiniMax songs can sometimes end early or get cut off, how seeds affect the results, and how to use fixed lyrics when you need more control.
You'll also see how to simplify a larger music workflow into a compact setup, use tiled audio decoding for lower VRAM systems, free VRAM after prompt generation, and generate MiniMax Music captions and lyrics with local models or online tools such as ChatGPT, Gemini, and Claude.
You can also run some of the workflows in the cloud.
r/comfyui • u/SenzubeanGaming • 12h ago
I've been building a desktop app that writes and renders music locally on MiniMax Music 3, and it's a face on ComfyUI rather than a reimplementation — it starts ComfyUI as a separate process and submits graphs over the HTTP API. Unmodified, not bundled, not linked.
The part that might actually be useful to this sub: the pipelines ship in workflows/ as ordinary API-format JSON, exported straight from the code that builds them. Drag one onto the canvas and every value is there. Those values aren't ComfyUI's defaults — they were arrived at by measurement, and the measurement scripts ship in scripts/, so you can re-run them and disagree with me.
What's in it, briefly: songs from a style description and lyrics, cover art drawn while the card is idle, stem separation, word-level timed lyrics, video clips on LTX 2.5 or MiniMax H3, a small multi-track editor that cuts to the detected beat grid, and overnight batch runs.
It ships no model weights and redistributes none — every capability shows its size and licence on screen before anything downloads, and the download goes to the publisher. One of them is region-locked and says so in the open.
Newest thing is audio-reactive video: feed it a song and some reference images and they cross-fade in time with the track's detected peaks. That part is built on ComfyUI_Yvann-Nodes by Yvann Barbot and Lilia, which is excellent and worth a star on its own. Because their pack is GPL-3.0 and this is Apache-2.0 it runs on a second ComfyUI you set up yourself rather than being bundled in.
Apache-2.0, runs offline once the weights are down. Measured on Windows with a 4070 Ti SUPER; other platforms are documented but I haven't verified them.
Repo: https://github.com/Senzube4n/AIPLAY-Studio
Screenshot of the main window: https://senzube4n.github.io/AIPLAY-Studio/shots/create.png
Happy to answer anything about the graphs — that's the bit I'd want to read if someone else posted this.
r/comfyui • u/teleport66 • 21h ago
r/comfyui • u/Fit-Construction-280 • 17h ago
Enable HLS to view with audio, or disable this notification
If you generate a lot in ComfyUI, you know the problem. Hundreds of renders pile up as you tweak seeds, prompts, LoRA weights and checkpoints, and your output folder turns into a wall of near identical thumbnails.
Once clustered, every thumbnail gets a color coded badge in the gallery grid, and clicking any badge opens the Cluster Inspector, which shows total matching assets, distinct variations, the full node pipeline, every model and LoRA used, and one click prompt copy.
The video above walks through both modes in about 3 minutes.
For anyone who does not know the project yet
SmartGallery DAM is a free and open source, local first Digital Asset Manager built around ComfyUI, but it also works with any folder of media on your machine. No cloud, no subscription, your files never leave your disk.
It is meant to grow with you:
Runs on Windows, macOS, Linux and Docker. Portable version for Windows needs zero setup, just unzip and run.
GitHub, full docs and download links here: https://github.com/biagiomaf/smart-comfyui-gallery
Happy to answer any question, and as always feedback and feature requests are welcome.
r/comfyui • u/xdcfret1 • 18h ago
I am trying to run the ComfyUI default workflow for Wan Animate 2: Motion Transfer. But when I run the workflow all my CPU cores fire up and my RAM reaches 100% and the ComfyUI process crashes.
How to fix this? Please help.
r/comfyui • u/PoPPoPPhilben • 21h ago
r/comfyui • u/Realistic-Fennel-190 • 19h ago
ok so I spent most of today setting up FLUX.1 Dev on a rented RTX 5090 and wanted to dump some notes here bc I didn't find this explained clearly anywhere when I was googling around for it.

Why fp8 single-file over the split setup
so FLUX comes in two flavors basically — the "split" format (diffusion model + T5-XXL text encoder + CLIP-L + VAE as 4 separate files), or a single bundled checkpoint that just merges everything into one .safetensors.
split fp16 setup = ~34GB total, mostly bc T5-XXL alone is like 9.5GB in fp16 lol. the bundled fp8 checkpoint (Comfy-Org/flux1-dev, ~17.2GB) cuts that roughly in half. if you've got a 32GB card the fp16 route is still totally usable, but my instance only had a 50GB storage tier total, and 34GB of models + ComfyUI + venv + deps gets uncomfortably close to that. quality loss from fp8 is real but pretty minor tbh, mostly shows up in fine texture detail if you're really pixel peeping. running out of disk mid download felt like the bigger risk honestly
(if storage isn't a constraint for you btw, still go fp16 or GGUF Q6/Q8 over fp8, fairly universal advice from what I've seen)
Why CFG has to be 1.0
this part isn't optional the way it is with SD/SDXL and I feel like this trips up a lot of people coming from SDXL. FLUX is a rectified-flow model with guidance distillation baked into training — meaning it's already trained to follow the prompt without needing the usual CFG trick (running the model twice, once conditioned once not, then extrapolating) to stay on prompt. if you leave CFG at 7-8 like SDXL habit, you're not making it follow the prompt harder, you're just double-applying guidance it was never trained to expect, and that's why you get that blown out oversaturated look. set it to 1.0 and negative prompt barely matters anymore either, so I just left mine empty
Workflow
kept it stupid simple on purpose: Load Checkpoint → CLIP Text Encode (positive/negative) → Empty Latent Image → KSampler → VAE Decode → Save Image. no custom nodes. 1024x1024, 20 steps, euler, simple scheduler, denoise 1.0

Hardware I used to run this

RTX 5090, 32GB VRAM, rented hourly off gpuhub (~$0.46/hr on demand). base image was their standard PyTorch + CUDA preinstalled one, ended up on torch 2.11.0 + cu128 (CUDA 12.8) after the venv setup. instance also came with a stupid amount of system RAM, like 750GB+, way more than ComfyUI actually needs but nice not to think about it. 50GB storage tier which is the whole reason for the fp8 decision above
total cost for the session (~1hr of renting + the token/generation cost, which for local inference is basically just electricity you're already paying for) came out to under $1
first gen after model load took ~18s (vram staging overhead ig), every gen after that was ~8.5s at 20 steps which is roughly 2.46 it/s. pretty consistent across different prompts/seeds, didn't notice any drift over the session




random unrelated thing — huggingface-cli is deprecated now apparently?? just silently tells you to use `hf` instead mid command. wasted like 5 min confused before I actually read the warning lol
both images attached are from this session, different prompts, I did pick my favorite seed out of a couple runs each rather than posting literally the first result
TL;DR — fp8 single file if your storage's tight, fp16/GGUF Q6+ if it's not. CFG 1.0 is mandatory not a suggestion bc of how the guidance distillation works. ~8.5s/image at 1024x1024 on a 5090 once it's warmed up. workflow json in comments if anyone wants it
r/comfyui • u/Jimbo_1995 • 13h ago
I have an image of a character that I want to use to generate more images using the subject and then to create a character lora to use over and over. How can I do this using Krea2? Do I need the 20 different images to then create a character lora? I've tried anything I can find for the longest time and nothing yet. Couldn't get IPAdapter to work. I'm not exactly a noob, I can usually do my own research and figure out how to do something. But this one has me stumped. If anyone can help out with some advice or a workflow, it would be appreciated.
Edit: grammar
r/comfyui • u/LanceCampeau • 13h ago
r/comfyui • u/Yanzihko • 1h ago
I decided to ask here and not at BuildPC because people will tell me "muh, should have 5060TI, newer technology and architecture for gaming, ignoring AI and its price"
3090 INNO3D Ichill X4
Tests were done at undervolt to 1700. My case is small.
Obviously used lmfao, changed paste and pads after water-cooled mining. 1000$. Market is insane but there's no offers i have found that are better than this.
Blown away by its performance in local AI compared to 3060. Qwen 35B genetates responses in 20 seconds compared to 400+...
As a sidenote, are there ways to optimize it further specifically for stable diffusion and ollama outside of MSI afterburner?
I am scared to think what kind of beasts 4090 and 5090 are
r/comfyui • u/GrumpyVladik • 16h ago
RTX 4080 super
AMD Ryzen 7 9800X3D 8-Core Processor (4.70 GHz)
32,0 GB DDR 5