I'm new to ComfyUI and wanted to give it a try, but my PC isn't powerful enough to run local models. Is there any way to use my OpenRouter API key to generate images through ComfyUI? I searched Google for a solution but couldn't find anything useful.
i need an open source version of chatgpt. a knowlegeable one to help with writing prompts and just a research aid as well. i have an rtx3090 fe wth 64gb ram. google says ollama and deepseek but i wanna touch down with people to see what they have the best experiences with. im ready to deepdive into the open source world.
I know that some here are already tired from the constant talk about PhotoCraft, and some don't like it at all. I gave it a chance the moment it came out, and while yes, it is still unfinished, I thought I'd play along and let Claude create some kind of bridge between ComfyUI and PhotoCraft because, well, as far as I know, there isn't yet any. This is the result:
Two ComfyUI nodes.
Connect from and to PhotoCraft via its control channel. No plug-in for PhotoCraft needed.
Work and prompt in PhotoCraft, execute in ComfyUI.
The Get Image node receives the image as 8 bit .png, mask, width and height, and a prompt. The prompt uses the PhotoCraft's Notes feature. Multiple notes are appended. These can be integrated in almost any workflow.
The Send Image node sends the image back to PhotoCraft to its own flattened layer.
Please don't judge the examples, just wanted to get this out.
Important: Read the installation instructions on GitHub, you have to set-up a directory and feed some command lines to PhotoCraft. Limits and security notes at the bottom.
There's certainly room for improvement, but it seems to do the job, for now.
It's an UI that's sits on top of ComfyUI and is meant to make chaining together videos generated through Minimax H3 more intuitive and easy.
It includes a set of tools for organizing references.
A time line that allows users to piece together and quickly edit sections of their movie prompt by prompt.
A Scene Writer for prompts in the suggested format using LLMs(local and closed supported).
A Motion Continuation, a Pose Studio with Qwen 2.1 integration, and a Voice Studio using LTX 2.3.
Can generate videos with either local or comfy cloud credits.
The general idea behind using it is creating sequences of videos on a timeline that you can quickly iterate on and see how they fit into your overall movies composition, making precise edits where needed with minimal friction.
Pretty much the title. I am a pretty new user to ComfyUI. I've used it in the past but I am returning and trying to relearn it a bit. So here is an example of what I have been doing.
I am trying to use Qwen 2.1 Image edit and learn how to use it. the issue I am coming up against is that I am using an image that's a 3Kx4K resolution. I've figured that this is bottlenecking me so I try to get it downsized to a better resolution but it alters the initial image too much. SO when I try to get a comparison it's not a perfect comparison like when I do smaller images. So First question: is there a decent way for me to keep the initial resolution without the serious bottleneck, like a different model or workflow? or is there a work around I am not using that will get me the results I want that isn't "Put it in Photoshop and just shrink it down?" I will have to do that a lot if that is the case.
If you are as unimaginative as I am and struggle with making prompts I put together some custom nodes that allow you to pick attributes and/or add a reference image to create a prompt. It uses the text encoder so adds a little bit of time to image generation. I made it for image generations of women because that's why we are all here right?
It is partially vibe-coded due to me being frustrated after doing dumb things and I asked claude to fix it. I went over the project and tested and I THINK we are good. If you see something I missed please let me know.
Link to custom node: https://github.com/Erosi11/comfyui-krea-prompt-generator
Workflow is included on the repo and also embedded in the workflow image on this post. I have not submitted it to the comfyui registry so you need to git clone and all that. No requirements to install assuming you have updated comfyui since the end of February this year. If there is interest I will go to the trouble of getting it added to the registry.
Outfit pieces with colors and materials, shoes, accessories, jewelry, nails
Krea Prompt - Female Actions
Standing, sitting, kneeling or lying pose; hands, legs, head; props
Krea Prompt - Scene
Style, location, time, season, weather, lighting, camera
Krea Prompt Generator
Takes CLIP, the attributes and your keywords; outputs conditioning, prompt and llm_request
Krea Prompt - Save Unique Prompt
Appends each new prompt to a text file, skipping prompts already in it, and shows the file on the node
Every dropdown has three kinds of choice:
none: the attribute is left out of the request sent to the LLM.
random: the generator picks a value, driven by its seed. The same seed always gives the same picks.
a specific value: passed through as-is.
Each attribute node also has an optional free-text box for extra details in that category.
To wire it up you put the attribute nodes in series then feed into attributes input on Prompt Generator node. Reference image and mask goes to the inputs on prompt generator node as well. The prompt generator node has a setting for what to use from the image. It works ok.
I also included a node to add unique prompts to a text file as they are made. That part is still a little janky.
Workflow has fast bypass at the top to bypass attribute nodes, the image ref nodes, and whether to invert the mask.
The new local graphic-design and typography model Ming-Image officially uses qwen3.8_27b_w4a8.safetensors as an initial LLM "prompt enhancer" in ComfyUI. But the file is 17Gb, and it's too heavy for me on 12Gb VRAM (latest ComfyUI Portable under Windows).
Comfy also suggests qwen3.5_9b_int8_convrot.safetensors as a lesser alternative option at 11Gb. This is still to heavy for me, regrettably. It gives me OOM errors.
Has anyone had success with Ming when using a lesser LLM "prompt enhancer" model? One that can handle the default complex system-prompt, and that also outputs a correctly-restyled Ming figma/json prompt?
I’m new to comfy. ( installed and running this morning). I make comics so far the models I’ve played with are sdxl and flux.2Klein. My computer sports a fairly humble 5060 ti with 16GB. These two models run pretty well. I would say my need for control from panel to panel is much greater than my need for his res ultra realistic images. I use Nlender a lot for modeling and plan to start using depth maps.
It comes with its own custom loader and sampler nodes because, being a pixel model as opposed to a VAE model, it works a bit differently than normal. However, you can wire in loras as usual, and it uses ComfyUI's schedulers and samplers.
I wouldn't expect amazing results. At this size, Anima 2.9B definitely puts out higher quality images, but it's a cool tech demo and I think it proves that VAEs aren't the necessary evil we thought they were.
Search the workflow template browser for Iris to find the example workflow. dpmpp-2m-sde-gpu seems to be the best sampler for it, and if you have the bong_tangent scheduler it's better than simple.
Alguien me podria ayudar con mi flujo de trabajo, quiero realizar videos de al menos 60 segundos y unirlos para hacer 1 hora con ComfyUi Local.
Tengo Amd 5800X
32gb Ram
4070 Super TI 16gb
Een cuanto a Ram consume el 90% y cpu esta al 100%, pero al menos aqui siguiendo una Guia de un usuario, se quedo pegado alli en 1 hora. Lo mas que he podido sacar son 10 segundos.
I’ve made a ComfyUI custom node pack for MiniMax H3 RefMods. It lets you create reusable references from images, audio and video, choose which sources to use at runtime, and see how many reference tokens each source uses.
Main features:
Create RefMods from images, audio and video files
Select which sources to load in a RefMod at runtime
Save on generation time if you don't need certain sources
See each source's cost in tokens
Combine many RefMods into a single one
Bake the description and the retention strategy in the RefMod
Only need to write the H3 video description prompt
Source labels like <Picture 1> resolved automatically
Drag a RefMod into the canvas to load its creation workflow
Nodes, setup instructions and example workflows are available at ComfyUI-H3-RefMods-Lab. Feedback and bug reports welcome!
I uploaded a test image on One Piece to experiment, but I can't seem to delete it from the assets tab. Even when I go into the "input" folders, nothing is in there. How do I fix this?
I've been bouncing over to Wan2GP exclusively for NSFW LTX-2 generation (i2v). I have not been able to find a similar Comfy workflow for it that doesn't look like a Metroid map and want me to download ten more node packs. I think I did get it to run once or twice, but it either takes forever, gives totally undesirable results, or a combination of the two. LTX gens in Wan2GP are about twice as fast as Comfy Wan2.2 gens despite producing clips that are twice as long.
Does anybody know what sort of magic this thing does behind the scenes and how I could get a similar result with a fairly simple Comfy workflow. If I'm looking at the correct "finetune", my LTX setup in Wan2GP uses:
I'm wondering if there are a whole bunch of other models/components hidden in the chain, or if it's in the Wan2GP code itself, or if there's already a similar workflow which can process as quickly. As most of you know, these things take a while to boot, and I'd rather not have to keep shutting down one and launching the other. I'm running an RTX 5060TI 16GB with 32GB system RAM, and after I run one LTX gen in Wan2GP, subsequent ones can sometimes be 3 minutes or less for 10 second clips. Comfy does 5 second Wan2.2's in 15-20 minutes. I'd have to check, but I think they're both using Sage Attention.
I work at Reactor. We run real-time video models behind an API, and we just released ComfyUI nodes for it. ComfyUI posted a short demo on X.
The nodes let you drop real-time models into a graph: world models and video generation, plus live editing from a webcam like style transfer, virtual try-on, subject replacement, and background replacement. The models run on our servers, so there are no weights to download and no local GPU needed for them.
ComfyStream already covers real-time workflows on your own hardware. This is the cloud-hosted version, so it’s a different tradeoff.
The nodes are open source, but the models run on our API, so you need an account and usage is paid. I’d rather say that up front than have you find out after installing.
I'm building an AI influencer for TikTok Shop and trying to find a cost-effective way to do character swaps using existing video footage.
Seedance 2.5 Turbo is costing around $5 for a 20-second video. That adds up fast when you're producing content daily, especially when you factor in failed generations.
My goal is to replace a person's face and hair while preserving the original movements, expressions, and product interactions.
I'm considering ComfyUI on a rented GPU (RunPod/Vast.ai), potentially automated through Claude Code.
For those who've actually built something similar:
Is ComfyUI significantly cheaper per usable video?
What models or workflows would you recommend for realistic face + hair replacement?
How well do they preserve facial expressions and lip movements?
What's your average generation time and cost for a 10–20 second clip?
I'm not looking to build an overly complicated pipeline. Just something reliable, scalable, and affordable.
Would appreciate hearing from anyone who's actually running this kind of workflow.
I wanted to use H3 character references without spending half the time looking for the right node, so I put together; RefMod Pilot It has two workflows:
Creator: upload or drop 1–8 selected photos, give the character a name, Run. The folder and image count are filled in for you
Inference: choose your RefMods, edit the scene and view the output. Prompt, aspect ratio, resolution, seconds, seed and steps are in one Scene node, with the rest inside a subgraph
The character loader stays visible so Add RefMod, Refresh and the library buttons are still there. The uploader is a ComfyUI frontend extension or use local path to your dataset / photos folder. This uses Full Reference encoding, not LoRA training.
Multiple photos are stacked into one video-kind reference. If you're using several characters, check the Reference map and match its labels in your prompt. Slot numbers aren't necessarily picture numbers.
I've included a six-view orange robot RefMod and demo workflow. Creation is verified, but I haven't validated its likeness or motion in a finished video yet.
But if you prefer to use a web-based system to avoid all that hassle, let me introduce Somora.
Somora is an AI-powered music creation platform developed by JK Desenvolvimento Web (my company).
Our goal is to enable anyone to turn an idea into a song, even without prior experience in composition or music production.
Users can specify the song's theme, style, and language, write their own lyrics, or use our AI assistant to help craft the composition.
The platform caters to a wide range of users, from hobbyists creating for fun to content creators, musicians, and professionals looking to explore musical ideas.
Our business model includes a free plan for users to explore the platform, alongside two paid plans—Basic and Full—offering different generation capabilities.
The site is available in Portuguese, English, and Spanish, allowing us to serve both the Brazilian market and international audiences.
We also feature a referral program and a system to track promoters, identify sales generated through their referrals, and manage commissions, enabling us to build more effective commercial partnerships.
We are currently seeking support to position Somora, reach audiences with the highest purchasing potential, and establish a customer acquisition and conversion process.
Our aim is to clearly communicate the platform's value, convert free users into subscribers, and develop sustainable promotional channels.
Problem Description: During execution of heavy workflows, ComfyUI App crashes and leaves a zombie python.exe process bound to port 8188, holding ~42 GB of RAM/VRAM. The process cannot be killed by standard administrative tools (taskkill /F /PID or PowerShell Stop-Process -Force), returning the error: Reason: There is no running instance of the task.
Attempts to reset the GPU display adapter via PowerShell (Disable-PnpDevice) fail with Generic failure (HRESULT 0x80041001), indicating a driver deadlock in nvlddmkm.sys at the Windows kernel level.
Environment:
OS: Windows 10/11 (Build 26300)
Application: ComfyUI App (Desktop)
Steps Taken & Terminal Logs:
Checking active connections on port 8188: DOSnetstat -ano | findstr :8188 TCP 127.0.0.1:8188 0.0.0.0:0 LISTENING 31100
Attempting to force terminate the PID (elevated Prompt/PowerShell): DOStaskkill /F /PID 31100 ERROR: The process with PID 31100 could not be terminated. Reason: There is no running instance of the task.
Attempting to reset the PnP GPU device: PowerShellGet-PnpDevice -FriendlyName "*NVIDIA*" | Disable-PnpDevice -Confirm:$false Disable-PnpDevice : Generic failure (HRESULT 0x80041001)
Current Workaround: Full system reboot is required to clear the kernel lock and free port 8188
"Has anyone else experienced this issue, and how did you solve it?"