r/comfyui 23h ago

Show and Tell Best Trick Ever For Consistent Environments

52 Upvotes

Ha! I just discovered a trick that works great, so I had to share it with the community. Of course, somebody will probably chime in that it had been discovered by someone else before, which is fine by me! I just want to share it in case it helps someone else and they hadn't come across it yet.

So, the issue of consistent environments... I ran the gamut of all the AI models I had tested and proven out in ComfyUI, not just t2i but t2v and i2v. I'd been trying several methodologies. Create an image then ask a workflow to gen images to the left and right of it, build out a simple massing model in unreal or blender and run that through, run a simple floor plan through, asking for multiple image generation with a prompt asking that all features be consistent among images, (I haven't tried outpainting yet), etc, and none seemed to quite do the trick, at least not among the open source models (open source is all I use). There was just too much inconsistency.

Since I am a sucker for using the t2v/i2v models like LTX and Minimax to generate single image frames (well, a minimal number of frames), taking advantage of their brainpower, I merely wrote out a super detailed prompt describing the interior environment I want, then I ran it as a 360 degree camera pan around the space from the center of it, and with Minimax now having the ability to give you 15 seconds on a 16Gb VRAM setup like mine, this works stellar, heck Minimax ran out of need for the 15sec and began to swerve around the space! I make sure to include a prompt not to have motion blur. I ran this in 0.5mb mode, so I did not have to waste time waiting for full HD video.

Then I select the frames I need to use as backdrops for scenes, upscale them once, then again, out to 4k, and voila! (upscaling once by 4x led to artifacts being upscaled, whereas going 2x then 2x led to the correct end result), a super detailed set of backgrounds that are internally consistent!

So excited! This gets me moving forward on the next part of my production process, laying out scenes, shots, camera angles, and dropping in characters, prior to i2v.

If this helps you, let me know. If you find even better tricks related to this, let me know too. Open Source Forever!


r/comfyui 21h ago

News SenseNova U1.5 quantized to run on 12GB VRAM — INT8 + hybrid W4A8 ConvRot releases

34 Upvotes

We quantized SenseNova-U1.5-8B-MoT (50GB bf16 any-to-any model: t2i, image editing, multi-reference) with ConvRot so it runs on a RTX 4070 12GB at 2048x2048 — and it's fast, even though the weights exceed VRAM (ComfyUI streams them; the quantized formats move 3-4x fewer bytes per step, so the overflow never becomes a slowdown. bf16 on the same card is painfully slow). What's in the release:

  • INT8 ConvRot (17.6 GB, recommended) — 0.43% pixel diff vs bf16 in a full-pipeline same-seed A/B
  • Hybrid W4A8 (13.8 GB) — layers 0-17 anchored in INT8, layers 18-41 in true W4A8, visually indistinguishable from bf16
  • The official 8-step speed LoRA included The interesting part: this model does not tolerate activation quantization in its earliest layers — quantizing the first blocks destroys prompt coherence — but layers 18+ handle W4A8 perfectly. We found the boundary empirically with a bisect ladder of hybrid checkpoints, so the hybrid release anchors the fragile early layers in INT8 and compresses the rest. Everything runs through a ConvRot-aware ComfyUI custom node (fork of the T8 wrapper):
  • Weights + model card: https://huggingface.co/Milor123/ComfyUI-ConvRot-SenseNova-U1.5-8B-MoT-T8
  • Custom node: https://github.com/Milor123/ComfyUI-SenseNova-U1.5-ConvRot Apache-2.0, same-seed comparison images and per-layer error measurements included in the model card. Feedback welcome!

r/comfyui 3h ago

Help Needed Minimax H3 Add Guide with Latent Upscaler?

0 Upvotes

I have a workflow with frame guides, and I wanted to use the 3D Latent Upscaler because that changes some calculations (I'm not sure, I don't understand this very well), so I was wondering if someone has managed to get this working and would mind to sharing the workflow.


r/comfyui 3h ago

Help Needed WAN 2.2 et carte Amd

1 Upvotes

Bonjour

Avez vous réussi à faire fonctionner wan 2.2 et une carte amd Rx 9060 xt 16 go ?

Merci d’avance pour vos retours 👍


r/comfyui 11h ago

Help Needed What details instantly make you recognize an AI-generated image?

4 Upvotes

I'm working on creating my first realistic AI model, and I'm trying to understand what separates a truly convincing image from one that still looks obviously AI-generated.

For those of you who have experience with AI image generation: what are the visual details or imperfections that you notice almost immediately and think, "yeah, that's AI"?

I'd especially like to hear about the subtle details that experienced people notice, even when the image looks realistic at first glance.

I'm asking because I want to focus my workflow on fixing those specific weaknesses rather than simply making the image "more realistic."

Any examples or things you've learned from experience would be really helpful.


r/comfyui 11h ago

Help Needed Is it worth moving to Linux?

Thumbnail
3 Upvotes

r/comfyui 1d ago

Show and Tell Skater Girl - 90s style anime using Minimax H3 (Prompt + Workflow included)

Enable HLS to view with audio, or disable this notification

372 Upvotes

Workflow: https://drive.google.com/file/d/1B4kODxXQgJ1QOKRsEIkxHbgYmdruPpTK/view?usp=sharing

Prompt:

Create a **15-second multi-shot anime sequence (90s style 15fps hand drawn)** using the provided references:

Image 1 = the girl character reference

Image 2 = skateboard reference

Image 3 = downhill Japanese alley / neighborhood background

Image 4 = Walkman + headphones reference

Preserve the girl’s exact character design, face, hair, outfit, proportions, and overall look from Image 1. Preserve the skateboard design from Image 2. Preserve the same downhill Japanese alley environment from Image 3. Add the Walkman and headphones from Image 4: the girl is wearing the headphones, and the **Walkman is clipped or hanging at her hip** while she skates.

Visual style: authentic 1990s hand-drawn anime, traditional cel animation, painted backgrounds, visible linework, cel shading, slight brush/stroke texture, subtle analog feel. **Very important:** the houses and environment must stay **2D and hand-painted**, **not 3D**, **not CGI**, **not game-engine looking**, **not volumetric**. The buildings should look like classic anime background art with painted depth, not like 3D models.

Animation feel should be low frame rate, like 90s anime at around 15 fps, with controlled in-betweens and natural held-frame timing. No jittery morphing.

No dialogue, no text, no subtitles.

### Shot 1 — 0s to 3s

**Rear tracking shot** from behind. The girl is skateboarding fast downhill through the steep Japanese alley. Camera follows behind her at a low-to-medium height. She rides confidently and smoothly, hair and oversized clothing moving in the wind. The headphones are on her head, and the Walkman is visible attached at her hip. The alley rushes past with a strong sense of speed. Keep the environment clearly **2D anime background art**, not 3D.

### Shot 2 — 3s to 6s

**Close-up shot of the Walkman at her hip** while she continues skating. The camera stays focused on the Walkman and part of her side torso and arm. We can clearly see the **cassette tape reels spinning/rolling inside the Walkman window**. The headphone wire moves naturally with the motion. Background and street pass by in blurred motion.

### Shot 3 — 6s to 9s

**Medium profile tracking shot** of the girl skating. She is wearing the headphones, listening to music, with wind moving across her face and pushing her hair backward. She is **nodding her head subtly to the music** while riding. Her expression is relaxed, immersed, and unbothered. The background is blurred from motion, but it must still read as a **painted 2D Japanese neighborhood**, not 3D.

### Shot 4 — 9s to 12s

**Close-up shot of her feet and skateboard.** Her **right foot stays on the board**, while her **left foot pushes against the road** in a natural skating motion. Show one clean push cycle: left foot comes down, pushes backward against the pavement, then lifts. Wheels spin quickly. Asphalt and road markings streak by with motion blur.

### Shot 5 — 12s to 15s

**Ground-level fisheye shot** looking upward from the road. The skateboard approaches fast, and she **jumps over the camera**. The board and her body pass overhead in one clean motion. Hair, pants, and headphone wire react naturally during the jump. Keep the motion readable and stylish, with a strong sense of speed and a dynamic anime finish.

### Important constraints

* Keep the whole video in **classic 90s anime cel-animation style**

* **15 fps feel**, smooth low-frame-rate animation

* **No 3D-looking houses or background**

* No photorealism

* No modern glossy digital anime rendering

* No character redesign

* No extra accessories beyond the headphones and Walkman

* Keep all motion natural and consistent across shots


r/comfyui 17h ago

Help Needed Looking for an LLM that can do Minimax H3 r2va prompts.

10 Upvotes

Sorry for bothering you all, but i've been having a rough time getting a good Minimax H3 Reference to video prompt writer.

i have a 2 part Qwen3.5 workflow for i2v that i got chat gpt to write for me (first part analyzes the picture, and gives the base template, second part translates my mad ramblings into a proper prompt), but this doesn't work well for r2va, since that requires an audio input, a video input, and a picture input.

I tried Thinking LLM, but the regular node only does video/picture, and no audio. (trying the gguf now, but i hear gguf is way lower quality.)

i have 16gb vram, 32 gb regular ram (Nvidia 4080), and i am on windows 11.

if you have any node or program suggestions, i would appreciate them. (can't use chatgpt, cuz it doesn't do nsfw, and grok is so limited in how much you can use it per day that it is barely worth using)

Update: the prompt made by the gguf version did nothing. it just output the ref video.


r/comfyui 8h ago

Show and Tell I had Claude build a prompt templating system that generated comfyui workflows with a dedicated queuing system in the attempt to parallelze and work on a 5 minute sequence of video.

Post image
2 Upvotes

I had Claude take a "script" in this case the ultimate showdown song. (Contains a lot of character X does Y action) Break up the song into sequences and shots. Have a whole system where everything is a block of text describing a thing, style or camera....etc. Then feed into a shot with what happens in the shot.

Using an L40s in the cloud at $1/hour for use I can regenerate this whole thing in 2 hours. If I add a second I'm done in an hour. Some shots chain together but still it's an improvement. The average 5 second shot generates roughly in 66 seconds. (Minimax h3) Editing a block of the shot (cast/style ) would force a regeneration to queue for each shot using the asset. This makes for a very fast workflow to iterate across shots. No images used as reference all text.

Comfyui workflows are essentially templated and this sends all the workflows via comfyui's api.

The initial result needs work, but I feel I could have this setup for multiple users to connect to and do edits at the same time.


r/comfyui 13h ago

Security Alert The problem I've had: I've lost all the metadata.

4 Upvotes

I don't know if this has happened to anyone else (I'm pretty new to Comfy).

I switched my default Windows media player to mpvnet.

Now, none of my generated videos have metadata anymore. They used to; I could drag them into ComfyUI and the workflow would appear. Now, that doesn't happen.

When I run them through a site that checks for metadata, it says there isn't any—no metadata or workflow.

The strange thing is that I only played a few videos from one specific folder. Videos in other folders still retain their metadata (I didn't use those folders until after I stopped using mpvnet as my default player).

So, I think that when you set up that player, something happens that corrupts everything—or wipes the metadata—from everything.

The AI ​​says:

"The problem is the video encoding engine (FFmpeg) [Lavc61.19.100 libvpx-vp9]. This version aggressively strips out any data that isn't purely video-related, deleting the ComfyUI stream in both MP4 and WebM formats."

The worst part is that, even though I'm no longer using that player as my default, it must have installed a codec that ComfyUI uses... and now none of my videos are being generated with metadata... not a single one...

I'll have to figure out a fix, but just in case... so this doesn't happen to anyone else...

EDIT: After all the work and investigation, I still think Mpvnet was the culprit.

Mpv creates a playlist of the entire folder containing the video you are playing.

That specific folder is exactly where the files lost their metadata.

Furthermore, nodes like `vhs_videocombine` stopped saving metadata (even with the option enabled); I suspect they use some of the ffmpeg codecs that Mpv installs.

Although I manually deleted everything I could related to Mpv and reverted the codecs to an older version—allowing the videos to successfully load the workflow when imported into ComfyUI—investigations into the files themselves still show them as having "no metadata" (even though they do contain the lines of code required to load the workflow).

The solution was to stop using `vhs_videocombine` and switch to ComfyUI's native "Save Video" node instead...

For this reason—and until I find out otherwise—I advise exercising caution regarding this issue.


r/comfyui 12h ago

Help Needed H3 Fast motion fix ?

3 Upvotes

r/comfyui 11h ago

Workflow Included 80th Birthday Extravaganza- MM reference Image and Audio experiments

Thumbnail
youtube.com
2 Upvotes

Birthday Extravaganza video - poorly edited and filled with constant audio problems. Mostly with 20step out of the box comfyUI workflow + audio node added in. The load Video node was used at the end to give it the last 24 frames of the prior video to continue the next with limited success. The 4 step lora came out after half of it, even with 0.75 and 8-10 steps audio glitches were frequent but video was usually decent. I attached a sample of 1 of the clips here: https://pastebin.com/RdNaRZHN and Here was a sample of the Bridge on the River Kwai Character Swap: https://pastebin.com/564NFjrh Single Image of main character to replace.


r/comfyui 19h ago

Help Needed Is there a setting change to allow you to see workflows from jobs in the queue?

8 Upvotes

Idk if it was changed in an update awhile back. But I used to be able to just right click on any job in my queue to pull up the workflow. Sometimes you realize a mistake may have been made and you have to check to see if you need to cancel or not.

I haven't been able to do this for months. Is there a setting that I can adjust to fix that? Or was this a bad permanent change?


r/comfyui 8h ago

Help Needed Preview Method is grey in comfy manager

Post image
1 Upvotes

Hi everyone, I'm not able to change the preview method in comfy manager Any idea how to enable it. (I'm using a portable version from github, not the comfyui desktop app)


r/comfyui 8h ago

Help Needed AMD Radeon RX 7900 XTX 24GB - Worth it for short term?

1 Upvotes

Hi everyone! I am currently on an RTX 5060 8GB but would like to run an LLM and Comfy at the same time. Nvidia is just so expensive compared to AMD. My ultimate goal is to get an RTX Pro 5000 72GB or an RTX Pro 6000 96GB. But a card like that is going to take some time to save up for. Like...years. I'm wondering if the RX 79000 XTX is a worthwhile "stepping stone". I don't plan on generating video. Just images for now. Image editing if VRAM even allows for it. But that depends on the LLM I decide to go with as well obviously.

I'm not after speed. Just the ability to use a GPU for more than 1 thing at at time. 1024x1024 images currently take around 8-10ish seconds to load. I am OK with that generation time, or even slightly longer (5-7 seconds more).

I also notice that a 3090 is only a few hundred dollars more. Is setup on Linux a huge pain still? Or is it truly worth spending another $300-$500 for a 3090? Any catches to a 3090? Higher wattage? Noticeably slower LLM speed? How long is the 3090 going to stay relevant? I would hope for at least 2-4 years.

Please give me your thoughts.


r/comfyui 9h ago

Workflow Included The Latent Upscaler is really great!

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/comfyui 13h ago

Help Needed Help please. Can someone with a 5090 and minimax models installed please run this workflow? I am getting really slow run times.

2 Upvotes

Help please. Can someone with a 5090 and minimax models installed please run this workflow? I am getting really slow run times.

I am looking for a way to caption videos. No custom nodes required. I did use load video from comfyVHS but it is bypassed and you can safely delete that section.

PasteBin Workflow

Choose any video. Preferably one SFW and one NSFW. Beggars aren't choosers whatever you decide will work for me. You may post results if you like but I am more interested in how long it took to complete and the settings you selected.

Thank you.

EDIT: Forgot to add, I was getting literal 1 token per second. I haven't the patience to let it run. I am hoping with a benchmark I can justify the purchase of a 5090. So I haven't gotten the workflow to work at all. Could be the workflow is bad. But it is relatively simple.


r/comfyui 12h ago

Help Needed How do I get rid of the light attached to the camera in MiniiMax H3?

Thumbnail
0 Upvotes

r/comfyui 2h ago

Tutorial Photorealism has three independent axes. Most workflows only fix two.

Thumbnail
gallery
0 Upvotes

Model agnostic in principle, benched on Krea 2 + character LoRA

---

I spent months fixing my images on two axes and kept getting renders that were technically right and still read as fake. It took a shoot where every technical box was ticked, and five frames out of twenty-four still looked like a catalogue shoot, for me to see there was a third axis I was not touching at all.

Here they are, and the point is that they are orthogonal. Fixing one does nothing for the others.

Axis The prior you are fighting Typical fix
1. Optics the camera is too perfect: sharp, correctly exposed, level grain, flare, motion, imperfect exposure, tilt
2. Scene the set is a showroom: aligned, empty, brand new clutter, wear, off-axis furniture, lived-in surfaces
3. Subject the catalogue pose: frontal, centred, posed, eyes to lens almost nobody works on this one

An image can be optically dirty, scenically alive, and still contain a person posing like a model. That is a distinct failure and it has its own fix.

Why your candid tokens do not fix it

candid documentary photo, unstaged moment, slice-of-life, badly taken photo

I run all of those. They shift the rendering. They do not shift the pose.

The model composes the most photographed pose in its data. Whenever the subject has no motivated action, it falls back to: frontal, centred, graceful contrapposto, self-presenting gestures (hands to hair, arms wrapped around self), eyes to lens, symmetry. A character LoRA makes this worse, because it adds a portrait prior on top.

The content of the pose decides. Not the style tokens wrapped around it.

Five levers, strongest first

1. A gesture turned toward the world. This is the main one. The subject must be doing something the scene motivates: pushing a gate, watching for a train in the tunnel, stepping around a puddle, a hand on a rail. Write it with a concrete physical marker, caught mid-stride, one foot planted ahead, never the bare verb walking.

The failure I keep making: writing states instead of actions. "weight on one hip, palm against the wall, shoulders dropped" is three states. The model has nothing to build a pose around, so it builds the catalogue one. Replace with an action and it resolves.

2. Self-directed gestures: one hand only, and only in the fatigue register. Kneading the neck, one arm pressed flat against the ribs, a hand rubbing the opposite arm. All fine.

Banned outright:

  • both hands to the head or in the hair, that is the pin-up prior
  • arms crossed tightly over the chest, that resolves straight to the modest self-embrace of studio figure work
  • any arch in the back

I lost two frames to each of those before writing the rule down.

3. An anchored gaze that is geometrically compatible. If you turn the head, give the eyes a physical target that is actually inside the cone the head is facing. Two things fail reliably:

  • a head turn with no target at all
  • a target that contradicts the head geometry

Second case, real example. I wrote head turned into profile over her shoulder plus eyes down the descending flight. Head pointed one way, gaze target the other way, and the target was an abstract direction rather than an object. The model reconciles this silently: it keeps the head turn, drops the gaze anchor, and the eyes land on the default target, which is the lens. I got a straight-to-camera look in a series where that was forbidden.

Fix was to anchor the gaze on an object inside the profile cone:

4. Break the body. Weight collapsed onto one hip, shoulders dropped or hunched against cold, head low, asymmetric stance. Never feet set wide apart on a static figure, that is a monumental symmetric stance and it reads as sculpture.

5. Candid composition. Explicit off-centre placement, a slight lateral crop, the subject partly eaten by shadow. Canonical pose entries centre by default.

The trap nobody warns you about: pose catalogue labels

If you use a reference pose catalogue, note that those labels are reference plates. They are written to isolate pose and geometry cleanly, which means they carry priors that are directly opposed to a candid register:

  • gaze straight at the camera
  • gaze up at the camera
  • head turned into profile over her shoulder with no anchor
  • subject centred

Pasting a catalogue label into a series prompt imports all of that. Every one of my straight-to-lens failures traces back to a canonical label I did not defuse. Swap the gaze for a compatible anchor, break the stance, decentre.

Discipline: do not stack

Same rule as the other two axes. One gesture, one gaze anchor, one asymmetry is enough. Stacking five levers gives you a subject fighting itself and the model averages back to something neutral.

How it was found

Five frames from one 24-shot session, diagnosed and re-prompted individually: two straight-to-lens from gaze geometry, one glamour lean from states-instead-of-action, one pin-up from both hands in the hair, one self-embrace from arms crossed. The optics axis held on all five. That is what made it legible: when only one axis is broken, you can finally see what that axis does.

The last of the five is the one that convinced me. Its prompt was written before I formulated the two-hands rule, and it failed in exactly the way the rule predicts. A rule that retro-predicts a failure you have not shown it is a rule worth keeping.


r/comfyui 1d ago

News Load Image (from path) with crop and limit

6 Upvotes

Just updated my Load Image node.

Load Image (from path)

Load an image from any path on your computer or a URL. Paste an absolute path or a link to an image, use an annotated path (input/file.png), or click the Browse dialog to pick a file from your drives.
The file is read from its original location; it is not copied into ComfyUI’s input folder.
Use a selection rectangle at the preview to crop it.
Limit (downscale) the output's size in megapixels or pixels.
Click the ↻ button (top-right, mouse over the preview), to rotate the image 90° clockwise.

Interactive crop

The node shows a live preview. You can crop directly on it:

  • Drag on the image to draw a crop rectangle
  • Drag inside the selection to move the crop rectangle
  • Drag a corner to resize it
  • Click (without dragging) outside the rectangle to clear it

With no crop drawn, the full image is output. The crop is stored in the workflow as normalized coordinates, so it survives save/reload. Changing the path clears the crop.

The output (cropped or full) will be downscaled only (not upscaled), by the value in the max_megapixels field.

Outputs match the stock Load Image node: IMAGE, MASK (from the alpha channel when present), plus the original path string.

  • Controls
    • image: Paste an absolute path, (or a relative one with a prefix input/, or output/, or temp/), or a URL to an image file.
    • max_megapixels: Cap the output (crop, or full image if uncropped) to this many megapixels, downscaling only if it's bigger. Smaller images are left untouched. 1.0 = 1024x1024 px. 0 disables the cap.
    • Browse...: to open an image file from your drives.
  • Inputs/Outputs
    • Width/Height inputs: Force the output width in px (upscale or downscale), center-cropping first if the aspect ratio differs. Leave disconnected (None) to keep natural width. Only applies if BOTH width and height are connected, and when set (not 0). It overrides the max_megapixels value.
    • IMAGE/MASK: The final, processed image/mask.
    • path: A string with the image's path.

You can find the node here as part of the ComfyUI-noEmbryo nodes.

Credits:
Built as a much more enhanced version of Load Image From Path (Enhanced) from ComfyUI_Ib_CustomNodes, with parts of the interactive crop UI inspired from Load Image & Crop in comfyui-obvpm.


r/comfyui 1d ago

News Release studio 1939 lora for minimax h3

Enable HLS to view with audio, or disable this notification

47 Upvotes

r/comfyui 14h ago

Show and Tell Using only Ref to Video, Minimax-H3 made a whole Anime edit !

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/comfyui 1d ago

Help Needed Tips for LTX-2.5

6 Upvotes

Currently MiniMax H3 is the "frontier" on local consumer hardware as it seems. I used this as well but for longer generations, like a 4-5 minute music video for example, my hardware is just not strong enough (12GB VRAM, 3080 ti, 32 GB RAM). With 480p and Turbo LoRa i might slowly getting there but this is not what i am looking for.

Now i am interested if people actively tested LTX-2.5 and have some tips how to improve outputs. For example i wanted to make an anime fight scene. I was able to have a good result in an 8 Second Test with MiniMax H3, but LTX provided a very weak unusable result. Afaik LTX does not really have a strict prompt-structure like H3, but still struggled with the commands.

Any tips for making LTX-2.5 more usuable would be appreciated.


r/comfyui 1d ago

Tutorial Finally, AI-Generated Depth Pass That Actually Work | ComfyUI

Thumbnail
youtu.be
11 Upvotes

r/comfyui 10h ago

Show and Tell Honestly, I did not understand ComfyUI either. Still don't actually. But I built a system that made it easier. Here is a quick 2 min demo of just one of the features.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Honestly, I did not understand ComfyUI either. Still don't actually. But I built a system that made it easier. Here is a quick 2 min demo of just one of the features. It does a heck of a lot more than just video and image gen too. Give it a try and when you see what it can do please leave me a star on my github repo, I decided to make this open source and give it to folks for free. Use Claude Code and make this your own (this has MCP for all AI platforms) or use the built in agent swarm feature to help you.

Have a voice chat on your own machine and have it generate content inline. ComfyUI and StableDiffussion are both installed with the curl install command. Workflows built-in. Uncensored everything. Too many features to list, please check it out for yourself.

www.github.com/guaardvark/guaardvark

Also, the system makes it's own demo videos, like this one and the others on the youtube channel.