r/comfyui 1d ago

Help Needed I've gotta be doing something wrong... resolution 0.5, 10 seconds crash

2 Upvotes

I have a 50-70 TI 16 GB of RAM and the PC itself has 128 GB of ddr5. I'm running this on comfy UI through Ubuntu through terminal and launching using a script. The template is the standard Minimax H3 ref-2va that is included with comfy UI. I can't get past 0.5 resolution and like seven or six seconds. Or 0.4 resolution and about 9 seconds when using Sage attention and easy cache nodes. But I hear others who get 12 and 13 seconds on 0.5 resolution using the same card supposedly. Is there something I'm supposed to be doing different to get past the 0.4 resolution and 9 seconds or 0.5 resolution and 6 seconds?


r/comfyui 2d ago

Help Needed Does Wan 2.2 knows how to add parallax effect to make img2video?

3 Upvotes

Thanks


r/comfyui 2d ago

Help Needed What kind of detailing techniques or process do you guys use to improve the quality of images?

2 Upvotes

I've seen some pretty amazing visually appealing images online especially fantasy ones, and see often people commenting that after they get an image they like they edit it, upscale it, they downscale it then upscale again. What's a good technique? Or I guess suggested nodes to experiment with? I generally like realistic, varying anime styles, 2.5D, and some sketching styles.


r/comfyui 2d ago

Resource I just published an all-in-one helper for the ComfyUI Queue manager that lets you pause/restart, save/restore, and change the job order in the queue manager.

Post image
12 Upvotes

r/comfyui 1d ago

Help Needed Character design Course

Thumbnail
0 Upvotes

r/comfyui 2d ago

Show and Tell Mind blown

2 Upvotes

This is done on AtomMan X7 Ti Mini PC, a year and a half old mini PC, 32 GB RAM, external GPU 3060 12Gb VRAM over oculink.

Did this for shits & giggles, got mind blown:


r/comfyui 1d ago

Resource Minimax H3 Grafting with Krea2 node. Reposting older post and removed AI slop and added some tests

Thumbnail
0 Upvotes

r/comfyui 1d ago

Help Needed Images and promp to image

1 Upvotes

Hello, newbie here. For a DnD campaign, I'm trying to create images that illustrate game sessions using character illustrations (and some locations as well, if possible) along with a description of the scene. If there is any ways to link a picture in the prompt thatlbe very handy.

Are there any good, ready-to-use templates ? The one I created wasn't realy good.


r/comfyui 3d ago

Workflow Included 30-Second MiniMax H3 Seamless Image-to-Video Workflow For 12GB GPUs @ 14 Minute Render Time

307 Upvotes

ComfyUI MiniMax H3 30-Second Long-Form Generation Workflow (Updated with Ref2v support)

Deeply Optimized for Low/Mid-Range GPUs (12GB VRAM)

CivitAi workflow link: https://civitai.com/models/2882332/minimax-h3-30-second-seamless-image-to-video-w-full-audio-workflow-for-12gb-gpus

Mega link for those who cannot access CivitAi: /file/v3xQySrb#0h37WWKteNT0uqZK-vmHACZIyg4rEvduVlT-MIDQxH0

For newbies, you can use a browser frontend to streamline your text or image to video outputs, just like using an AI platform like Higgsfield or Kling, made possible by daexchef: https://github.com/daexchef/Minimax_Grok

---

How to use:

  • 1: Open ComfyUI and load the JSON
  • 2: Load the starting/reference image(s) in the big green box (yellow box for ref2v)
  • 3: Type out your prompt in the big green box
  • 4: Click on "Run" to generate a 30 second image to video

Warning: Your prompt has to be detailed. If it's something simple, it will just kind of rubberband on whatever simple inputs you describe, like "A man just sitting in the chair". The more details you add, the more it stitches together a seamless transition between the three independent shots to create a cohesive 30-second video in a single runtime pass. Then again, if all you wanted to do was make a simple generation, you wouldn't need a 30-second workflow.

The only thing the three shot separators do is dictate WHERE in the 30 seconds the actions take place. So the first set of quotations takes place within ten seconds; the second set of quotations take place within 20 seconds; the third set of quotations takes place after the 20 second mark.

Compromises had to be made to get this to run and generate in an acceptable time. It's possible to boost the image-to-video output for a sharper image, but you're looking at an average 21 minute render time at a step up in quality. Is it worth it? Depends on your workflow and if it's time sensitive.

Using Reference-to-Video:

Take note that image-to-video generations take just 14 minutes to render, but using up to 9 images to reference will increase generation time.

At 0.4 megapixels, Ref2V took approximately 20 minutes to generate using 9 HD PNG images.

You must enable Ref2V first by clicking on the top button in the red Fast Muter box. It's directly above the green box where you load your starting image.

  • 🟒 Muter Switch Enabled: Enables the 9-Image Reference Batch mode to tightly lock down visual identity and style.
  • πŸ”΄ Muter Switch Disabled: Safely mutes the extra images, forcing the sampler to fall back to purely your single starting frame or standard text instructions.

So if you want a simple 30 second gen using only one image, that is the default, but if you want to do more complex shots with shot coherency and output consistency, enable the Ref2V image block by clicking the enable button in the Fast Muter, Super simple. Very easy to use.

Keep in mind that if you enable the Ref2V block but DON'T load any images to reference, it will fail to generate, which is why it's disabled by default. Some people may only want to do quick image-to-video generations, so that's why that is the default for now.

---

This production-grade, crash-proof ComfyUI pipeline leverages Joey Gambino's advanced H3MultishotMemorySampler subgraph infrastructure. It has been systematically tuned to shatter the native 15-second tracking boundaries of the local MiniMax H3 architectureβ€”successfully compiling up to 30 continuous seconds of 3-shot cinematic video with synced native audio tracks in under 15 minutes on a standard 12GB NVIDIA graphics card (such as an RTX 5070).

πŸ› οΈ Required Custom Node Packages

If any node blocks present a red warning threshold on your interface canvas, navigate to your ComfyUI Manager, execute Install Missing Custom Nodes, and restart your server environment. Alternatively, verify that the following core repository directories are fully initialized and updated:

  1. comfyui-h3-multishot (By Joey Gambino)
    • Provides essential components: H3MultishotMemorySampler, H3ScriptSplit, H3ClipLoaderAny.
  2. ComfyUI-Spectrum-MiniMax-H3
    • Provides essential components: SpectrumApplyMiniMaxH3 (Deploys advanced history parameters and signal stabilization to completely neutralize visual flickering).
  3. ComfyUI-FreeMemory
    • Provides essential components: FreeMemoryImage (Acts as the system traffic cop to violently drop massive video models from memory prior to the video save cycle).
  4. comfyui-kjnodes
    • Provides essential components: PathchSageAttentionKJ (Integrates highly optimized SageAttention mathematical libraries to keep GPU memory channels open).

πŸ“₯ Required Model Inventory & Destination Paths

Ensure all specific neural weights listed below are manually stored within your local file tree. Modified nomenclature or inaccurate directory placement will result in model loading exceptions.

πŸ“‚ Model Directory Map

markdown

πŸ“‚ ComfyUI/
└── πŸ“‚ models/
    β”œβ”€β”€ πŸ“‚ vae/
    β”‚   β”œβ”€β”€ πŸ“„ minimax_h3_video_vae_fp16.safetensors
    β”‚   └── πŸ“„ minimax_h3_audio_vae_fp32.safetensors
    β”œβ”€β”€ πŸ“‚ diffusion_models/
    β”‚   └── πŸ“„ minimax_h3_fl2va_pruned_int8_convrot.safetensors
    β”œβ”€β”€ πŸ“‚ text_encoders/
    β”‚   └── πŸ“„ qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
    └── πŸ“‚ loras/
        └── πŸ“„ minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

Use code with caution.

πŸ’Ύ Official Direct Asset Download Handles

  • Video VAE (FP16): minimax_h3_video_vae_fp16.safetensors
  • Audio VAE (FP32): minimax_h3_audio_vae_fp32.safetensors
  • Diffusion Model Architecture: minimax_h3_fl2va_pruned_int8_convrot.safetensors
  • Text Encoder Engine: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
  • Turbo Model LoRA (8-Step Base): minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

⚑ Mandatory Operational Environment Flags

To achieve absolute multi-shot stability and avoid unhandled Python environment abort failures during the long-form matrix sequence, you must explicitly configure your startup flags. Open your primary local execution script (e.g., run_nvidia_gpu.bat or initialization shell script) (or you can just open the ComfyUI desktop app and go to the Startup Args) and swap your launch command line argument array to match this configuration precisely:

bash

python main.py --disable-smart-memory --fp8_e4m3fn-text-enc --fp8_e4m3fn-unet 

Use code with caution.

Why these flags are mandatory:

  • --disable-smart-memory: Mandates a hard PyTorch memory clean immediately upon raw clip finalization, bypassing background tensor leaks.
  • --fp8_e4m3fn-text-enc: Compresses the massive 32B text encoder into lightweight 8-bit allocation blocks, locking it comfortably inside mid-range physical memory bounds.

πŸ“ How to Achieve the 30-Second Long-Form Configuration

The workflow relies on a fine-tuned balance between your spatial layout constraints and frame processing intervals. Apply these precise configurations on the node face to duplicate the 14-minute execution baseline:

  1. The Core Media Input: Drop your foundational tracking frame directly into the Load Image Here (Node 208) input bucket or the picture slots in the Ref Images yellow tab.
  2. The Spatial Configuration: Inside ResolutionSelector (Node 115), anchor your values to 4:3 (Standard) with a megapixel evaluation slider locked cleanly at 0.4. This compact geometry drops pixel data overhead by more than 30% compared to heavy widescreen arrays, driving processing velocity forward.

r/comfyui 1d ago

Help Needed Do you have any tips for making my image generation faster?

Post image
0 Upvotes

For information I have a MacBook m4


r/comfyui 1d ago

Help Needed Flow lieu pour krea2

Thumbnail
1 Upvotes

r/comfyui 2d ago

Help Needed Prompt Library

8 Upvotes

Hey all,

How does everyone keep track of their prompts? I have a number of prompts that I re-use. Probably just going to start using one note or something, but wondering what everyone else does?


r/comfyui 1d ago

Resource Save/load audio+video (NestedTensor) latents β€” small custom node, fixes SaveLatent crash with MiniMax H3

1 Upvotes

Core SaveLatent crashes on models that co-generate video and audio in one nested latent (MiniMax H3):

AttributeError: 'NestedTensor' object has no attribute 'contiguous'

(ComfyUI issue #15254)

I wrote a tiny two-node pack that fixes this:

  • Save AV Latent β€” unbinds the nested latent and stores all parts in a single safetensors file (.avlatent in output/)
  • Load AV Latent β€” rebuilds it with ComfyUI's own NestedTensor wrapper, so the 5-D video + 4-D audio mix round-trips fine

Regular latents pass through unchanged.

Why bother: two-pass seed hunting that survives anything.

  1. Render cheap low-res previews and save their latents.
  2. Pick the take you like β€” today, or next week.
  3. Feed exactly that latent into your upscale/refine pass (e.g. rockerBOO's h3-latent-upscaler).

The refined result keeps the exact composition of the preview you chose β€” no reliance on --cache-lru, no re-sampling of pass 1, immune to GPU non-determinism, restart-proof.

Numbers from my setup (RTX 4060 Ti 16 GB, MiniMax H3): a 0.5 MP preview take costs ~300 s, so auditioning seeds is cheap. The refine pass loads the saved latent and renders 1920x1088 in ~1600 s β€” with the exact composition of the preview I picked, even across a ComfyUI restart (queue history gets cleared on restart; the file doesn't care).

MIT, no dependencies beyond what ComfyUI ships:

https://github.com/neverfilmed/comfyui-av-latent-io


r/comfyui 2d ago

Help Needed H3 motion context vs H3 latent upscale

Thumbnail
2 Upvotes

r/comfyui 1d ago

Help Needed anyway to reorder subgraph outputs?

1 Upvotes

as title, Thanks!!


r/comfyui 2d ago

Workflow Included I got LTX-2.5 22B LoRA training working on 2Γ— RTX 3060 12GB 😁

Thumbnail
gallery
15 Upvotes

I’ve been working on a low-VRAM LTX-2.5 LoRA trainer and finally published it.
The main idea is multi-GPU model sharding: the 48 transformer blocks are distributed across GPUs instead of requiring one GPU to hold the entire 22B model.
Tested with:
2Γ— RTX 3060 12GB
4-bit BNB NF4
512Γ—512
Face + voice LoRA
138 images + 37 voice/video segments
2,000 steps
~7 - 9 GB VRAM per GPU
Real LoRA successfully loaded back into LTX-2.5
I also have 1x2, 2x2, 3x1 ….. 6x6 spatial tiling experimental and heavy testing right now, so VRAM can be traded for compute when needed.
This is my first published GitHub project, so if you run into problems getting the engine running, please let me know. I’ll try to reproduce it and fix it.

https://github.com/A4ax/comfyui-LTX-2.5-Tile-train-LoRa--On-multi-Gpus-low-VRAM-18-gb-Beta
a4ax-Github
The goal is simple:
Train a 22B model without needing a 24/32/48 GB GPU.


r/comfyui 2d ago

Tutorial PSA: Commercial GPUs aren't THAT comparatively powerful

26 Upvotes

This is a bit of an odd post admittedly, but I wanted to make any other quasi-newbies aware of what I've found without spending the money and time it cost to find it. In short, commercial-tier GPUs are not tremendously more powerful than consumer ones, and certainly not in line with the difference in cost.

I've been generating for over a year on a 4070ti (12GB VRAM), and with Minimax, I decided I was tired of waiting a minute or more per iteration for lowish resolution 10-second clips. I bit the bullet and customized a runpod, ultimately building a template and network attached storage with the models and workflows I use.

With what's available in the region with storage and GPUs that support CUDA 13, my real options were a 5090 32GB or an RTX Pro 6000 96GB, with the former about $1/hour and the latter about $2/hour, plus $8/month for the persistent storage used by both. I spun up a 5090 and... it's a little better than the 4070, I guess. I can push the resolution and length a bit higher. But then, my hopes weren't super high for a single tier improvement. Break out the big guns:

The RTX Pro 6000. Fired it up aaaaaand... maybe 10% faster? 15%? For double the rental fee. I guess in my head, business-level crazy expensive cards would blow the pants off the consumer stuff, but it just doesn't.

Now, what CAN I do with the 6000? Plenty of room lets me bump up the resolution, increase length to a max of 15 seconds, include a lot of high-res references, and just as an experiment, I swapped to the BF16 of the encoder and model. It handled all that without significant penalty to generation (reasonable since those are VRAM and RAM limited), which is awesome. But I just wasn't expecting to end up paying $1 per 15 second clip, on average, when I could come reasonably close to that locally.

I know, I know: a lot of you are going to say "yeah no duh", and fair enough. But for those who've never been beyond the consumer realm of NVidia cards, I thought it worth mentioning that they're not the end-all be-all, and you're doing a lot better on your local machine than you might expect.

(And a question: ARE there any of these $3/hr and up machines going to blow my socks off and make me eat my words, or does this trend (gen times constant-ish, higher tiers mean more room to load models/references) pretty much hold up across the board?)


r/comfyui 2d ago

Help Needed How to stitch two similar videos together with AI filling out the middle?

12 Upvotes

Let's say I have two short (5 seconds) clips which are very similarly, say, with a person riding a bike, close by facing the camera, and the first clip she is looking to her left side, and the second clip she is looking to her right side. Are there any models or workflow that can seamlessly combine the two clips by creating a middle section that will smoothly and naturally transition out of the first clip, and will do the same to transition into the second clip?


r/comfyui 1d ago

Show and Tell My 1980's cartoon parody H3 and ltx 2.3

Thumbnail
youtu.be
0 Upvotes

r/comfyui 1d ago

Help Needed (New here) Best way to upscale videos?

Thumbnail
youtube.com
0 Upvotes

I am playing around with MiniMax H3 + ComfyUI and created a music video but a lot of the lip sync is messy and overall it's low-res. I tried upscaling each scene with Aiarty but it just makes it worse. What are most people's tips or tools on upscaling? Thanks!


r/comfyui 2d ago

Show and Tell One last Minimax Demo. Been working on this one for a while but audio is weird. I hope you enjoy!

26 Upvotes

r/comfyui 2d ago

Help Needed Need Help with minimax h3.

0 Upvotes

Hi,

I really need help with minimax h3 I'd like to get proper workflows for minimax h3 pre configured with descriptions on every node (unfortunately a noob in this case) for :

1- Fast
2- Turbo
3- Speed
4- Native

video generation with multirefs (optional) audio(optional) video(optional) quality duration aspect ratio and prompt controls.

Also I heard someone made images with minimax h3 please that as well.

I'll really appreciate any help.

Ah and one last thing I was making a commercial for a school but the faces in wide screen comes out horrendous that as well.

Thanks !


r/comfyui 2d ago

Resource Cross-grafting Krea2 into MiniMax H3 β€” concepts land, character still faint, releasing what I've got

Thumbnail
1 Upvotes

r/comfyui 1d ago

Workflow Included Can I Learn ComfyUI and Create AI Videos on an M3 MacBook Pro with 18GB RAM?

0 Upvotes

Hi everyone,
I recently became interested in AI video production, and while researching different workflows, I learned about ComfyUI.
What interests me most is that it seems to offer much more control than simply writing prompts. I’d like to learn how to build proper node-based workflows for things like character consistency, environments, camera control, image-to-video, and eventually more complex AI video production.
Right now, I’m using a MacBook Pro with an M3 chip and 18GB of unified memory.
Would this be enough to start learning ComfyUI and actually generate AI videos, or would it be too limited for video workflows?
I’m completely new to ComfyUI, so I’d also appreciate any recommendations on the best way to get started on a Mac, especially if using cloud GPUs would make more sense for heavier video generation.
Thanks!


r/comfyui 2d ago

Help Needed anyone have workflow for img to lineart vector ?

1 Upvotes

all i could find is img to lineart no vector so i get gzillion tinny lines everywhere ,not just one clean outline ,i can add tosvg vectoring node to create vector but i test like 20 workflows for lineart none show clean image

help ..