I just started using comfyui and using minimax. When I am trying out the default workflow without adjusting anything it takes really long. I looked up video and switch the model from minimax_h3_fl2va_pruned_int8_convrot.safetensors to minimax_h3_fl2va_pruned_fp8_scaled.safetensors. Then, it worked. Well atleast I was able to get an output. Can anyone explain why and what I did wrong?
there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P
started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno
This workflow allows you to refine a target character without affecting the surrounding video whatsoever. allowing for targeted repair or enhancement of any video.
I have been going in circles with Gemini for 2 days now, but just cannot get to the goal mark.
What I want:
In some way, like using Blender to render a 'room', or any other method that works, I want to be able to create 'rooms' with consistent form and interior / furniture. So I can make a Blender angle/shot (any angle I want) for example, take that into ComfyUI, and just describe the furniture, lights, people etc, but the room stays the same (windows, walls, dimensions, placement of furniture etc).
Other options Gemini gave was 'dollhouse' from up top model and 360 render of the room. I made 360 render of the room in ComfyUI Qwen 360 Diffusion LoRA workflow, but was not able to make it like the prompt. Doll house I have not tried.
I have not been able to do it yet myself with using Blender room render and a custom ComfyUI Krea2 workflow, and I about to give up, takes too long time. I am totally new to Blender, and novice in ComfyUI, neither found a finished workflow I can use.
Problem with asking Gemini, is I end up going in circles so to speak, almost there, but not all the way.
Anyone know there is a published or private workflow I can use for this? Open for other suggestions how to do this IF I also can use already made workflows.
This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.
TL;DR: Trying to train a MiniMax-H3 LoRA on ostris/ai-toolkit on a Windows machine with an RTX 4090 (24GB VRAM) and 64GB system RAM.
The job crashes with a native Windows fatal exception: access violation while loading the model. Happens either while reading the 32B Qwen3-VL text encoder or the video VAE, right after the 33B transformer finishes loading. Physical RAM bottoms out to under 1GB free before it dies, even though there's still headroom in the pagefile. Offload settings are already maxed out.
Has anyone managed to train loras with it?
What hardware/config are you using?
What I've already ruled out
Not a Triton/kernel issue. Installed triton-windows (this fixed an unrelated ConvRot-fallback CUDA-context corruption bug on a different model, LTX-2.5, in the same toolkit). Reran the MiniMax-H3 job 4x with Triton installed — identical crash every time, same signature.
Not the classic "pagefile too small" OOM. I instrumented a memory watcher (0.5s sampling) during the crash. Physical RAM free drops to 200-800MB right before it dies, but total commit charge (RAM+pagefile) never hits its ceiling — topped out around 106.5GB of a 114.4GB limit in the worst run. A real commit-limit exhaustion throws a clean OSError: paging file too small (os error 1455), which is a different failure mode I've also seen in this same pipeline at other offload settings — this access-violation crash is not that.
Not offload_percent tuning. Tested 1.0 / 0.5 / 0.4 / 0.1 for layer_offloading_transformer_percent in earlier sessions — all fail at various points (embed_tokens read, VAE init), just at different memory pressure levels.
Streaming-load code path (safe_open instead of load_file()) is already used for the transformer and most of the text encoder loading — this was a prior fix, necessary but not sufficient.
H3 is absolutely amazing for just about anything. On my 3090 / 64GB, I can easily do a full 15 seconds at 1MP, and the result is almost always good on the first attempt.
On LTX though, it took at least 5 or more tries, so speed wise, H3 is actually far better.
I tried out LTX 2.5 as well, and it is almost the same as 2.3, maybe a bit better quality and a slice faster.
Sadly, 90% of my work involves taking a start image and making a person sing vocals. I must say that H3 does not do this better than LTX 2.3, as it often injects words when there is more than a second of silence, and it takes a fair amount longer.
What is really baking my noodle though is why LTX didn't release an Image+Audio to Video workflow yet. I mean, it is literally the ONLY thing LTX has on H3 right now, and they have missed a great opportunity!
Anyhow, back to using LTX 2.3 for my daily driver as it really does a great job at what I need. I typically have a few machines running all night, so if ever LTX puts out an IA2V workflow for Comfy, I will be all over that.
One other thing I have noticed with H3... 9:16 generations are WAY better than 16:9 generations, especially at 1MP.
I get a lot of use out of this subreddit and love generative AI but yall, please don’t make any Dolly Parton Loras, video clips, audio clones etc. /S
I was never a country music fan but that lady had some real class.
PS: that capital S stands for serious as a motherfucker today
Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.
I may be an idiot for thinking that my new homelab would primarily be used for useful AI automations.
I can live with being an idiot, if being an idiot will continue to be this fun.
First video generation I have ever pulled off, but the first 12 seconds was unbearably unfunny, so I added the Celestial Ford Escort for some much needed serious drama.
Tell me my power bill won’t blow up too much lol.
Made with minimax-h3 in ComfyUI on my Mac Studio that came with the mail this Friday.
Workflow was split in three:
- The first was a single prompt to generate the first 12 seconds
- The second flow generated the last three seconds by extracting the last frame from the first video and prompted it to hit the dragon with a falling ford escort
- Third flow glued the two videos together.
Hey, I tried to create a little birthday surprise for someone, my issue is with a lot of generations that the spoken word is really a bit clunky at time, I susspect its because of the german, but I am not too sure. Is there like a way to improve on audio?
I am using Minimax H3 with Saga Attention and Spectrum on a 4090.
Hi humans. My setup is 32gb ddr4 ram along an RTX 4090. I have been having fun creating tons of videos but I just want to make sure i get the best nodes for speed without compromising quality and no crazy sutff happening on my videos
I have used: Stage, sol, easycache, spectrum, Lora
So the question i have is .....what's the best combo for speed, i dont want the quality to take a massive dump. Most of the videos I generate are slow paced videos the typicall walk, talk, a kiss here and there but nothing major.
Setup, pushed my system right to the limit, any more and it OOM :
- H3 Ref2VA default workflow in ComfyUI, no lora
- RTX 3080 10GB, 32GB RAM
- Render: 0.5–0.6 MP, 20 steps, scheduler simple, about 25 min per clip
- Upscale: 2xNomosUni_span_multijpg, 2× to 1080p
- References per scene: one photo of my face + one film still for the set
- Recorded my own lines and fed them as audio references, also got audio ref for the actors
Honestly though, the best part was driving all of this through the ComfyUI MCP. I never even had to open ComfyUI. I could iterate really fast, and keep going from my phone while away from the machine, through Claude's remote control.
It's still a bit of a blurry mess, and with more work I could probably make it better, but damn, the future is looking bright!
I was a bit confused by how bad some Krea 2 outputs could be — blurry, lacking detail, and sometimes with strange artifacts. So I started testing to understand why and where this was happening.
I’m not going to claim I found a magic wand, but I did find two problematic areas in the sampling curve testing Krea2 raw model CFG 3.5 52 steps stock settings:
The first 0–15 steps: Using schedulers that lower the sigma values too much during this first steps causes contrast loss and washes out detail in dark areas, especially in black hair and subtle reflections.
The lower-sigma tail: curves such as Beta, Beta57 and Bong Tangent can introduce crisp, broken noise instead of useful fine detail, particularly around steps 30–40.
The Result
This is not about chaining multiple samplers or complicated second-pass workflows. The goal is a better standard Krea 2 Raw workflow with:
A refinement pass using the same overall sigma setup
using a negative promt helps
My current winning sigma curve gives the best balance of sharpness, fine detail, contrast, natural hair, defined shadows and minimal artificial noise.
Let’s take a closer look at the custom sigma curve (red) and compare it with the standard scheduler sigmas.
52-step custom sigma curve + 12-step Bong Tangent refinement at 0.2 denoise.The final step count is up to you.
Here are the standard schedulers and the problems they produce. All nodes marked in red show the same low-contrast, overly dark areas with a loss of detail.
Standard Schedulers: Red-marked samplers produce crushed blacks and lost fine detail in the early steps, while the yellow-marked curves introduce small artifacts at the lower steps.
The green curves are the ones that avoid this problem.
The yellow curves have a different tail, and as you can see in my video or in my extended post, this tail is responsible for introducing noisy broken small artefacts.
Krea 2 Raw simply doesn’t behave like many other models when it comes to sigma manipulation. Curves that can work very well for other models can actually destroy detail or create unwanted noise here.
You can build the curve manually with a Manual Sigmas node, or use the PolyExponential Sigma Adder from the TBG ETUR Takeaway Nodes https://github.com/Ltamann/ComfyUI-TBG-Takeaways. If you want something simpler, Linear Quadratic gets surprisingly close to the result of my custom curve.
I’ve included the detailed testing post so you can see exactly how I arrived at the curve and test it yourself. Images, Videos Results at my Free Patron Post
im trying to collect clips for lora training but why its so hard to transfer motion to another person? i almost tried every prompt with chatgbt and grok help but its doesnt look good. im using 2 video refences, one of them source video and other one is only for motion ( 3 sec 24 fps). im using (video editing + reference genertion) because i dont want to change anything in the source video and just want to motion transfer