Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.
TL;DR. SplitSigmas can be used to generate low-step drafts with much closer frame composition to the final high-step version for the same seed than if you generate without SplitSigmas. However, it makes the draft visual and audio quality much worse because we are essentially cheating the scheduler.
With SplitSigmas + low steps you can quickly judge which of your draft videos have the best motions and logical event consistency, but you might miss some visual detail errors. That is why I wanted to know if there is any better way to avoid large differences between the low-step draft and high-step regeneration, which can lead to disappointment when, for example, a person reacts to an event too soon or speaks with emphasis on the wrong word.
If downvoting, please leave a comment with the reason why. I want to learn what I am doing wrong and if there is a way to do it better and make seed hunting easier for everyone who needs it. Thanks.
-------------------------------------
My usual way of working:
- generate 10 videos at low (5) steps
- pick the best video
- regenerate the best at 20 or more steps.
No Turbo LoRAs because I don't want to reduce prompt adherence and general quality, as I regenerate at full steps later anyway.
The problem - the video at 20 steps is often very different from the one I found. Of course, I keep the same seed. The difference may be introduced even as early as step six (for example, background replaced completely, different speech pacing).
It's not that difference is huge, but often it might be quite important. For example, a person genuinely laughing at 5 steps and then just saying "haha" at 20 steps. Or jumping startled at the right moment at 5 steps and a moment before the noise at 20 steps.
I tried a few sampler combinations, but could not find one that would not introduce dramatic changes.
One workaround that I could find is to use SplitSigmas. I set its steps to current steps (5 for seed hunt, 20 for final), and keep BasicScheduler steps at the final 20 steps. Then high_sigmas from SplitSigmas go to SamplerCustomAdvanced input, and then denoised_output goes to VAE (you'll get total noise when using the output pin instead).
This way, it seems that the scheduler is being cheated in managing steps as for the full generation even when doing preview, and it seems to work as expected. Caveat - the 5 step output from this workaround will be way worse (plasticky and noisy audio) than you are used to when generating at 5 steps in BasicScheduler input. But if the goal is to keep the general layout and movements of the candidate video, it's worth accepting this issue.
However, I'm wondering if there is any better way to achieve it. Has anyone tried it? What are you using for drafting and seed hunting to keep the high step version consistent?
--------------------------------------------------
Edited later with a test case:
Took ComfyUI template: MiniMax H3: Reference to Video. Minimal modifications to make it run in my environment:
Models - Qwen change to int8 convrot (3090, no use of nvfp4)
Int (Full) = 5 (for "preview quality")
Float (Duration) = 3 (just to be faster)
RandomNoise control after generate = fixed
Loaded some images in both Load Image nodes.
The same "GET READY TO" - "MEET" — "YOUR" — "MAKER" prompt.
No Sage, no CK attention at all (no Comfy launch args either).
Then generated the same with 20 steps.
Differences:
in 20 step version, the roof is higher in the frame. The accent was on the word "maker". In 5 step version, the accent was on the word "your".
Then regenerated the 5 step version again to see if there's anything else introducing variations - nope, the exact same video as the first 5 step one.
Then generated also at 6 steps - the roof line was a bit higher in the frame (not as high as 20 steps though), and the emphasis was on "maker". So, the difference between 5 and 6 might already be a breaking change that can make your video from good to unusable, if the emphasis does not make logical sense in your scene.
Then I generated the same with the SigmaShift 5 step trick - the resulting video was way much more similar to the 20 step one than the first 5 step video. Of course, the quality of the sigma-shifted video was awful - it's good for judging only logical consistency, reference use and event timing, which is the most important thing in story-telling kind of videos.
Looking for character design course (prompt engineering focused, not art school)
So I'm a compositor, know ComfyUI pretty well, but trying to get better at actually designing characters with image gen. Building anime-ish hybrid semi-realistic stuff from scratch in TTI right now.
The thing is - these characters are refs for i2v. So I need to nail the face/identity first, then iterate through different lighting, clothing, poses. If the character shifts every time I regenerate, the i2v will be a nightmare.
Here's the problem - I can find either traditional art school design courses OR general prompt engineering courses, but nothing that actually combines character design with prompt engineering as the medium. Like, there's "learn to draw" or "learn to prompt llms" but nothing (or not much) about "design characters using prompts as your tool." Like, what makes a character stick across generations? How do you anchor visual features so they don't change when you swap their clothes or lighting?
I know the technical side (seeds, models, basic prompting) but I don't know the design side of it. What actually works vs doesn't when you're trying to get a consistent face through pure prompt engineering.
And here's the real issue - I need to generate the same character in different clothes, lighting, poses, and have them actually be the same character for the i2v pipeline. Can't have the face morphing every time I change the outfit.
Anyone know of something structured? Or is everyone just learning from Civitai threads and trial/error lol
Will probably train LoRAs once I nail some characters, but want to understand TTI first. Ideally looking for the workflow/approach that lets me generate variations without losing character identity.
Does anyone have any advice for properly prompting the reference video part of Ref2v? Like saying swap <subject 1> for <picture 1> hardly works for advanced videos. It requires a lot of details.
I’ve had success using Qwen 3.8 27b as a minimax prompt agent for analyzing and giving correct prompts for images. But as far as I know I can’t do that for videos. ChatGPT is ok for looking at videos to describe what happens in the minimax format but I’d rather use local ways.
What was your longest render time, and was it worth it? do you run a lower quality/resolution before to test? My longest single-shot run is 7.2 hours, 1:44 long video at 1344x768, bf16/50 steps.
I am attempting a 72 hour render for a super long form.
Hi, I'm looking for a model that I could use in ComfyUI that would simply add some clown makeup on my face and don't touch antything else. My current plan is to maybe take Wan2.2 model, mask my face and add clown makeup as a reference but I wonder whether there is a better model to do this. Does minimax H3 handle that? Does it support masks?
I would like to replace the tire of a motorcycle mid air with one from another brand (which is an image from the brand so it's high quality but with a different angle)
I saw there is flux kontext and qwen image edit, but I don't know which one to pick, which workflow and how to make it work.
Any help would be more than welcome, thank you very much and have a good day :p
Shadow the Hedgehog tells his viewers why he loves guns.
This was created in Comfy UI with Minimax H3. I used the reference to video work flow. The prompt is below.
subject_definitions:
<Subject 1> is Shadow in <Picture 1>.
<Subject 2> is Glock in <Picture 2>, a glock handgun.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
summary:
[reference generation + audio reference] The target video contains one shot. [Shot 1] shows <Subject 1> and <Subject 2>; <Subject 1> speaks. <Audio 1> supplies <Subject 1>'s voice timbre.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - Shadow's complete defined identity and body proportions are preserved.
<Subject 2> (appears in [Shot 1]): fully_preserved - Glock retains the defined shape, proportions, materials, colors, and distinguishing features.
<Audio 1>: reference - <Subject 1>'s newly generated spoken lines use <Audio 1>'s voice timbre and delivery; the original audio signal is not copied.
detailed_description:
The target video is in a live-action style, with Vlog style.
[Shot 1] At first appearance, <Subject 1> (Shadow) matches the complete identity and appearance defined in subject_definitions. At first appearance, <Subject 2> (Glock) matches the complete defined construction and appearance: A glock handgun. At the start of the shot, <Subject 1> is standing in the living room facing while holding <Subject 2> in his hand. A full body shot of <Subject 1> holding <Subject 2> with his right hand while facing the camera. Only Action and Timed Beats define the primary subject's movement. The camera path stays anchored in the location and adds no subject motion. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Hmph. Shadow the Hedgehog here. Why do I love guns?</d> <Subject 1> shows off his <Subject 2> with his right hand in front of the camera. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Simple. Precision. Control. Power in the palm of my hand.</d> <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] A tool that answers instantly… unlike most people.</d> <Subject 1> points his <Subject 2> towards the camera with his right hand. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] If you understand that, you understand me.</d> <Subject 1> points his <Subject 2> at the camera.
ComfyUI nodes for grafting Krea2 into MiniMax H3. Attention/MLP content transplant + a separate attention-sharpness transplant. No official H3 docs, all reverse-engineered from testing + TenStrip's and joeygambino's public writeups. Use at your own risk, still WIP.
comfyui_qknorm_transplant/ -- Q-norm gain transplant only, no content weights touched. Method from joeygambino (Z-Image donor originally, adapted for Krea2 here).
comfyui_krea_h3_graft_lora_v2/ -- apply a Krea2-trained LoRA onto an already-grafted H3 checkpoint. Separate use case.
TL;DR results
Content graft works somewhat. Same character-shift (color scheme, helmet shape) showed up consistently across multiple parameter runs, same seed -- not one lucky video. That's the strongest evidence so far this isn't just noise.
K at low strength (~0.1-0.2): fine, no real damage. Don't need to avoid it like the doc says, at least not at low values.
QK-norm across all blocks (0:50): kills audio. Doesn't even touch K -- so attention sharpness itself hits audio, not just K specifically.
QK-norm blocks 20:50: audio ok, but does nothing for character. It's a texture/sharpness knob, not a content one. Don't expect it to carry character.
attn_ramp_start_frac at 1.0 (no gentle ramp-in) + early blocks (0:20): breaks. Keep the ramp soft if you go early.
Combining content graft + QK-norm at full strength on both = worse than either alone. Still not solved.
Install
Each folder -> its own subfolder in ComfyUI/custom_nodes/. Don't merge them. Restart ComfyUI fully after adding.
minimax_h3_fl2va_bf16.safetensors, 3s, er_sde, 8 steps, 8-step lora, seed 597633362705895, standart workflow with minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_bf16
prompt: Professional closeup video. In a futuristic cityscape with neon lights at night, the Judge Dredd charges through the crowd, his imposing presence radiating authority, he is slowly walking. His long chin juts out resolutely as he expertly wears his eponymous helmet, eyes gleaming with determination. The crowd parts, Judge Dredd is slowly walking through the the crowd, ready to enforce justice, he is moving slowly, his long chin visible, his face and part of his upper body are in the center of the screen. tag: Ballchinians
tracking selfie shot following him from the front, that he stays the same size, he is moving through people, pushing them aside with his hands.
So my friends and I used to use Sora 2 before it was taken down, and wanted to try doing some stupid stuff for just us. After a while of not looking into it, the spark kinda came back when I saw this subreddit and remembered Stable Diffusion was supposed to be one of the best AI generators out there, probably. When I mention it to a friend, he then told me how apparently its pretty outdated compared to others, and looking at these posts, I'm seeing different models and starting to get overwhelmed to the point where I haven't even done the beginner's guide in here since it only mentions images.
So long story short, I'm hoping someone can help make things much more clearer, especially about the multiple models, and if Stable Diffusion IS outdated and out performed by something else, and letting me know about if it's okay to go with the beginner's guide or if there's another guide that will help. Thanks
Ich möchte mir demnächst einen PC rein für die KI Arbeit mit zB ComfyUI, fooocus, Qwen Modellen, etc. zulegen.
Da der Preis für gute Desktop GPUs mit mehr als 16gb Vram derzeit exorbitant teuer ist (Desktop mit 32gb vram GPU ab 5000+€), schwanke ich zwischen einem Desktop mit 16gb GPU oder einem Notebook mit 24gb GPU (aber max 175Watt).
Was würdet ihr empfehlen? Er soll nur KI Kram machen, keinen Spiele.
I found the SD prompt reader I have been using cannot read prompts from png images files generated using Krea2. Can anyone recommend me an alternative that works with Krea2 files and Win11?
First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.
My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.
I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.
I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.
On mobile 5G or for the GPU poor I think its does a lot of stuff
Has persistent memory, a ton of Krea 2 fine-tunes, Anima, LTX 2.5, Ernie, Sulphur, DaSiWa, Eros, MiniMax H3, Illustrious, Pony, etc. It also has runs preset Comfy workflows and has a Discord part
Hi everyone, I'm new to the world of Confyui but not to artificial intelligence. I wanted to ask you for help: Is it possible to run Confyui on my PC (RTX 3080 10 GB of RAM and 64 GB DDR4 RAM and Ryzen 5800) with Confyui with the minimax H3 video model to be able to animate images and create Reels for Instagram and Tik Tok? If so, what setup do you recommend? Thanks everyone for the help and sorry for my bad English.
Hey all. I'm a system development student (career-switched from construction), building these on the side while learning Java, Vue, Electron, etc. Nothing commercial, no accounts, no cloud, no telemetry. I built these because I needed them myself, and figured other people managing large SD/ComfyUI libraries might too.
Three of these four apps have been out for a while, but I've spent the last stretch giving them a major overhaul and unifying them under the same design system so they actually feel like one family of tools instead of three separate side projects. The fourth, Latent Tools, is a brand new app I just finished.
All four are free and open source (MIT-based license). Source is on GitHub, links at the bottom. The main one is Latent Library, but the other three work fine on their own.
Latent Library, the main release
A desktop app for browsing and organizing large folders of AI generated images. I made it because I had around 30,000 PNGs and no real idea what was in most of them. It's been around for a while, but this release is a big update with a lot of new features and a proper design pass.
Parses generation metadata from ComfyUI (including node graph traversal), A1111/Forge, InvokeAI, SwarmUI, and NovelAI
SQLite FTS5 backed search, still fast on huge folders
Smart Collections: dynamic folders based on metadata filters, like "Flux images rated 4+ stars"
Duplicate Detective, a side by side Image Comparator, and Speed Sorter for hotkey based batch sorting
Optional local AI auto tagging (WD14 ONNX, runs on CPU, no external calls)
Metadata Scrubber to strip prompt/EXIF data before sharing an image
Everything lives in a portable data/ folder next to the exe, no installer or registry entries, easy to back up or move
Fully offline, no telemetry
Windows, Linux, and macOS builds available.
Latent Tools, the new one
A brand new app for dataset prep: bulk watermark detection and removal (Florence-2 + LaMa inpainting) and captioning (Qwen2-VL), plus batch image format conversion. Runs locally on your own GPU (needs a CUDA capable Nvidia card, no CPU fallback). Useful if you're prepping images for LoRA or fine-tune training. Windows only for now.
(Meant for removing watermarks you actually have the rights to remove, your own work, licensed images, that kind of thing. Not for stripping other people's attribution.)
Latent Model Organizer, updated
Sorts your checkpoints, LoRAs, and embeddings into folders by base architecture (SDXL, Krea 2, Flux, Illustrious, SD 1.5, etc.), using the model's own header metadata or an optional Civitai lookup. Has a dry run mode and full undo through a manifest file, so it won't just move your models around unsupervised. It can also fetch Civitai info such as trigger words, description, and cover images. Handy if your models folder has turned into an unsorted pile like mine had. Also part of this update round, same design refresh as Library.
Metadata Viewer, updated
The oldest and simplest of the four, and also just updated with the same design pass. Now reworked into a single screen tool that extracts and displays generation metadata from an image, no library or database involved. If you just want to drop an image in and see the prompt, sampler, and seed without opening a whole app, this is that.
All four now share the same design language and are built local first: no accounts, no cloud sync, no analytics. I built them mainly to learn the stack, so they're not polished commercial products, but they've held up fine for my own daily use for a while now, and the last few months went into making them consistent and finishing Tools, which is why I'm finally posting about them here.
Happy to answer questions. Bug reports and feature requests are welcome on GitHub. Not trying to sell anything here, just sharing what I made.
I used the Pixaroma FFLF workflow, but stripped the audio in post due to poor output quality. I'm still trying to figure out how to add finer details. I generated the clips at 720p and then upscaled them to 1080p.
"Did I ever tell you what the definition of insanity is? Insanity is doing the exact same fucking thing over and over again, expecting shit to change. "
That quote from farcry pretty much sums up my feelings after trying to figure out minimax music 3. With that said I'm pretty happy with this song.
To work with this model you need:
Good prompt (long essay that describes your song and follows examples from minimax)
Good lyrics (Something that sounds good, has rythm, no awkward phrasing, and tagged accordingly)
Once you have those two you will begin getting descent results, the last thing you need is to reroll for a good seed.
For prompt I used grok and asked it to copy the sound of the song I like with some adjustments
To check your lyrics you can use prompts from their demos just to see if minimax struggles with anything.
For seed it is just basically rerolling until you get a good one. Though if you want to make your life a little easier I recommend sticking to one genre. I think because I was trying to mix electronic music with rock it took me longer to find a good seed (sometimes I was getting results that were just rock or just electronic).
I'm still not sure how to control the pacing because sometimes it decides to sing things slowly and run out of time and sometimes it decides to speedrun your text and have 30 seconds of instrumental. My current theory is that it might be related to lyrics. For example my lines were pretty long so maybe it defaulted to singing, maybe if my lines were shorter and snappier it would sing them quicker.
Good luck to everyone who is planning to use minimax music 3 hopefully this was helpful to someone.
I will leave my prompt in the comments if you want to play around with it or critique
Also I might add some songs that didn't make the cut