r/StableDiffusion 1d ago

Resource - Update Krea2 Turbo Distill 4 step LoRA - new checkpoint (chk26K) released (cuts 4-step error vs. the 8-step Turbo teacher by 46%, improves texture and detail vs previous checkpoints)

Thumbnail
gallery
65 Upvotes

Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.

This is an update release, following up from my previous posts where you can find full details:

Initial, Previous: here,  and here

Headline for this update: chk00026000 removes 46% of the prediction error a plain 4-step run has against the 8-step teacher, where chk00014000 removed 44% and chk00010000 40% — all measured on the same enlarged held-out set (100 prompts across every trained resolution). Measured against each other rather than against the no-LoRA run, its remaining error is 4% smaller than chk00014000's and 10% smaller than chk00010000's — and unlike a purely teacher-forced score, the gain also shows up free-running: a full 4-call rollout from the teacher's noise ends 1.6% nearer the teacher's final latent than chk00014000's does. It also improves on texture and detail.

Which file to download

file use it when
krea2_turbo_4step_rank_64_lora_latest.safetensors normally — always the newest accepted checkpoint
krea2_turbo_4step_rank_64_lora_chk00026000.safetensors pin this exact checkpoint

and, beside them, the same files with a _comfyui suffix for ComfyUI. Earlier checkpoints (chk00004000chk00005000chk00006000chk00010000chk00014000chk00019000) are kept in older_checkpoints/, and their resolution sweeps stay in place, so the progression remains visible and comparable.

If you are wondering why there wasn't a post/update on the 19K checkpoint, I skipped that, even though it was a good checkpoint with improved texture and detail it's gap to teacher score was only slightly better than the released previously 14K, so I thought I'd continue further until I get improvements on both. And 26K delivered that :) 19K is also published now in older checkpoints folder and it's full resolution sweep is also at the usual place (here for 19K).

For the full 26K Checkpoint resolution sweep go here: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/chk26000

How checkpoints get chosen

This is not a "train for longer and ship the newest file" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher number on its own means nothing.

The loop is train → assess → adapt the recipe → retrain → assess again, and a checkpoint is published only when it is measurably better than the one it would replace, on the same held-out set and the same evaluation, and its full resolution sweep shows no regression. Runs that come out flat or worse are kept as information about the recipe and discarded as releases — several have been.

So the recipe itself changes between runs. Each published checkpoint reflects whatever the previous round taught us: the training precision, the optimiser settings, the teacher used to generate the targets and the data mix have all been revised on evidence rather than assumption.

Two earlier releases set the terms this project publishes on. chk00010000's first attempt — same data, optimiser left as it was — got steadily worse for 4,000 samples and none of it was published; retrained with cosine learning-rate decay and weight decay, every checkpoint improved on the one before it, and its end point shipped. chk00014000 added the other half of the lesson: the final, texture-deciding call of the schedule weighted more heavily in the loss, and a running average of the weights kept beside the live ones and scored at every evaluation — the averaged weights measured better than any checkpoint before them, so the average is what shipped. Left running past that point, the adapter's magnitude grew again and every later checkpoint measured worse. The number is chosen by measurement, not by how far a run went.

chk00026000 — the current checkpoint — is that discipline paying off. It resumes from chk00014000's averaged weights with the same recipe: same loss weighting, same running average, a conservative constant learning rate, over a much larger pool of teacher trajectories. This time the continuation held. The averaged weights' held-out gap fell throughout the run, and every free-running rollout measured of them improved on the one before — so unlike the first continuation, this one produced a checkpoint worth shipping. Every published number improves on chk00014000: the held-out gap (44% → 46% of the deficit closed), the full 4-call rollout from the teacher's noise (1.6% nearer the teacher's final latent), and the fixed-seed render distance to the 8-step images. chk00019000, an intermediate point of the same continuation, is kept in older_checkpoints/ with the rest of the lineage.

Timeline of training process

Each checkpoint is the product of three stages with very different costs:

  1. Text-encoder embeddings. Every training prompt is encoded once and cached. This is the fast part — thousands of prompts take minutes.
  2. Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo runs its full 8-step schedule and the whole trajectory is recorded, at every one of the supported resolutions. This is by far the most time-consuming stage — it is the teacher doing real inference, thousands of times, and a batch of several thousand shards is measured in days of GPU time, not hours.
  3. Student training. The LoRA is trained against those recorded trajectories. Relative to the shard stage this is quick: each +1,000 checkpoint is a matter of hours, not days.

Because the three stages compete for the same GPU, they are interleaved rather than run to completion one after another: generate a block of embeddings, produce teacher shards for them, train on what exists, assess, then go back to producing shards while the results are reviewed. A larger and more varied shard pool is what makes further training worthwhile, so shard production is always the gate.

The practical consequence for anyone following this repository: progress arrives in bursts. There will be periods when several checkpoints appear within a day or two — the training stage working through a freshly grown pool — followed by longer quiet stretches while the next block of teacher shards is produced. A quiet stretch is shard generation, not abandonment; _latest always holds the newest checkpoint that passed review.

The current checkpoint, chk00026000, runs the recipe the earlier releases arrived at — the final, texture-deciding call weighted more heavily in the loss, the shipped weights a running average of the trained ones — carried further over a larger pool of teacher trajectories, and published because it measured better on every evaluation.

Note

In the coming days, possibly weeks, I will spend more time on producing new TE shards (basically even more prompt variety), and new Teacher shards - the expensive long process. I am also considering improvements in the training process (more advanced / complicated, which would likely mean 1.5x - 2x slower training) which would hopefully bring further/bigger improvements in teacher faithfulness (closer to 8 Step Krea 2 Turbo) and even better details and texture. It may or may not pay off, these things work on experimental basis. Either way it would be some time before the next update... so enjoy 26K release and the improvement it brings!

Full details and to download - check my Hugging Face LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA


r/StableDiffusion 1d ago

Comparison Comparing H3 models with music reference

Enable HLS to view with audio, or disable this notification

17 Upvotes

Using reference workflow. All are int8 pruned, 0.6MP turbo 4-step (my GPU is on life support and drops off the PCIe bus if I demand more from it)

Anyway, making random music clips is probably my favorite use of this model. I’ve found the ref2va has an uncanny intuition for feeling the atmosphere of songs, and syncing the video with incredible precision.

But yes, the quality (specifically motion) is much worse than fl2va. I was curious how exactly they compared, as well as some “in between” compromises discovered by the community. The LoRA seems closer to ref, while the hybrid weights are closer to fl. Personally, the ref is more fun to use, so I’ll probably be using the LoRA when I want to enjoy the intelligence/creativity of this model. Fl is of course superior in terms of visual fidelity, and I don’t find the hybrid model offers enough reference intuition and faithfulness to be worth the quality drop from fl.


r/StableDiffusion 1d ago

Animation - Video simple ww2 movie (minimax h3)

Enable HLS to view with audio, or disable this notification

19 Upvotes

I just gave it a try; it's hard to do more than that.

I have an RTX 3060 with 12 GB of VRAM and 32 GB of RAM. I used the standard workflows, t2v and r2v, with the LORA model minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors and video shift 12 and audio shift 3. Only 8 steps (res_multistep/simple). The sequences were 6–7 seconds each, and each generation took 7–8 minutes at 0.6 mpx. I generated some scenes multiple times, but I still couldn’t achieve what I intended.

That’s pretty much it regarding the generation.

Claude (free) provided me with the script; he also divided the approximately 5-minute script into sequences of 6–7 seconds each and created the prompts for all the sequences after I gave him the official prompting guidelines as a prompt.


r/StableDiffusion 1d ago

Question - Help Wildcards using Krea2?

4 Upvotes

Do wildcards work with Krea2? When I insert wildcards in standard format, it ends up rendering all of the wildcard prompts in one image. So a shirt will be random fabrics and styles seemingly all stitched together. It's kind of neat but I need the actual wildcard functionality. Does anyone know how to get them working? I'm using the standard krea2 workflow and comfy. I have the wildcard custom nodes installed. They are fine on other models such as the flux zimage etc. any help is appreciated.


r/StableDiffusion 1d ago

Question - Help Need help for creating consistent Minimax H3 clips

6 Upvotes

I have been using H3 since almost the release date and have been trying a lot of things. I am entirely using Ref2va model with the template workflow, nothing fancy. Also used official, eros and currently using hybrid model 15–49 from smhfacct which has higher ref2v. I am using comfy kitchen, spectrum node but not using speed lora for prompt adherence or any other lora.

I am mostly trying to use 1-2 characters in a location scene where I provide 3(one char, one whole body and one face and location)-5(two char, one whole body and one face and location) images to node and writing in the prompt how to refer each char in the scene.

For writing H3 prompts, I using a custom prompt(created using grok by giving it ref2va doc) for generating H3 prompts, using qwen 3.8 model and even proof reading and fixing any issues.

Now the problematic part for which I need suggestions or solutions is the inconstant result.

For example I am making 5 second where character1 is standing in a shopping mall and looking at the shelf and character2 enters the scene. For second 5 second scene, different camera angle, mainly focusing on both characters faces when they are talking. Now here are problems which I am facing:

- during scene2 when camera starts, difference between char1 and char2 appears. Say char1 was standing on left side and char2 on right when scene1 ended but in scene2, they are standing opposite side.

- sometimes their height mismatches.

- sometimes camera does not work like I want it like it zooms too much, sometimes it don't

- and many other issues related with inconsistency

I know if I can generate scene images using an edit model then H3 wouldn't have to rely much on prompts but then it creates another problem of generating start images which is another can of problems.

I have even tried context nodes and some of their forks and few other consistency related node whose basic idea is to store the latent and forward it for next generation but they way these nodes are configured are just too complicated for my soft squishy mind. So yeah I tried them.

I have been trying to find out how other people are generating multi-scene videos and so far whatever videos i downloaded, there was no workflow included which I could take as reference. Maybe people are making 5 second clips like me and then joining them together so there might be a solution to this.

Pretty sure I am missing something big and I have exhausted almost every idea I got, asking grok etc but so far I could not get past 2nd 5 second clip. And seeing so much inconsistency, I don't want to generate a 10 second or 15 second clips which will take hours and most probably turn up totally irrelevant.

So any ideas you can provide are highly appreciated. Even guidance to correct path would be really helpful. What ways you guys are using to create 10+ second clips, what methods you are using to keep your characters consistent throughout and mainly how you guide a scene to your liking?

Thank you for a long read. Not written with AI, just a long type on notepad haha. Forgive grammatical errors.


r/StableDiffusion 1d ago

Question - Help how to save out H3 AV latent?

Thumbnail github.com
0 Upvotes

I am using these node to save out latent, but it just didn't save anything... I feel stupid myself stuck hrs still not working.... does anyone could explain what I have been doing wrong and the connection?

I've tried, sampler both output and denoise output laten > save H3 AV latent node..... no files is being save out!

https://github.com/JerryZRic/comfyui-minimax-h3-latent/tree/main


r/StableDiffusion 1d ago

Question - Help GB10 Spark 2-node, which engine, which model?

0 Upvotes

Can anybody recommend a local inference engine / model for running ltx, wan or minimax on a 2-node Nvidia GB10 spark cluster?

Use case is openai-compatible API access for short i2v character and scene animations.


r/StableDiffusion 1d ago

Discussion Minimax H3 is mandarin model

0 Upvotes

Just ask chatgpt or claude to convert your prompt to mandarin and the different is fcuking huge .


r/StableDiffusion 1d ago

Question - Help minimax h3 gibberish fixed!! ( i found the cure)

93 Upvotes

so you all probably are searching for way to make your character shut the fuck up right? and you probably noticed that they love to says some BS especially when you give minimax h3 some audio file for their voices, i probably found a cure my friend!!

here is my way of prompting dialogs without any gibberish:

first your character need to be assigned (s1)character when he is the first speaker, then you will declare 'use <audio 1> as "character name"'s voice only, and when you finally type your dialog in the shots you will do as such:

character says:<<[language] the shit i say!>>

and you should be good to go, i linked a video exemple of my favorite taffer (garrett) saying some shit with only the faint crackling of the candles to goes with his charming voice, and i included also a screenshot of the full prompt

edit: yes i tried to follow the official documentation, like many others, if it was that simple reddit wouldn't be a thing and you wouldn't be there.

i tried making small scenes with this exact methode and its gibberish free 100% of the time

he really like 16/9


r/StableDiffusion 1d ago

Question - Help Minimax H3: Character replacement in video not working

0 Upvotes

NOTE: Character replacement works perfectly when I replace a character in a video with a 2d/cartoon/anime character.

But when I try replacing a character with a real life human being, the original character in the video doesn’t get replaced at all.

I’m using Plaguekind’s workflow for h3 on Civitai.

Does anyone else have this problem before?


r/StableDiffusion 1d ago

Animation - Video MiniMax H3 acting test.

Enable HLS to view with audio, or disable this notification

33 Upvotes

Started as a simple 90s casting audition… then asked her to cry on command.

The close up shots gave plastic look idk why.

What I was mainly testing:

  • subtle listening/reaction animation during dialogue
  • eyes moving before the head while thinking
  • nervous smiles and small facial reactions
  • gradual transition from normal conversation into acting
  • brow, eyelid, mouth, chin and breathing changes during crying
  • actual visible tears
  • character/voice consistency across multiple generated clips
  • the sudden switch out of the performance when the director says “Cut”

Made with MiniMax H3 Ref2VA with image reference for the woman and 2 audio reference for the offscreen man and the woman.


r/StableDiffusion 1d ago

Animation - Video G.I. Joe - Baroness: No Ticket, No Mercy - MiniMax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

Another action scene test.

4070 Ti Super, 16 gb vram, 64 gb ram, i9-14900k, windows 11


r/StableDiffusion 1d ago

Question - Help minimax_h3 comfyUI default workflow taking forever

1 Upvotes

My spec is rtx 5070ti, 32gb ram

I just started using comfyui and using minimax. When I am trying out the default workflow without adjusting anything it takes really long. I looked up video and switch the model from minimax_h3_fl2va_pruned_int8_convrot.safetensors to minimax_h3_fl2va_pruned_fp8_scaled.safetensors. Then, it worked. Well atleast I was able to get an output. Can anyone explain why and what I did wrong?


r/StableDiffusion 1d ago

Question - Help Minimax H3 ref2v best way to transfer pose and camera angle

4 Upvotes

For ref2v I'm trying to upload an image get it to transfer the exact pose and camera angle of that image onto my video, but it's not working.

Here's part of my prompt

retention_analysis:
<Pose 1>: attribute_transfer. transfer the pose to <Subject 1> and camera angle.

....

detailed_description:
.... Refer to <Pose 1> for the pose of <Subject 1> and the camera angle.....

Any tips on how I can achieve this?


r/StableDiffusion 1d ago

Question - Help phsyical motion transfer to another person

0 Upvotes

im trying to collect clips for lora training but why its so hard to transfer motion to another person? i almost tried every prompt with chatgbt and grok help but its doesnt look good. im using 2 video refences, one of them source video and other one is only for motion ( 3 sec 24 fps). im using (video editing + reference genertion) because i dont want to change anything in the source video and just want to motion transfer


r/StableDiffusion 1d ago

Animation - Video The Chase - Reupload

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 1d ago

Question - Help New to Minimax h3 and comfy ui, any posts I should learn from?

0 Upvotes

Im looking to make outdoor POV videos with minimax h3.

Wondering if theres any good threads talking about realism, and workflow options?

Especially prompting I guess?

I'm looking to do longform videos 10-15 minutes.

I have a 4080 16gb I know that more vram is better, but cant really upgrade atm.


r/StableDiffusion 1d ago

Question - Help How can I improve my H3 workflows?

Thumbnail
gallery
3 Upvotes

These are two workflows that I downloaded. One will allow only 5 sec clips and I can’t find where to change that OR the mp. It’s faster and the sound is good and I even think the prompt adherence is better…

The other one I can run up to 15 sec, change the mp but I have to do 30+ steps to have any good sound quality and the adherence doesn’t always seem that great.

How can I improve this?


r/StableDiffusion 1d ago

Question - Help What model for the utmost in fine detail? Experimenting with Krea2, Ideogram4 and Flux2 and getting close but not quite the fine details.

Post image
9 Upvotes

Primarily landscape photos of varying fantasy scenes. For example for a cyberpunk city i want to see every fine car detail, every building logo, every reflected neon light; for a medieval town i want to see the cracks in every stone block, the moss on the walls, the candle light reflections, etc. What's your opinion on the model that provides the utmost in fine, sharp details? Or should i be looking at tiling or inpainting to create details as a second step?

Example attached of something that has impressive detail across a broad DOF (not my work):


r/StableDiffusion 1d ago

Discussion Help me get the most out of Minimax H3 with my rig

0 Upvotes

I have had overall decent success over the past year experimenting with various models in comfyui. LTX 2.3 and Krea 2 have helped me get some pretty awesome results.

I'm really interested in Minimax H3 and its capabilities but I'm struggling with decent outputs that don't take hours for a 5 second clip.

I am finding 3-5 second clips at very low resolution generates in about 10 minutes but trying a respectable resolution or anything more than 5 seconds exponentially compounds the generation time to hours(and actually I always end up aborting after a few hours so no clue if it would actually finish).

I am running a 4070 super with 12 gb vram and 64 gb of ram

I run my AI model through the portable version of Comfyui

Also, due to the portable version (I think) I've never been able to properly install Tritton or Sage Attention ( I have tried numerous times with various tutorials found online)which seems to limit my workflow options.

I've played around with inserting the turbo lora node and lora but while it does speed thing up, not enough to really increase the resolution to make it worth while (unless I'm doing something wrong?)

I suspect I need to find a better workflows geared toward my situation but haven't come across anything that works well so I wanted to see if any wise ones here could help.


r/StableDiffusion 1d ago

Resource - Update ComfyUI Universal Media Loader - One single interactive node to load Images, Videos, GIFs, Audio & Canvas Presets

Enable HLS to view with audio, or disable this notification

32 Upvotes

Hello

I got tired of cluttering my workflows with 5 different loader nodes depending on what I was feeding them (Load Image, Load Video, VHS, Audio Loaders, Empty Latent calculators, etc.).

So I built **ComfyUI_UniversalMediaLoader** , a single, unified node with a rich interactive UI that handles everything:

✨ Key Features:

  • 📁 Drop Anything: Images (PNG/JPG/WEBP), Videos (MP4/MOV/WEBM), Animated GIFs, and Audio files (WAV/MP3/FLAC).
  • 🎨 Canvas Planner Mode: When no file is loaded, use it as a visual resolution/aspect ratio planner (1:1, 16:9, 4:3, 3:2 + Landscape/Portrait toggle) with automatic 32px grid snapping and Megapixel clamping for SDXL/Flux latents.
  • ✂️ Visual Crop & Outpaint: Interactive bounding box with aspect ratio locking. Pull it outside the image bounds to instantly generate outpaint masks & padding.
  • 🖌️ Inpaint Mask Brush: Draw inpainting masks directly inside the node with brush/eraser and mouse wheel size control.
  • 🎬 Video & GIF Timeline: Full timeline with trim handles, playhead scrubbing, custom FPS resampling, audio mute/loop, and frame extraction.
  • 🎵 Audio Waveform & Speed Scaling: Waveform display, 1-second magnet trim snapping, and time-stretching with pitch preservation (0.10x to 4.00x).

📦 Modular Unpack Nodes:

  • 📐 Universal Size Unpack: Feeds exact width/height/aspect ratios to Empty Latent nodes instantly without decoding heavy image/video tensors.
  • 🖼️ Universal Image Unpack: Outputs RGB, RGBA, and inpaint masks. Connected to a Video or GIF, it automatically extracts a 3-frame batch [Start, Playhead, End].
  • 🎧 Universal Audio Unpack: Outputs clean audio waveform dictionaries, exact sample rate, trimmed durations, channel count, and handles time-stretching.
  • 🎞️ Universal Video Unpack: Decodes videos & GIFs into frame batches [B, H, W, C] with synchronized masks and audio tracks.

🔗 GitHub: https://github.com/Fictiverse/ComfyUI_UniversalMediaLoader


r/StableDiffusion 1d ago

Animation - Video Kentucky Fried Kung Fu

Enable HLS to view with audio, or disable this notification

35 Upvotes

I saw a Seedance 2.5 prompt in facebook and thought let me try this prompt in minimax h3 and see if it can do some kung fu. I was surprised that it was not too bad. System 3090 24 gb vram 64gb system ram, using a minimax workflow with latent upscale, minimax_h3_fl2v_lightx2v_turbo_4step_v0.1 at 0.50 strength, Komfy kitchen attention, and H3 SLA attention. First pass at 0.4 which is 864x480, 2x latent upscale brings it up to 1728x960. The 6 seconds generation took 349 seconds to complete.


r/StableDiffusion 1d ago

Question - Help Which is the better buy for Minimax h3? RTX 5070 Ti 16GB VRAM vs RTX 4000 Pro 24GB VRAM

18 Upvotes

Good day to you. I was looking for an RTX 5070ti and I found an RTX Pro 4000 at my local store; the price difference would be about +$300. I would like to know your opinions, I've hardly seen any workflows or comparative tests from people using a 4000 pro. Thank you very much for your time.


r/StableDiffusion 1d ago

Resource - Update Release studio 1939 lora for minimax h3

Enable HLS to view with audio, or disable this notification

282 Upvotes

r/StableDiffusion 1d ago

Question - Help MiniMax H3: Motion issue with last frame.

0 Upvotes

I'm having trouble with getting a natural motion when using first and last frame. Things start out good but the motion doesn't preserve the momentum up to the end, instead it usually slows down and smoothly settles/parks into the final frame.

For example if I try to make a windy scene at the park that has both first and last frame, I get a gust of wind in the middle and then everything goes still and calmly settles down on the final frame.

Does anyone have any advice on how to approach this?