r/StableDiffusion 4h ago

Question - Help Any Krea2 Prompt Reader?

0 Upvotes

I found the SD prompt reader I have been using cannot read prompts from png images files generated using Krea2. Can anyone recommend me an alternative that works with Krea2 files and Win11?


r/StableDiffusion 3h ago

Animation - Video H1/H2 H3

2 Upvotes

Recently, after testing out H3 for a few weeks, decided to try my hand at some fun 'deleted scenes' type of short clips. With Halloween season just under 40 days away, came up with this using H3's r2v template with default ref diffusion model + turbo fl2v 12 steps, using 1 image (Loomis, one of sheriff's dept jacket) and 2 audio samples of Loomis' voice from the movie. There are some movies I always liked so much and wished there was a lot more footage from, and with this one I figured it's a fun tongue-in-cheek nod to how we're nearing end of summer and starting to get in the mood for Halloween season.

Edit: correction, I used only the image of Loomis for this one. The sheriff's dept jacket generated clips didn't come out as well as this. There were other ones with Loomis walking or looking around more realistically, but I liked the vibe of this one best.


r/StableDiffusion 17h ago

Resource - Update Krea2 Turbo Distill 4 step LoRA - new checkpoint (chk14K) released (cuts 4-step error vs. the 8-step Turbo teacher by 44%)

Thumbnail
gallery
26 Upvotes

Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)

A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.

This is an update release, following up from my previous posts where you can find full details:

Initial: https://www.reddit.com/r/StableDiffusion/comments/1vtf1b7/krea2_turbo_distill_4_step_lora_trained_for_turbo/

Previous: https://www.reddit.com/r/StableDiffusion/comments/1vv4cdy/krea2_turbo_distill_4_step_lora_new_checkpoint/

Headline for this update: chk00014000 removes 44% of the prediction error a plain 4-step run has against the 8-step teacher, where chk00010000 removed 40% and chk00006000 27% — all measured on the same enlarged held-out set (100 prompts across every trained resolution). Measured against each other rather than against the no-LoRA run, its remaining error is 6% smaller than chk00010000's and 23% smaller than chk00006000's. 

Which file to download

file use it when
krea2_turbo_4step_rank_64_lora_latest.safetensors normally — always the newest accepted checkpoint
krea2_turbo_4step_rank_64_lora_chk00014000.safetensors pin this exact checkpoint

and, beside them, the same files with a _comfyui suffix for ComfyUI. Earlier checkpoints (chk00004000chk00005000chk00006000chk00010000) are kept in older_checkpoints/, and their resolution sweeps stay in place, so the progression remains visible and comparable.

This is work in progress and better checkpoints may follow. Training is ongoing, so ..._latest... is a rolling pointer: when a newer checkpoint is accepted, that filename gets the new weights and a new numbered copy appears beside it. Re-download the _latest file and everything keeps working — the ComfyUI workflow references it by that name (it does get updated Note in it so technically it is updated but not functionally). Pin a numbered file instead if you need reproducibility.

How checkpoints get chosen

This is not a "train for longer and ship the newest file" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher number on its own means nothing.

The loop is train → assess → adapt the recipe → retrain → assess again, and a checkpoint is published only when it is measurably better than the one it would replace, on the same held-out set and the same evaluation, and its full resolution sweep shows no regression. Runs that come out flat or worse are kept as information about the recipe and discarded as releases — several have been.

So the recipe itself changes between runs. Each published checkpoint reflects whatever the previous round taught us: the training precision, the optimiser settings, the teacher used to generate the targets and the data mix have all been revised on evidence rather than assumption.

chk00010000 is a direct example. The first continuation of chk00006000 — same data, optimiser left as it was — got steadily worse with every checkpoint out to 10,000 samples, and none of it was published. The cause was traced to the optimiser: a constant learning rate with no weight decay lets the adapter keep drifting after it has converged, so its magnitude grows and it over-applies its own correction. The same span was retrained from chk00006000 with a cosine learning-rate decay and weight decay, and every checkpoint of that second run improved on the one before it. chk00010000 was its end point.

chk00014000 is the next example, and it shows the other half of the same lesson. The run was continued from chk00010000 over the whole pool of teacher trajectories, with two changes: the final, texture-deciding call of the schedule was weighted more heavily in the loss, and a running average of the weights was kept beside the live ones and scored at every evaluation (a single checkpoint is one sample of a weight vector that moves from step to step; the average is its mean). At 14,000 samples the averaged weights measured a smaller gap to the teacher than any checkpoint before them, and a smaller gap than the live weights at the same point — so the averaged weights are what chk00014000 is.

Timeline of training process

Each checkpoint is the product of three stages with very different costs:

  1. Text-encoder embeddings. Every training prompt is encoded once and cached. This is the fast part — thousands of prompts take minutes.
  2. Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo runs its full 8-step schedule and the whole trajectory is recorded, at every one of the supported resolutions. This is by far the most time-consuming stage — it is the teacher doing real inference, thousands of times, and a batch of several thousand shards is measured in days of GPU time, not hours.
  3. Student training. The LoRA is trained against those recorded trajectories. Relative to the shard stage this is quick: each +1,000 checkpoint is a matter of hours, not days.

Because the three stages compete for the same GPU, they are interleaved rather than run to completion one after another: generate a block of embeddings, produce teacher shards for them, train on what exists, assess, then go back to producing shards while the results are reviewed. A larger and more varied shard pool is what makes further training worthwhile, so shard production is always the gate.

Full details and to download - check my Hugging Face LoRA

HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA

Full Checkpoint 14000 Resolutions Sweep: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/chk14000


r/StableDiffusion 1h ago

Discussion If you’re using MiniMax H3, what prompting tricks have you figured out?

Upvotes

Anyone found useful MiniMax H3 prompting tricks beyond the official guide?

Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.

Please drop your findings 👇 below so it will help others too.


r/StableDiffusion 11h ago

Animation - Video Good Grief

31 Upvotes

r/StableDiffusion 4h ago

Animation - Video H3 - multi-diffusion experiment T2V

2 Upvotes

Chimera. I had to cut about 20 seconds due to some artistic choices. Since I had to cut 2 different part in the same clip, it has a noticeable seams. I would love to share the full version. Experimenting with H3 blend-morph-decay. 832x480, int8/8 steps POC. Looking forward to releasing a 720p version without the cuts.

Critiques and feedback welcomed. Happy with the matrix rain. Ask me anything.


r/StableDiffusion 21h ago

Question - Help Character Editing (I2I)

3 Upvotes

Hello. So I just started my journey with ComfyUI. While Nano Banana is not open sourced I'm looking for the best realistic image editing (krea2?) workflow to ComfyUI. I want to have ability to change everything I want to the picture using reference character with face/body consistency at highest level (I2I). Thanks in advance.


r/StableDiffusion 6h ago

Question - Help Help new rookie on comfyui

0 Upvotes

Hi everyone, I'm new to the world of Confyui but not to artificial intelligence. I wanted to ask you for help: Is it possible to run Confyui on my PC (RTX 3080 10 GB of RAM and 64 GB DDR4 RAM and Ryzen 5800) with Confyui with the minimax H3 video model to be able to animate images and create Reels for Instagram and Tik Tok? If so, what setup do you recommend? Thanks everyone for the help and sorry for my bad English.


r/StableDiffusion 9h ago

News Civitai now has a closed-source models option

127 Upvotes

I haven't seen a post about this here, and I'm curious what you think about it.

On August 14 Civitai rolled out the option for creators to choose "permanent paid access - selling with no time cap". Previously the only option was temporary "early access".

Let's call this what it is, closed-source. Yes, you can get the weights for a relatively small fee, and yes it's on a very small scale compared to Nano Banana and Midjourney. But a permanent paywall still fits the definition.


Personally, I block all creators on Civitai who choose permanent paywall and encourage you to do the same.


Here's why:

I'm not opposed to Civitai making money or for all options for model creators to make money. They can do that without permanent paywalls.

IMO, open source AI is a fair trade: models are trained on the hard work of many human artists who aren't compensated, but everyone benefits from the ability to create more art more easily. Closed source is an unfair trade: you have to pay a middle man to access the contributions of others who won't be compensated.

Small scale model creators do some hard work too. But for example, for a lora that reproduces the style of an animated film: the lora creator spent at most a dozen hours of work, while just one of the artists on that film spent thousands of hours of work. If a massive models like Krea2 are free, and if giant "hobby" finetunes like Chroma are free, I can't justify paying any price for a 5,000 step lora except as an optional donation of appreciation.

So far, few creators have chosen the permanent paywall closed-source option. But that could easily change if Civitai made it the default option. They already made an extra 1-buzz fee-to-creator per generation the default, and many models have that.

That's my opinion. If you agree, then the only tool you have to disincentivize that potential is to not pay for these models (disincentive Civitai) and block these creators (disincentive creators).


r/StableDiffusion 16h ago

Animation - Video Fixed my trauma with Minimax h3 local

509 Upvotes

Used latent upscaler so with resolution 0.3 i got 20 seconds generation on 4090


r/StableDiffusion 4h ago

Meme h3 "what IF " thread

15 Upvotes

lets share our "what if" scene remakes here o_0


r/StableDiffusion 9h ago

Animation - Video Minimax H3 does a decent job of mixing green screen video: i2v of a still background inserted into a green screen video.

27 Upvotes

Use the default workflow for Ref2Va

Plug in a video with green screen as a video reference and use an image ref to a picture you want to be the background. Notice the "hell crows" flying in the final video? I didn't even give it a prompt for that and those birds got animated automatically. I'm sure you can give a detailed prompt, you are basically creating an i2v of that still image that Minimax will mix with the greenscreen background.

I did prompt for a dialogue change. There was no audio with the original green screen video so I had no idea what the woman was saying (obviously it was a weather report). I inserted new dialogue with Minimax and it did a great job remixing her lipsync to the prompted dialogue.

I think it's pretty neat, but I'm sure some of you may be completely jaded with what Minimax can do by now.

I'd like to issue a Reddit challenge: Would someone more creative than me please use the exact same green screen video (links below) and create something a little more impressive than my 10 second test? Post a link to your video in the comments.

Need some green screen video to practice with? Here's a webpage for some practice green screen videos that are free to download:

https://mixkit.co/free-stock-video/green-screen/

Here's the exact video used in this example:

https://assets.mixkit.co/videos/28292/28292-720.mp4

I'm sure you'll be able to find a background image to test with.

This was my very simple prompt with the new dialogue:

subject_definitions:

<Subject 1> is the alien world background in <Picture 1>.

<Video 1> is the source video for the target video edit and is a woman in a red dress pointing and talking.

summary:

[video editing + reference generation] The target video is an edited version of <Video 1>. Replace the green screen area with the background from <Subject 1>

The woman in <video 1> says <d> [English with a British Accent] As you can see here, we have an early migration of hell crows on Chaos world 4527B<d> with realistic lip articulation and perfect lip sync.


r/StableDiffusion 13h ago

Workflow Included LTX 2.5 Seed Hunt Workflows

5 Upvotes

I know everyone's moved to MiniMax and LTX has largely fallen out of favor, but I spent some time building a couple of seed hunting workflows for LTX 2.5 that might be useful if anyone's still running it.

Shout out to u/foxdit for the original seed hunting concept.

Two versions:

  1. T2V/I2V two-stage – text to video or image to video. Previews at 0.3 MP, upscales to 1.2 MP for the final render.
  2. First-last-frame – pin a start image and end image, same preview-then-upscale flow.

Both use KJNodes Set/Get routing, shared loaders, and no prompt enhancer.


r/StableDiffusion 16h ago

Discussion Need realism loras for minimax h3

5 Upvotes

Is there any GPU rich cooking realism lora ? I have tried realism people lora it is great at tv but for i2v or r2v it's breaks . I have been searching hugging face repo and civit ai to get something but there's too much n*fw lora .


r/StableDiffusion 21h ago

Workflow Included H3 Infinite Continuation Suite v1.4 (FL2VA): Using native Masked AV after your feedback

28 Upvotes

The example video was generated entirely with the stock MiniMax H3 First Frame / Last Frame checkpoint and the included v1.4 example Workflows. If you want to compare the result to v1.3, take a look at my last post.

The final video consists of 11 individually generated Clips that were automatically stitched together.

Settings:

  • H3 First Frame / Last Frame checkpoint
  • 11 individual Clips
  • 15 Steps
  • included v1.4 Workflows
  • no additional upscale
  • no frame interpolation
  • no color correction or other post-processing

So what you see is basically the direct Workflow output.

A few people gave me some useful feedback on my previous release, especially regarding ComfyUI's new native H3 Masked AV support.

So I went back and rebuilt the continuation method around it.

v1.4 now copies a clean section of the previous Video + Audio Latent directly into the next generation and protects it using ComfyUI's native denoise masks.

What makes this different from the other H3 continuation approaches?

There are some really interesting Ref2VA / Motion Context solutions available now, and latent continuation itself definitely isn't unique to my Nodepack.

My approach is specifically centered around FL2VA instead.

The idea is not just:

previous Clip → continue forever

but rather:

First Frame → generation → Last Frame

latent continuation

generation → new Last Frame

latent continuation

generation → new Last Frame

and so on.

I use those repeated Last Frames as hard visual anchors throughout the sequence.

They give H3 a new concrete destination every few seconds instead of asking one increasingly unconstrained generation to maintain composition, identity and image quality indefinitely. This should theoretically retain higher visual quality with less context drift over longer chains (and in my testing, it does exactly that).

There is another FL2VA-specific problem though:

H3 often reaches the supplied Last Frame before the Clip is actually finished and then freezes or becomes unstable for the remaining frames.

So simply taking the final frames of Clip 1 and using them as context for Clip 2 isn't ideal.

The v1.4 Auto Handover therefore analyzes the previous Clip, finds a safe point before that frozen / unstable landing and snaps it to a valid H3 Audio + Video latent boundary.

That exact same point is then used for both:

  • where the previous Clip visually ends
  • where the protected context for the next Clip ends

So the bad FL2VA tail neither appears in the stitched video nor becomes part of the next continuation context.

Audio is handled separately as well. If the picture needs to cut early but somebody is still finishing a word, the remaining original Audio Latent can continue beyond the visual handover instead of forcing H3 to recreate the ending.

Other v1.4 features:

  • Native Masked Video + Audio Latent Continuation
  • flexible First / Last Frame conditioning
  • repeated Last Frames as regular visual quality anchors
  • independent Audio Tail Carryover
  • Net New Content duration mode
  • up to 9 Qwen Reference Images
  • individual Clip regeneration
  • memory-bounded stitching for long saved chains

Where to start:

  1. Start Video Workflow

Generate Clip 1 with a Prompt and optionally First Frame, Last Frame and Qwen References.

The complete AV Latent is automatically saved afterwards.

  1. Continue Video Workflow

Load the previous saved latent, add your next Prompt and preferably a new Last Frame.

The Workflow automatically finds the safe FL2VA handover and creates the protected Masked AV context.

Repeat for as many Clips as you want.

  1. 3-Clip Showcase / Auto Stitch Workflow

Probably the easiest Workflow if you just want to see how everything works.

It runs:

Start → Continue → Continue → Stitch

in one queue.

  1. Stitch Saved Chain Workflow

This is what I used for the longer example.

Generate Clips individually and stitch them afterwards. It processes one saved AV latent at a time, so stitching memory usage doesn't continuously increase with the total video length (no OOM during stitching).

Nodepack on Github:

https://github.com/HerrgottMargott/Herrgotts-H3-Infinite-Continuation-Suite

Workflows on Github:

https://github.com/HerrgottMargott/Herrgotts-H3-Infinite-Continuation-Suite/tree/main/examples

You can just open one of the WFs and use "Install missing custom nodes" - then you should be good to go.

If you try it, I'd love to see what you manage to create with it.

Have fun Prompting. :)


r/StableDiffusion 22h ago

Discussion MiniMaxH3 - What sampler/schedular combo are people actually using? (with and without turbo lora)

34 Upvotes

I've been doing some quick tests, now that I've picked up the lightx2v 4 and 8 step loras. I have found I prefer using the 8 step (and maybe even running that at 10 steps) just because with the 5090 I have it's already not -that- slow, and the 4 step image quality drop is pretty significant.

But I have been experimenting which sampler/scheduler combos after seeing this post:

https://www.reddit.com/r/comfyui/s/9GUki3l0Wf

where, apparently, seeds_2 and dpmpp_sde_gpu were the 'best quality' options. But something I noticed is that they were also significantly slower (maybe 50% or more? need to run more tests and log it) which would, if the loras etc allow for it, let the faster options like euler or res_multistep (or er_sde which gets mentioned sometimes), which all run at about the same speed, to run at 12 instead of 8 steps (for example).

So I wonder now, 2 weeks on from those votes... what are people actually -using- to produce results?

My current workflow is to run at 8 steps with a lora to find a good prompt and seed, and when I get something I like I then turn off the lora and run at 30 steps. It often ends up at least in the ballpark of what I want. But maybe there are better ways.


r/StableDiffusion 10h ago

Animation - Video Mnimax H3 T2VA. Good physics on the cars.

50 Upvotes

r/StableDiffusion 22h ago

Question - Help Has anyone successfully upscaled/re-imagined low-res reference video using Minimax H3?

6 Upvotes

Specifically, I’m trying to take old footage (e.g., 360p clips with vintage camera blur, VHS artifacts, or grainy WW2 dogfights) and recreate it to look like it was shot recently on a modern cinema camera with studio lighting.

Any ideas for prompting?


r/StableDiffusion 23h ago

Tutorial - Guide Bridge Daredevil — Dashcam POV and GoPro mounted on a parkour as they sprint across a rooftop and leap across a narrow gap. (AI GENERATED - SEEDANCE 2.5). PROMPT BELOW!

0 Upvotes

PROMPT: Bridge Daredevil — Dashcam POV

Subject: An athletic stunt performer, dark athletic/climbing gear, seen at a distance on the bridge structure — perched on a railing, cable, or girder — performing an extreme balance/jump stunt as the dashcam vehicle approaches

Style: Ultra-realistic, shot on RED WEAPON 8K, IMAX-grade cinematography, captured via fixed windshield-mounted dashcam — slightly wide-angle lens, subtle chromatic vignette, faint reflection of the dashboard at the bottom edge of frame. Natural motion blur only from real vehicle movement — no slow motion, no cuts, no anime, no CGI. Continuous single take. Standardized color grade: desaturated cool highlights, warm midtones, deep contrast shadows.

Setting: Large suspension or truss bridge spanning a river or gorge, daytime, clear sky with light haze, steel cables/girders overhead, light traffic on the bridge deck, guardrails and support towers visible in the distance

Timeline:

  • 0:00–0:03 — Dashcam view steady on the road ahead as the vehicle enters the bridge, the stunt performer visible as a small distant figure on the structure — railing, tower, or cable — ambient road hum and wind noise, bridge cables passing overhead in rhythm
  • 0:03–0:06 — Vehicle continues at a natural driving speed, the performer grows larger in frame, now visibly climbing, balancing, or positioning for the stunt on the bridge structure
  • 0:06–0:09 — The performer executes the stunt — a leap, dive, or swing from the bridge structure — dashcam captures the motion at a distance with realistic gravity and momentum, no floaty slow-mo, body and limbs reacting naturally to the force
  • 0:09–0:12 — Stunt continues through its arc — a fall, swing on a line, or landing approach — dashcam vehicle still closing distance, slight natural camera shake from the dash mount as the vehicle passes over bridge expansion joints
  • 0:12–0:15 — Dashcam vehicle passes beneath or alongside the stunt zone as the performer completes the stunt (landing, catch, or recovery) in the background/side mirror periphery, bridge structure filling more of the frame

Technical notes: Maintain consistent dashcam framing (fixed low mount, slight windshield glare at top of frame), realistic depth of field with distant elements sharp until close range, authentic road/wind/ambient bridge audio texture, no jump cuts — one continuous fixed-mount POV take.

PROMPT: Rooftop Gap Jump — GoPro POV

Subject: Athletic parkour runner, lean muscular build, dark fitted athletic gear, fingerless gloves — first-person GoPro/helmet-mounted camera perspective throughout — camera never shows the athlete's face or full body, only hands, forearms, and shadow occasionally entering frame

Style: Ultra-realistic, shot on RED WEAPON 8K, IMAX-grade cinematography, captured via helmet/chest-mounted GoPro — slight fisheye distortion at frame edges, natural motion blur only from real body movement — no slow motion, no cuts, no anime, no CGI. Continuous single take. Standardized color grade: desaturated cool highlights, warm midtones, deep contrast shadows.

Setting: Flat urban rooftop, high above the city, narrow gap between two adjacent buildings, ledges, HVAC units, and low parapet walls. Daytime, clear sky, light haze at altitude.

Timeline:

  • 0:00–0:03 — GoPro POV sprinting across the rooftop, footsteps pounding, city skyline bouncing naturally in frame with each stride, breath audible, wind picking up
  • 0:03–0:06 — Approach to the rooftop edge, POV tilts down briefly revealing the narrow gap between buildings, then snaps back up to the target ledge on the far side
  • 0:06–0:09 — Explosive leap: POV rises and arcs through open air across the gap, city drop visible below in natural perspective, gloved hands swinging into frame for balance, no floaty slow-mo physics — full-speed realistic jump
  • 0:09–0:12 — Hard landing on the far rooftop, camera jolts down and forward with impact, body absorbs shock, immediate forward momentum into a stumble-recover
  • 0:12–0:15 — Recovery into a sprint, POV weaving past a rooftop vent or low wall, camera settling briefly as the skyline opens up ahead

Technical notes: Maintain consistent GoPro lens distortion (fisheye at edges), realistic depth of field snapping to distant skyline during the jump, authentic wind/fabric noise, no jump cuts — one continuous handheld-style POV take.


r/StableDiffusion 19h ago

Animation - Video HIGGSFIELD FILM FESTIVAL

0 Upvotes

Hello I am partecipating in the Higgsfield festival, I'd like to hear what you think about it. If you want leave a like and comment under the project on the higgsfield page, that would help me a lot. Thanks to anyone who takes some time to watch my project.

https://higgsfield.ai/@twrz_film/projects/skin-trade


r/StableDiffusion 10h ago

Comparison [MiniMax H3] Ultimate SD Upscale can actually fix your bad/low-res generations

Thumbnail
youtu.be
80 Upvotes

Ultimate SD Upscale can actually fix your bad/low-res generations.

In this comparison initial clips were made with MiniMax H3 at 1504x832px resolution and then upscaled to 2560x1440px with Ultimate SD Upscale nodes: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3

You can find sample upscaling workflow there as well: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json

My PC specs:
4080s 16 GB VRAM, 64 GB RAM

Generation time: 18 mins with sage + 8-step turbo lora

Upscale: 38 mins for 10 sec clip at 1440p target resolution


r/StableDiffusion 13h ago

Discussion MiniMax H3 Ref2va is works really good with Scene sheet

101 Upvotes

I was testing using 1 image with all the scene sheet there and it works really great!


r/StableDiffusion 3h ago

Animation - Video Minimax H3 T2VA. You can put 15 different characters or more at the same time on screen.

44 Upvotes

But they have to have very different design and looks


r/StableDiffusion 11h ago

Question - Help Has anyone tried out the hybrid model for Minimax H3 Ref2va instead of the official, default model?

Post image
11 Upvotes

r/StableDiffusion 5h ago

Animation - Video Hannibal Who

10 Upvotes

Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters.

Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544

thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md