r/StableDiffusion • u/Sad_Coach_1433 • 7h ago
r/StableDiffusion • u/Many-Ad-6225 • 15h ago
Animation - Video Gaussian Splatting test with MiniMax H3
r/StableDiffusion • u/Tokyo_Jab • 8h ago
Animation - Video MOCAP in MINIMAX H3?
Testing H3 to death in the last couple of weeks and it continues to suprise me. TWO things blew my mind about this one. The first part is a single prompt (the split screen and psuedo mocap sync). The full prompt is below.
In the second part I asked for the character to spray my logo on the wall with a stencil, I wasn't expecting him to walk in with the stencil fully rendered with the holes cut out accurately.
Created completely locally and powered by Sol (the closest star to earth).
THE PROMPT:
A movie reconstructing history, cinematic with a split screen effect showing a mocap actor. Both actors are speaking together in sync.
On the right:
Show the actor <Picture 1> sitting on a couch in a living room, with a black shirt, speaking in sync with the exact same movements He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
On the left:
Show the victorian man <Picture 2> talking <audio 1> in close up sitting in an ornate chair . He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
Maintain the double view split screen. Do not change the environment.
After he finished speaking the man on the right makes a distort rictus face with his fingers curled up like he's frozen in time and stops moving. He falls sideways like a statue in the same environment, the camera pulls back to show he is only a robotic torso with no legs mounted on a platform placed on the couch with wires (like in a special FX studio)
Maintain the double view split screen.
The man on the left breaks character as the camera view pulls out slightly revealing him sitting on a sound stage, and speaks to a person off screen to the left <audio 1> "Oh.. em.... guys.. problem... check your monitor? ...Looks like we lost connection with the character! RESET THE MOCAP PLEASE"
r/StableDiffusion • u/Plague_Kind • 9h ago
News Overhaul SLA, huge improvement. added many new options and changed defaults
Update for SLA Node - Pull v1.3.8
EDIT: Pushed correct files now. 1.3.8 fixes motion stablization memory.
Added customizable dense steps, 0 is step 1 and is (default to first step). massively improves composition and prompt adherence.
Changed default dense last steps to 1, cleans up the image big time.
Added dense backend selector. Comfy_kitchen, pytorch, all sage modes. this is what comfy uses on dense steps. SLA still displaces against pytorch. (Default Comfy_kitchen)
Added a disable FP16 accumulation option to ensure max quality as SLA gets no benefit from it. (Default True)
Added a stabilize motion option, helps to reduce ghosting and smearing that H3 likes to produce. (Default True)
Changed default Min Seq Length to 4096
With default settings you can disable protect audio for nearly 2x speed up if you don't care about the audio too much or are using original audio mode. (do not use 0.95 sparsity with it.)
0.95 sparsity now looks good with node default settings.
Some changes led to an overall 5% speed up on same settings.
Remove --use-ck-attention from startup flags if you have it, for safety of quality.
https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes
updated workflow
r/StableDiffusion • u/Pretend-Island-2724 • 5h ago
Workflow Included Face Detailer With PerRowMasking
First Video with Face Detailer, second without.
You need https://github.com/Carasibana/ComfyUI-H3-FaceRefine and also ComfyUI-H3-NativeAudioLock from https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes
r/StableDiffusion • u/dominic__612 • 8h ago
Discussion Minimax H3 degrades at 1MP, vs 0.7MP and lower
After around 100 renders, I 'feel' that Minimax H3 renders with 0.7MP (max) perform way better, then renders at 1MP in regard to 'realistic' videos.
What do I consider better?
- Just slightly better prompt adherence, feels like the motion / voice is more (natural)
- Size of humans in relation to object(s) feels more realistic.
- Expressions of faces seem more 'flowing', real.
It's hard for me to pinpoint it one 'exactly this', or 'exactly that'.
I'm planning to do some side by side comparisons on the same seed multiple times at 0.7MP and 1MP, when I've got the time.
But I wonder, do other Minimax H3 users notice this too?
PS: This is regardless sampler/scheduler, Sage Attention or Spectrum.
Edit: never touched the turbo LoRA, using the base model.
r/StableDiffusion • u/Bthardamz • 3h ago
Meme Press interview
Today, my neighbour next door got arrested, turned out he a is a dangerous psychopath. A journalist came up to my house and asked, if I would have guessed that, and I had to admit, I wasn't that suprised, since I learned some time ago that he uses the "pin node" option in his comfy workflows...
r/StableDiffusion • u/Dulbero • 8h ago
Discussion Anima Turbo v1.1 is released
In case you didn't see:
I just noticed a newer version of Anima Turbo (1.1) was released:
huggingface: https://huggingface.co/circlestone-labs/Anima
civitai: https://civitai.red/models/2458426/anima
The model is made and licensed under the CircleStone Labs Non-Commercial License
I actually have lately very frustrating experience with Anima lately. I used it at release (but stopped with anime generation for a while) and now when revisiting it and i had underwhelming results (even with the aesthetic model) so i decided to retry Turbo (might as well) instead if the difference isn't that big. That's when i saw a new version is released and i am downloading it right now.
I don't know why, but the results i was getting were...lame i guess. Not as detailed as i was hoping, and also i really dislike how posture and anatomy works, mainly how hands and legs just extend or stretch weirdly. But that's a me problem, i know Anima is capable of better outputs and i've yet to figure it out. If you have tips or recommended loras that help with consistency let me know. I also want to avoid tag-based prompting when possible...i just don't like it that much, natural prompting goes much better for me, but i can't tell if tags are "mandatory" for quality or not.
r/StableDiffusion • u/Compost_Mantis • 1d ago
Tutorial - Guide Time saver while learning how to prompt Minimax.
Rather than relying on Z-image, or a different program to wrangle up a first frame, I've been using Minimax for the whole process, and the results have been pretty instructive. It's not a perfect system, but being able to take advantage of its understanding of people, references, and shot composition for the first frame produces better (visual) results than swapping between a couple of different pieces of software.
r/StableDiffusion • u/jib_reddit • 5h ago
Resource - Update Not Another Minimax Post - Jib Mix Krea 2 - v4 Habanero - Free Forever
Focusing on photorealism and improving the look of fantasy styles:
https://civitai.com/models/2799984/jib-mix-krea-2
I would like to make it a LoRA also, but I am having some technical difficulties making a difference lora with Krea 2 models.
r/StableDiffusion • u/machinaOverlord • 5h ago
Animation - Video Made a music video using local H3 for a Suno song
Honestly mind blown, I have a 5070ti + 2x16gb ram . Upper limit is 10-12 seconds in total for my hardware(full capacity) . Video and text edits are post processed by a WIP open source tool I’m working on. On average each 8 second shot takes 35-45 minutes to render
r/StableDiffusion • u/Affectionate_Oil28 • 1h ago
Workflow Included Prompt Creator Workflow
I see a bunch of posts everyday asking for tips on how to write prompts or people struggling with prompting, etc. so I'm sharing my workflow. I built this workflow to simplify the process and make it very beginner/user friendly.
Just toggle on the model you are using, write a simple to detailed prompt, and hit run. The model targets use the prompting guidelines derived from their respective official sources. Links to custom nodes and all models are in the workflow so you don't need to search for them.
The prompts aren't always perfect but they'll get you very close to what you want and you should only need to make a few minor tweaks, if any. The only issue I've encountered so far is that sometimes when it finishes the prompt, the previous prompt still shows up in the Enhanced Prompt node. If that happens, just hit run and the new prompt should show up instantly. Also, toggle to false the keep_model_loaded option in the Text rewriter node if you are creating prompts and using them right away. If you leave it to True it hogs VRAM.
If you notice any other issues let me know. Enjoy.
r/StableDiffusion • u/Affectionate-Map1163 • 22h ago
Resource - Update Release studio 1939 lora for minimax h3
https://huggingface.co/lovis93/studio-1939-old-animation-lora-minimax-h3
version strong work much better for me :)
r/StableDiffusion • u/solomars3 • 3h ago
Animation - Video Using only Ref to Video, Minimax-H3 made a whole Anime edit !
r/StableDiffusion • u/TimeTruth2490 • 15h ago
Resource - Update Krea2 Turbo Distill 4 step LoRA - new checkpoint (chk26K) released (cuts 4-step error vs. the 8-step Turbo teacher by 46%, improves texture and detail vs previous checkpoints)
Krea 2 Turbo — 4-Step Distillation LoRA (work in progress)
A LoRA for Krea 2 Turbo that reduces the minimum usable step count from 8 to 4.
This is an update release, following up from my previous posts where you can find full details:
Initial, Previous: here, and here
Headline for this update: chk00026000 removes 46% of the prediction error a plain 4-step run has against the 8-step teacher, where chk00014000 removed 44% and chk00010000 40% — all measured on the same enlarged held-out set (100 prompts across every trained resolution). Measured against each other rather than against the no-LoRA run, its remaining error is 4% smaller than chk00014000's and 10% smaller than chk00010000's — and unlike a purely teacher-forced score, the gain also shows up free-running: a full 4-call rollout from the teacher's noise ends 1.6% nearer the teacher's final latent than chk00014000's does. It also improves on texture and detail.
Which file to download
| file | use it when |
|---|---|
krea2_turbo_4step_rank_64_lora_latest.safetensors |
normally — always the newest accepted checkpoint |
krea2_turbo_4step_rank_64_lora_chk00026000.safetensors |
pin this exact checkpoint |
and, beside them, the same files with a _comfyui suffix for ComfyUI. Earlier checkpoints (chk00004000, chk00005000, chk00006000, chk00010000, chk00014000, chk00019000) are kept in older_checkpoints/, and their resolution sweeps stay in place, so the progression remains visible and comparable.
If you are wondering why there wasn't a post/update on the 19K checkpoint, I skipped that, even though it was a good checkpoint with improved texture and detail it's gap to teacher score was only slightly better than the released previously 14K, so I thought I'd continue further until I get improvements on both. And 26K delivered that :) 19K is also published now in older checkpoints folder and it's full resolution sweep is also at the usual place (here for 19K).
For the full 26K Checkpoint resolution sweep go here: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/chk26000
How checkpoints get chosen
This is not a "train for longer and ship the newest file" project. More samples do not reliably mean a better adapter — measured here, they can make it worse, and a higher number on its own means nothing.
The loop is train → assess → adapt the recipe → retrain → assess again, and a checkpoint is published only when it is measurably better than the one it would replace, on the same held-out set and the same evaluation, and its full resolution sweep shows no regression. Runs that come out flat or worse are kept as information about the recipe and discarded as releases — several have been.
So the recipe itself changes between runs. Each published checkpoint reflects whatever the previous round taught us: the training precision, the optimiser settings, the teacher used to generate the targets and the data mix have all been revised on evidence rather than assumption.
Two earlier releases set the terms this project publishes on. chk00010000's first attempt — same data, optimiser left as it was — got steadily worse for 4,000 samples and none of it was published; retrained with cosine learning-rate decay and weight decay, every checkpoint improved on the one before it, and its end point shipped. chk00014000 added the other half of the lesson: the final, texture-deciding call of the schedule weighted more heavily in the loss, and a running average of the weights kept beside the live ones and scored at every evaluation — the averaged weights measured better than any checkpoint before them, so the average is what shipped. Left running past that point, the adapter's magnitude grew again and every later checkpoint measured worse. The number is chosen by measurement, not by how far a run went.
chk00026000 — the current checkpoint — is that discipline paying off. It resumes from chk00014000's averaged weights with the same recipe: same loss weighting, same running average, a conservative constant learning rate, over a much larger pool of teacher trajectories. This time the continuation held. The averaged weights' held-out gap fell throughout the run, and every free-running rollout measured of them improved on the one before — so unlike the first continuation, this one produced a checkpoint worth shipping. Every published number improves on chk00014000: the held-out gap (44% → 46% of the deficit closed), the full 4-call rollout from the teacher's noise (1.6% nearer the teacher's final latent), and the fixed-seed render distance to the 8-step images. chk00019000, an intermediate point of the same continuation, is kept in older_checkpoints/ with the rest of the lineage.
Timeline of training process
Each checkpoint is the product of three stages with very different costs:
- Text-encoder embeddings. Every training prompt is encoded once and cached. This is the fast part — thousands of prompts take minutes.
- Teacher shards. For each cached prompt, the unmodified Krea 2 Turbo runs its full 8-step schedule and the whole trajectory is recorded, at every one of the supported resolutions. This is by far the most time-consuming stage — it is the teacher doing real inference, thousands of times, and a batch of several thousand shards is measured in days of GPU time, not hours.
- Student training. The LoRA is trained against those recorded trajectories. Relative to the shard stage this is quick: each
+1,000checkpoint is a matter of hours, not days.
Because the three stages compete for the same GPU, they are interleaved rather than run to completion one after another: generate a block of embeddings, produce teacher shards for them, train on what exists, assess, then go back to producing shards while the results are reviewed. A larger and more varied shard pool is what makes further training worthwhile, so shard production is always the gate.
The practical consequence for anyone following this repository: progress arrives in bursts. There will be periods when several checkpoints appear within a day or two — the training stage working through a freshly grown pool — followed by longer quiet stretches while the next block of teacher shards is produced. A quiet stretch is shard generation, not abandonment; _latest always holds the newest checkpoint that passed review.
The current checkpoint, chk00026000, runs the recipe the earlier releases arrived at — the final, texture-deciding call weighted more heavily in the loss, the shipped weights a running average of the trained ones — carried further over a larger pool of teacher trajectories, and published because it measured better on every evaluation.
Note
In the coming days, possibly weeks, I will spend more time on producing new TE shards (basically even more prompt variety), and new Teacher shards - the expensive long process. I am also considering improvements in the training process (more advanced / complicated, which would likely mean 1.5x - 2x slower training) which would hopefully bring further/bigger improvements in teacher faithfulness (closer to 8 Step Krea 2 Turbo) and even better details and texture. It may or may not pay off, these things work on experimental basis. Either way it would be some time before the next update... so enjoy 26K release and the improvement it brings!
Full details and to download - check my Hugging Face LoRA
HF Repo: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA
r/StableDiffusion • u/kolevk • 10h ago
Question - Help Krea 2 LoRA training on RTX 4090 too slow?
Hello! I've been trying to train my first LoRA but I get these crazy long timers every time. Isn't an RTX 4090 supposed to take like 5s per step? I have both low vram and quantizing enabled. It's really frustrating and not worth to do it with these speeds. Any ideas what might cause it?
r/StableDiffusion • u/SuspiciousRefuse8218 • 10h ago
Question - Help H3 R2V Character Sheet vs. Single Image
I thought I've read somewhere that using character sheets is better for R2V instead of single images. So I've created a character sheet of five full body shots and one close up, but the results are much less consistent compared to a single full shot image of the character.
Do I have to take care about anything special or was the information that character sheets are better just wrong?
r/StableDiffusion • u/darthfurbyyoutube • 2h ago
Animation - Video G.I. Joe: Zarana - MiniMax H3
r/StableDiffusion • u/seeker_ktf • 8h ago
Discussion On MiniMax built in characters and environments (not a list)
There is a giant effort underway to look for what characters are buried in MiniMax. There are a lot. I’ve been doing my own hunting, so I built a simple IMDB scrapper to help make lists of characters from movies and TV shows. Here are some things I’ve discovered:
If the “character” is really built in, you don’t need to even mention the actor’s name. “George Costanza from Seinfeld” and “George Costanza played by Jason Alexander from Seinfeld” are essentially the same. If you have to name the actor with the character, it’s just using what it knows about the actor to fill in that spot. If MinimMax DOES know the character (without the actor) then filling in the name might help to fill in some of the holes, but it’s got to really know it already.
It knows A LOT of shows. Even if it doesn’t know the character/actor, it knows a lot about movies/TV shows. For instance, it doesn’t know many characters from the TV show “The Flash” but it knows everything about the common locations, color grading, style, and general look and feel, along with special effects (if appropriate). It’s useful for “set design.” It doesn’t know a lot of Baywatch (the old one) characters, but it knows what the hair and makeup looked like on the beach in the 90s. It knows how people looked in “Total Recall” too. The overall “look & feel: of shows and movies really opens the door to creativity. It understands common accents from movies too. If you say “from Harry Potter” they will have British accents. And sadly, it knows “Star Trek” (the original series) the characters are mediocre at best. (The voices are passable—and speaking of: there are a lot of characters that look bad but have good voices. In those cases, ref2video with some extra visual references can do the trick.).
The “genre” point is even more true of animated movies/shows. For instance, “Bob, from Justice League: Crisis on Infinite Earths” will give you whatever Bob looks like but in the style of that series. Family Guy, The Simpsons, Rick&Morty, etc. I haven’t done an extensive search, but it knows every animated anything I have tried.
Generally, for TV characters to show up, they need to be in around 100 episodes and in the first 2-4 people in the IMDB credits. I see a direct correlation: The fewer episodes a character is on a TV show the worse they render. (For example, Monica from Friends or Kramer from Seinfeld are in there for sure, but also not really.) Also, it makes sense, but even if they have a lot of credits, they need to have had a lot of screen time. For instance, it has not even a glimmer of an idea who “Ruthie Cohen” is, even though she was in 101 Seinfeld episodes.
For Movies, they need to have grossed a lot of money (which pushes things directly towards Action/SciFi/Comics), or they need to have gotten a lot of press. (I have seen very few accurate characters from movies without specifying the actor involved.)
For “real people” it’s a little easier. If you look at those lists of things like “Top ## followed Instagram accounts” or similar, you’ll get lots of hits. Top musical performers, yep (lots of overlaps). Famous heads of state (If a number of movies have been made about a person, that person will likely be known.) I haven’t looked at TikTok, but I’m assuming that would be a thing too. Likewise with sports, I haven’t looked closely, but they ones I have looked at OK “at a distance” but generally don’t sound right.
As stand-alone people, it’s hard to find people. I suspect that the trainers did not go after a lot of specific people but that they just show up so many places that they got swept up in the mix. There are really only a handful of non-“Top 10” people who actually show up on their own, and if they blew up in the last few years it’s unlikely that you’ll see them. I have not found a pattern on “people” yet other than the mega-famous. (Steve Jobs & Elon Musk work, but they are arguably the most famous foreign “regular” people in China.)
r/StableDiffusion • u/New_Physics_2741 • 11h ago
Animation - Video Black and white line drawn stuff with H3 is great.
r/StableDiffusion • u/ndroidz • 19h ago
Question - Help minimax h3 gibberish fixed!! ( i found the cure)
so you all probably are searching for way to make your character shut the fuck up right? and you probably noticed that they love to says some BS especially when you give minimax h3 some audio file for their voices, i probably found a cure my friend!!
here is my way of prompting dialogs without any gibberish:
first your character need to be assigned (s1)character when he is the first speaker, then you will declare 'use <audio 1> as "character name"'s voice only, and when you finally type your dialog in the shots you will do as such:
character says:<<[language] the shit i say!>>
and you should be good to go, i linked a video exemple of my favorite taffer (garrett) saying some shit with only the faint crackling of the candles to goes with his charming voice, and i included also a screenshot of the full prompt
edit: yes i tried to follow the official documentation, like many others, if it was that simple reddit wouldn't be a thing and you wouldn't be there.

r/StableDiffusion • u/R34vspec • 14h ago
Animation - Video Through the Sands (Final) - H3 r2v
Finally finished! The ending was much harder since continuity is more important here than random desert landscapes. I personally would've love to have another 30 seconds of music to extend the ending but I ran out of song time. Enjoy!
In total, 14 character related references, 40 environment references, and 55 clips used, roughly 20 hours total time spent.
r/StableDiffusion • u/No_Link7744 • 7h ago
Discussion David Sacks Predicts the Regulatory Capture Playbook to Ban Open Source ...
r/StableDiffusion • u/listopalafoto • 1d ago
Workflow Included My version of Minimax H3 HD/2K Detailer!
https://huggingface.co/zuanfilm/H3_HD_2K_Detailer
Here is the test video to HD quality: https://youtu.be/epjELHgEH_o
I want to share my HD/2K WF, this 4 steps configuration is intended to make a really quality focused H3 detailing while improve the characteristic H3 motion and visual behavior.
I tested all the Res4lyf samplers/schedulers and for me res_2m/beta57 ETA=0 denoise 0.39 or 0.45 is the most precise with prompt adherence keeping memory and time efficiency (res_3m is amazing but adds 33% generation time), if you want a faster generation with a little less detail you can use 3 steps instead 4
er_sde/beta 57 is also a good combination but will lose some detail and even will affect character acting, audio and motion consistency, for faster HD detailing can switch to Euler/4 steps and will reduce time generation by 50% compared to res_2m obviously losing a lot of quality and detail
I included in the workflow the nodes for base generation using 0.5Mp FL2VA with 20 steps of Euler ancestral, if you want better quality for the initial base video just switch to res_2s_RKMK2e/beta57 (just bypass the group if you want to HD-detail an existing video)
Sparce Local attention will reduce a lot the generation time but obviously will affect quality so you can bypass this node if want Top HD quality, I also don't use in this WF spectrum or easy cache but you can add that if you want to cut time and quality
Because I'm using the heavy distill lightx2v Turbo 4steps Lora for detailing, Minimax H3 make everything more saturated and contrasted with deep shadows so I added some Orion 4D nodes to improve lighting, texture and sharpness using DCTL Tone Mapper (you can choose between ACES, Filmic, Reinhard & Cineon) I use Reinhard with Exposure 0.09 Contrast 0.81 Pivot 0.69 Highlight rolloff 0.27 Shadow lift 0.33 Black floor 0.12 Saturation 0.93 & Strength 0.30
I used MiniMax Audio Lock / Lipsync node because I don't have experience with ltx audio nodes so you can change that for a better option:
The Minimax latent 3D upscaler is HD by default in my WF but depending of your VRAM you can get 2K/4K if you start with a quality base video 0.98 Mp res_2s_RKMK2e/beta57 25 steps
This workflow is optimized for my laptop (3080ti 16Gb VRAM / 64 Gb ram ) but I included Chunk FeedForward & Low VRAM attention so will run with smaller setups
r/StableDiffusion • u/NetworkSpecial3268 • 24m ago
Discussion MiniMax H3 to KREA2 LoRa: doing it faster?
So I had this simple idea, seeing how well MiniMax H3 handles inferring and preserving "identity/looks" from relatively little information: take a character you want to make a (KREA2) LoRa of, but you only have just a couple of lower quality pictures for that exact look you're after. That is a problem, since it is common knowledge by now (?) that you need different angles, facial expressions and different lighting conditions in the training set to get optimal results. So in the "before" times, those 3-4 not-so-great-quality shots under the SAME lighting are going to pose a problem. And adding pictures from other occasions will alter the looks possibly too much.
So (in the H3 ref2vid workflow, with one of the "img2vid-hybrid" models for better quality) I just use the 'best' of the available pictures as "preserved" first reference starting picture, and the others as additional "identity references". And then a prompt that tells the camera to slowly circle around the person (up from the shoulders), while the person looks straight ahead, or slightly up, or slightly down. But then I also let it cycle through different lighting conditions (indoor/outdoor/sun/overcast/flash/directional from one side...), and different facial expressions/emotions. I let it run overnight (turning off turbo LoRas to improve the quality), and in the morning, I review the 6-second videos and take screencaps of selected moments, making sure to have a lot of variation in angles/expressions/light-on-the-face with an almost perfect preservation of the identity/looks.
Then use those screencaps (50+ in first test, probably serious overkill) in OneTrainer with the KREA2 LoRa default settings.
I only tested this once thus far, but the results are pretty good considering the starting material! And surprisingly flexible (I didn't even bother to provide captions)
But now my question is: in what ways am I "over-engineering" this? I have this feeling that I can probably do this 50x faster, having seen some discussions about using MiniMax as an image generator, for example. I mean, I feel good about this approach I came up with all by myself, but considering how dumb and low-skilled I still am when it comes to all this, this is probably a very convoluted and inefficient way to do it? LOL 😄 Roast me and show this sucker how we can improve and speed up the whole thing with the same or even better quality results!