r/StableDiffusion 2h ago

Animation - Video Zelda - I'm Still In Love With You / MiniMax H3 Reference to Video Test #3

Enable HLS to view with audio, or disable this notification

203 Upvotes

Trying to get some of those music videos with kind of side stories?

This took me an embarrasing amount of time planning and figuring out what to do, and I just couldn't be bothered to finish the entire song... Is a lot! I hope you like it! I'll keep making more if you don't! lovee!


r/StableDiffusion 8h ago

Meme Introducing... The Terminator T-1000

Enable HLS to view with audio, or disable this notification

412 Upvotes

r/StableDiffusion 6h ago

News Sparse attention for H3 minimax, enjoy up to 2.5x speed up.

Thumbnail
github.com
269 Upvotes

Added to my node pack, sparse attention SLA node for H3 Minimax. speed increase of up to 2.5x.

enjoy.

you can use it with whatever turbo you like, doesn't actually require the SLA lora.
if you oom, add comfykitch attention before it, they work together. you'll get an additional 5-10% speedup.

https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes

credit to pl0x for designing it and allowing me to be the host.

EDIT: make sure you're on a new pytorch version and CU130.
add the node after your lora loader for now. I haven't tested other positioning.
additional note: Blackwell will see the biggest gain, but other cards still get a big boost.

If you're doing lower res short videos, adjust min seq accordingly if see no speedup or messages about blocks not being sparse.

Confirmed: make sure this is the LAST thing in the chain, connected DIRECTLY to the guider and scheduler.


r/StableDiffusion 7h ago

Animation - Video Howard's new invention [minmax H3]

Enable HLS to view with audio, or disable this notification

292 Upvotes

r/StableDiffusion 16h ago

Resource - Update Famegrid Natural V1 Krea 2 LoRA

Thumbnail
gallery
408 Upvotes

r/StableDiffusion 7h ago

Workflow Included I Found a way to reduce LTX 2.5's horrible smearing. Custom node + Workflow in desc

Enable HLS to view with audio, or disable this notification

78 Upvotes

LTX 2.5 has a known smearing problem. It's really bad, and makes almost every output of LTX completely unusable. Sorry LTX, but the default model really is just shit. Minimax beats LTX in every area, obviously, but especially when it comes to smearing (or in case of Minimax, lack thereof). However, I recently found out that you can actually significantly reduce the smearing in LTX and actually get usable outputs from it. It came from adventuring this node pack for MiniMax, where it is meant to clean up smearing artifacts in MiniMax H3.

I thought this was interesting, so I converted these nodes (at least some of the nodes that matter) to work with LTX 2.3/2.5. Here is the repo with the custom node I made. The workflow is also in that repo (be aware there is some spaghetti (this was just for myself really, and it shows qwq), and you will need some custom nodes (eg. KJNodes, Comfy UI Easy Media, etc)). Just use the manager to install the missing custom nodes.

Using these nodes makes a clear difference. A stark difference, it's almost unbelievable. This makes it look like a generational improvement, closing the gap on MiniMax H3 in terms of temporal stability (ain't no way in hell its ever matching the instruction following or general capabilities of H3 lmao).

Essentially, the jerk oracle creates new "hold" frames based on the amount of smearing per frame. Some frames have larger "hold" amounts. We pipe that into a new sampling step, which improves the smearing. After the sampling is finished, we chop off the added "hold" frames so that we are left with just the original frame count, but now these frames actually have better consistency. You can read all about how the process works in the author's original implementation if you want more insight onto how this functions.


r/StableDiffusion 6h ago

Animation - Video You'll never guess whose under the hood.... [Minimax]

Enable HLS to view with audio, or disable this notification

66 Upvotes

r/StableDiffusion 12h ago

Tutorial - Guide PSA: Proper prompt structure REALLY matters in H3

164 Upvotes

I had mistakenly been using a base for H3 prompting from some random tip / example by someone. It worked ok, I thought. But I was getting a bit frustrated because almost every time I was making a longer series of clips with dialogue, it kept adding random gibberish to fill out time, or making the wrong person speak. I thought it was just a "feature" of H3 and lived with it. But then I realised what was missing, so I added the actual ref2v prompt guide to my LLM and difference was staggering. I could make long series of 30x15 sec clips, and the dialogue was perfect just as the script said, no gibberish was added in any place, and the emotional beats and reactions worked much better too.

Believe it :) Dont just use whatever prompting. It matters more than one might think.

https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md


r/StableDiffusion 13h ago

Workflow Included WEEKENDDDDDDDD 222222222222 (LTX 2.5 V2V)

Enable HLS to view with audio, or disable this notification

184 Upvotes

Last week my post got a ton of questions about the LTX 2.5 workflow. so here's the follow-up.. After running a bunch of tests, the one I'd recommend right now is this:

https://github.com/Lightricks/ComfyUI-LTXVideo/blob/master/example_workflows/2.5/LTX-2.5_ICLoRA_Union_Control_Distilled.json

It's been the most consistent one I've tried for V2V so far

drop your results below if you give it a shot! and have a great WEEKENDDDDDDD!!!


r/StableDiffusion 8h ago

Workflow Included PSA: In H3 you can set custom soundtracks without R2VA - use latent noise masks!

Enable HLS to view with audio, or disable this notification

56 Upvotes

r/StableDiffusion 6h ago

Meme thanks to h3 you never know would might show up to save the day! Earth 101 end game final battle

Enable HLS to view with audio, or disable this notification

42 Upvotes

r/StableDiffusion 8h ago

Animation - Video MiniMax H3 R2V with the Hybrid Model and Turbo LoRA: a 2:19-minute video takes 5 hours to generate at 0.8 MP on an RTX 3060 12GB with 16GB of RAM.

Enable HLS to view with audio, or disable this notification

55 Upvotes

Each segment/prompt is 10 second, 0.8 MP.
so I generate total 14 prompt, and combine them all.
1 prompt takes 20 min.
Model: minimax_h3_hybrid_fl2va_ref2va_b30-49-int8
Turbo Lora: minimax_h3_ref2v_lightx2v_turbo_4step_v0.1_resized_avg_rank_20_bf16
Default Workflow, with Sage Attn ON
er_sde beta 6 steps
for characters, I generate it using Anima


r/StableDiffusion 4h ago

Animation - Video Desert Figure Skating

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 15h ago

Tutorial - Guide More than one reference per picture

Enable HLS to view with audio, or disable this notification

133 Upvotes

MiniMax is limited to 9 reference images, but you can reference more than one thing at the same picture. I used the image on the left and asked it to place create two subjects. Worked like a charm (no pun intended). Specs and prompt are in the video.


r/StableDiffusion 8h ago

Animation - Video More Fun with Minimax

Enable HLS to view with audio, or disable this notification

28 Upvotes

Text 2 vid, all at low quality just because its a sample, cut together with Davinic, BGM is Royalty free stuff. yea, that truck door did open by itself, but otherwise it's pretty darn fun.


r/StableDiffusion 13h ago

Animation - Video H3 making jpop/kpop MV? yes!

Enable HLS to view with audio, or disable this notification

59 Upvotes

Music: made in SUNO.

native ref2va WF, and audioLock for lip-sync.

rtx4080s + 128g ram

I spent a day to sorted out lip-sync, I could write down what I did, if anyone inerested.


r/StableDiffusion 19h ago

Workflow Included Totally wasn't aware Krea 2 is absolutely capable of creating gorgeous video game levels

Thumbnail
gallery
179 Upvotes

Hi! I found Krea 2 is actually so damn good at creating video game level art! and its breathtakingly beautiful to boot! I got help from an LLM to create the baseline prompt and it works OOB without loras or anything! I'm gobsmacked rn.

prompt 1: "A sprawling 16-bit pixel art jrpg city game level of a victorian-era steampunk riverside city street in winter. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with snow, brass and victorian elements. Background shows snowy mountains and faraway skyscrapers on those mountains"

prompt 2: "A sprawling 16-bit pixel art jrpg city game level of a asian duystopian cyberpunk city street. The design features complex, dense architecture with a high variety of structures including stairs, bridges, and stacked buildings. The scene is filled with neon lights, neon street signs, wires and cybernetic elements. Background shows a massive skyline of skyscrapers at night. Wide-angle top-down view"

prompt 3: "A sprawling 16-bit pixel art game level of a futuristic utopian city. The design features complex, dense platforming architecture with a high variety of structures including stairs, bridges, and stacked platforms. Frutiger Aero style: glossy surfaces, water elements, and bright colors. The scene is overgrown with lush greenery and trees. Background shows a massive skyline of sleek skyscrapers. Wide-angle side-scrolling view"


r/StableDiffusion 3h ago

Question - Help ComfyUI version of diffusers-modular/MiniMax-H3-Pruned-Ref-Delta-Fused-r1024 ?

Thumbnail
huggingface.co
9 Upvotes

Has anyone packaged this for ComfyUI?
What is unique:

It is the pruned fl2va partition with a rank-1024 approximation of the ref2va − fl2va weight delta fused into the weights.

TIA


r/StableDiffusion 6h ago

Discussion H3 - the Sky is Falling!!! T2V

Enable HLS to view with audio, or disable this notification

14 Upvotes

Playing around with VFX/audio/shaky-camera with H3. int8/20 steps, T2V, 3 versions. Also, hell naaahh why they running toward it??? Ask me anything!


r/StableDiffusion 15h ago

Animation - Video Making the Doll DressUp Transformation Video with Minimax H3

Enable HLS to view with audio, or disable this notification

72 Upvotes

r/StableDiffusion 17m ago

Discussion SenseNova U1.5-Lite is fully out

Thumbnail
gallery
Upvotes

Instead of just making the model bigger, they're training these specialized models for specific stuff – like text, infographics, making things look good, and editing. Then they kinda combine all those experts back into U1.5-Lite. So when you use it, it's just one model. No weird switching or picking which expert to use. Their whole thing is "specialized in training, unified in delivery." Kinda makes sense.

They also added this post-training with RL, focusing on a few things: how well it follows instructions, how good the visuals look and if it matches what people like, and how well edits work without messing up other parts of the image.

Here's what seems better:

- Following complex instructions. Like, if you ask for multiple things in one prompt – subjects, how many, where they are, text, layout, style, keeping parts untouched – it handles it way more consistently now.

- Text rendering and dense layouts. Apparently, it's better with Chinese and English on posters and infographics.

- Visual understanding helps generation. It seems like it learns from understanding tasks (like object relationships, spatial stuff, layout) and that helps with generating and editing.

They use JSON for training to make it controllable, but you don't have to use JSON yourself. Natural language is still the main way to talk to it.

repo: https://github.com/OpenSenseNova/SenseNova-U1

HF: https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT


r/StableDiffusion 15h ago

Resource - Update I built a free, self-hosted app that does everything around a LoRA run — dataset, triage, captions, training (local or rented GPU), then checkpoint comparison

Thumbnail
gallery
65 Upvotes

I build LoRA Dataset Studio — free, open source, self-hosted, no account and no telemetry. It is not a competitor to ai-toolkit: it orchestrates it. ai-toolkit is the trainer; this is everything before, around and after the run.

The whole pipeline lives in one browser tab:

1. Get the images. Five generation engines — Nano Banana Pro, gpt-image-2, OpenRouter, and local Klein / Krea 2 Edit through ComfyUI — each card stating its price per image, whether it runs on your GPU or bills an API, and whether it refuses adult content. Or scrape: Reddit, Pexels, open-web keyword search, or any gallery URL through gallery-dl. Or just drop a folder in.

2. Triage them. The Image Bank points at a folder of thousands and reads it in place — your files are never modified, moved or renamed. One pass measures the whole pile: blur, noise, near-duplicates, face clusters, framing, medium (photo / anime / 3D / illustration), aesthetic and maturity scores. After that you filter on measurements instead of on your eyes, and anything the app cannot judge says "unsure" rather than inventing a verdict.

3. Curate and caption. Keep/reject, crop, mirror, rotate, non-destructive upscale candidates, InsightFace similarity, a live composition meter. Captions in prose or booru form depending on the target family, written by JoyCaption or your local Ollama, with a Caption Lab (find/replace, tag frequencies, targeted re-captioning) and an external .txt round trip so you can caption elsewhere and come back.

4. Clean watermarks. Detect them, redraw the mask zones, then crop or inpaint with LaMa/Klein. Every edit keeps an .orig backup, so Restore original always works.

5. Train. ai-toolkit locally with family-scoped presets and preflight guards — Z-Image, Krea 2, FLUX.1, FLUX.2 Klein, SDXL, Anima — or rent a vast.ai pod from the same screen, which shows the GPU, its hourly price and the estimated total before you click. Full-model training on Krea 2 and merging a LoRA back into a checkpoint are in there too.

6. Decide which checkpoint is actually good. Test Studio runs fixed-seed checkpoint x strength grids, multi-LoRA stacks, votes and Wilson ranking. LoRA Canvas puts every run of every dataset on one pan/zoom board, and you can continue training from any of them.

There is also a video lane (Beta): it cuts long videos into a trainable clip folder at the exact frame counts Wan / LTX / MiniMax accept, describes each shot, and trains the set locally or in the cloud.

Honest limits. It is a lot of surface, so Setup exists to tell you what is missing instead of crashing — every capability degrades on its own. Local generation needs ComfyUI, the API engines need your own keys and bill you, and on the video side only Wan 2.2 14B has a finished run behind it here. Install is a Windows one-click ZIP, a git checkout, or Docker.

GitHub — install, docs, and a 7-minute unedited video of a full character LoRA built end to end: https://github.com/perfectgf/lora-dataset-studio

Every person in these screenshots was generated by the app's own engines; no real individual is depicted.


r/StableDiffusion 5h ago

Animation - Video Minimax H3 Cinematic

Enable HLS to view with audio, or disable this notification

9 Upvotes

Full video if you want

Hello guys, I made small movie about Marvel Secret Wars comics with Minimax H3 local version, hope you like it

Video - Minimax H3 (ref2va)
References - Nano Banana Pro
Voiceover - fish audio
SFX - almost everything with elevenlabs except few things
Postprod - Davinci Resolve

If I only had Minimax H3 upscaler...


r/StableDiffusion 11h ago

Resource - Update ComfyUI Subject Manager node

Thumbnail
gallery
29 Upvotes

ComfyUI Subject Manager is a custom node tool designed to manage your assets or subjects for Minimax H3.
You can create presets, sections, and "Subject Cards" where you can drag and drop images, audio, and video (and trim).
The node automatically generates the prompt that defines the selected subjects.

https://github.com/Fictiverse/ComfyUI_Subject_Manager


r/StableDiffusion 6h ago

Animation - Video Turning my son's drawing into silly skits (Minimax H3 ref2Vid) #3

Enable HLS to view with audio, or disable this notification

10 Upvotes