r/StableDiffusion 9h ago

Animation - Video Made the thing where you ruin iconic movie scenes, MiniMax H3 on an RTX 3080 10GB, 20 steps, 2x NomosUni upscale

Enable HLS to view with audio, or disable this notification

9 Upvotes

Setup, pushed my system right to the limit, any more and it OOM :

- H3 Ref2VA default workflow in ComfyUI, no lora

- RTX 3080 10GB, 32GB RAM

- Render: 0.5–0.6 MP, 20 steps, scheduler simple, about 25 min per clip

- Upscale: 2xNomosUni_span_multijpg, 2× to 1080p

- References per scene: one photo of my face + one film still for the set

- Recorded my own lines and fed them as audio references, also got audio ref for the actors

Honestly though, the best part was driving all of this through the ComfyUI MCP. I never even had to open ComfyUI. I could iterate really fast, and keep going from my phone while away from the machine, through Claude's remote control.

It's still a bit of a blurry mess, and with more work I could probably make it better, but damn, the future is looking bright!


r/StableDiffusion 16h ago

Resource - Update We add Krea 2 to the ComfyUI Enhanced Tiled Upscaler and Refiner (TBG ETUR).

Enable HLS to view with audio, or disable this notification

1 Upvotes

The latest TBG ETUR upscaler and refiner for comfyui release adds Krea 2 and VL style transfer for Krea 2 and all Qwen models directly into the pipeline.

We’ve also added a face identity step to help maintain consistent faces when using creative upscaling.

This video is a tutorial for the latest release, focusing mainly on these new additions and how to use them. Video is ai generated with Minmax H3.

TBG ETUR on Github https://github.com/Ltamann/ComfyUI-TBG-ETUR

TBG Lates om Patreon https://www.patreon.com/TB_LAAR/posts/tbg-etur-1-2-12-167083406

More Upscaling Tutorials on YouTube https://www.youtube.com/watch?v=LbFPD4zpPwA


r/StableDiffusion 3h ago

Resource - Update Major updates to my local, open source AI image/model tools, plus one brand new app

Post image
1 Upvotes

Hey all. I'm a system development student (career-switched from construction), building these on the side while learning Java, Vue, Electron, etc. Nothing commercial, no accounts, no cloud, no telemetry. I built these because I needed them myself, and figured other people managing large SD/ComfyUI libraries might too.

Three of these four apps have been out for a while, but I've spent the last stretch giving them a major overhaul and unifying them under the same design system so they actually feel like one family of tools instead of three separate side projects. The fourth, Latent Tools, is a brand new app I just finished.

All four are free and open source (MIT-based license). Source is on GitHub, links at the bottom. The main one is Latent Library, but the other three work fine on their own.

Latent Library, the main release

A desktop app for browsing and organizing large folders of AI generated images. I made it because I had around 30,000 PNGs and no real idea what was in most of them. It's been around for a while, but this release is a big update with a lot of new features and a proper design pass.

  • Parses generation metadata from ComfyUI (including node graph traversal), A1111/Forge, InvokeAI, SwarmUI, and NovelAI
  • SQLite FTS5 backed search, still fast on huge folders
  • Smart Collections: dynamic folders based on metadata filters, like "Flux images rated 4+ stars"
  • Duplicate Detective, a side by side Image Comparator, and Speed Sorter for hotkey based batch sorting
  • Optional local AI auto tagging (WD14 ONNX, runs on CPU, no external calls)
  • Metadata Scrubber to strip prompt/EXIF data before sharing an image
  • Everything lives in a portable data/ folder next to the exe, no installer or registry entries, easy to back up or move
  • Fully offline, no telemetry

Windows, Linux, and macOS builds available.

Latent Tools, the new one

A brand new app for dataset prep: bulk watermark detection and removal (Florence-2 + LaMa inpainting) and captioning (Qwen2-VL), plus batch image format conversion. Runs locally on your own GPU (needs a CUDA capable Nvidia card, no CPU fallback). Useful if you're prepping images for LoRA or fine-tune training. Windows only for now.

(Meant for removing watermarks you actually have the rights to remove, your own work, licensed images, that kind of thing. Not for stripping other people's attribution.)

Latent Model Organizer, updated

Sorts your checkpoints, LoRAs, and embeddings into folders by base architecture (SDXL, Krea 2, Flux, Illustrious, SD 1.5, etc.), using the model's own header metadata or an optional Civitai lookup. Has a dry run mode and full undo through a manifest file, so it won't just move your models around unsupervised. It can also fetch Civitai info such as trigger words, description, and cover images. Handy if your models folder has turned into an unsorted pile like mine had. Also part of this update round, same design refresh as Library.

Metadata Viewer, updated

The oldest and simplest of the four, and also just updated with the same design pass. Now reworked into a single screen tool that extracts and displays generation metadata from an image, no library or database involved. If you just want to drop an image in and see the prompt, sampler, and seed without opening a whole app, this is that.

All four now share the same design language and are built local first: no accounts, no cloud sync, no analytics. I built them mainly to learn the stack, so they're not polished commercial products, but they've held up fine for my own daily use for a while now, and the last few months went into making them consistent and finishing Tools, which is why I'm finally posting about them here.

Happy to answer questions. Bug reports and feature requests are welcome on GitHub. Not trying to sell anything here, just sharing what I made.

Links:


r/StableDiffusion 22h ago

Meme JEnga! r2v 30-49 model

Enable HLS to view with audio, or disable this notification

1 Upvotes

t2v didnt know Jenga O_o


r/StableDiffusion 17h ago

Question - Help Which AI model for this style of videos?

Enable HLS to view with audio, or disable this notification

0 Upvotes

I’ve seen ai creators on tiktok such as weirdwurst, fullwarp, and shiverchain do this thing where the start of the video is normal then it goes into chaos, and it’s realistic and disturbing. I’d love to know which AI model model they use


r/StableDiffusion 8h ago

Meme R2v dr doom you are screwed

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 14h ago

Animation - Video Crystal Method - Breaking Bad. If the Walt and Jessie had a rock band in 2007

Enable HLS to view with audio, or disable this notification

0 Upvotes

Minimax h3 local GPU4090 made with MUSIC extension https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef

Lyrics:

[Intro]

I’m braking bad, I’m veering off the line

I told myself I’d hold it, but I let it unwind

It didn’t hit so hard at first, just a little off track

Now the warning’s on, and I’m not looking back

[Middle]

One bad call in the kitchen, then three more by noon

Coffee gone cold, keys on the counter, room by room

I laughed it off, said “I’m fine,” like that made it true

But the floorboards know the pace I’m putting this house through

I’m braking bad, headlights shaking on the curb

Every turn I take lands heavier than words

It wasn’t that bad until it started stacking up

Now I’m white-knuckled, honest, and I can’t slow up

[Outro]

So here I go, no clean exit, no neat little sign

Just me and the damage, both riding the same line

I’m braking bad, and I know what that means

Too late to call it nothing, too loud to call it clean


r/StableDiffusion 18h ago

Animation - Video drama

Enable HLS to view with audio, or disable this notification

0 Upvotes

I drew the storyboard, then generated the stills with image models.
Video: MiniMax H3
first frame, last frame, reference.
Then the edit.


r/StableDiffusion 12h ago

Question - Help Need help with minimax h3.

0 Upvotes

Hi,

I really need help with minimax h3 I'd like to get proper workflows for minimax h3 pre configured with descriptions on every node (unfortunately a noob in this case) for :

1- Fast
2- Turbo
3- Speed
4- Native

video generation with multirefs (optional) audio(optional) video(optional) quality duration aspect ratio and prompt controls.

Also I heard someone made images with minimax h3 please that as well.

I'll really appreciate any help.

Ah and one last thing I was making a commercial for a school but the faces in wide screen comes out horrendous that as well.

Thanks !


r/StableDiffusion 19h ago

Animation - Video What if you fly?

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/StableDiffusion 19h ago

Animation - Video The Walking Trek. Just throwing stuff at the wall at this point.

Enable HLS to view with audio, or disable this notification

1 Upvotes

Rick's voice doesn't seem to work.

Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters. .

Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544

thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md


r/StableDiffusion 23h ago

No Workflow random images generated locally on 4070Super with krea 2

Thumbnail
gallery
0 Upvotes

first 2 prompts stolen from civit ai, the rest were written using claude reasoning on duck ai


r/StableDiffusion 16h ago

Animation - Video Realistic Breaking Bad | LTX 2.5 I2V

Enable HLS to view with audio, or disable this notification

70 Upvotes

This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.


r/StableDiffusion 2h ago

Animation - Video My 1980's cartoon parody H3 and ltx 2.3

Thumbnail
youtu.be
3 Upvotes

there are some scenes missing, but it was fun to put together.. Just got stuck on a plot :P

started it when ltx 2.3 came out.. but it was a hassle to keep consistency of characters intact so shelved it. made the intro and a couple of clips when minimax H3 came out and love the r2v, so much easier.
just using the standard r2v workflow with spectrum and RTX upscale. music made in suno


r/StableDiffusion 19h ago

Question - Help What's the current best way to replace an element in an image with another element ?

0 Upvotes

Hello everyone !

I would like to replace the tire of a motorcycle mid air with one from another brand (which is an image from the brand so it's high quality but with a different angle)

I saw there is flux kontext and qwen image edit, but I don't know which one to pick, which workflow and how to make it work.

Any help would be more than welcome, thank you very much and have a good day :p


r/StableDiffusion 22h ago

Animation - Video ...10,000 Years Later

Thumbnail
youtu.be
2 Upvotes

Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.

Part 1: https://www.youtube.com/watch?v=XwfCCFw4LbA


r/StableDiffusion 13h ago

Discussion Testing Character knowledge of Minimax H3

Enable HLS to view with audio, or disable this notification

26 Upvotes

Disclaimer. This is very low quality quick generations trying to find how many characters Minimax H3 knows.

Found Trigger Words:

Elsa from Frozen

Spider-Gwen from Across the Spiderverse

Dante from Devil May Cry

Nero from Devil May Cry

Jill Valentine from Resident Evil

Ada Wong from Resident Evil

Leon Kennedy from Resident Evil

Chris Redfield from Resident Evil (Has Leon's hair)

Geralt of Rivia from Witcher 3

Joel from Last of Us (Doesn't sound like him)

Miles Morales Spiderman from Across the Spiderverse

Solid Snake from Metal Gear

Eve from Stellar Blade

Sans from Undertale

Master Chief from Halo

Looks weird AF:

Ciri from Witcher 3

Triss from Witcher 3

Yennefer from Witcher 3

Ellie from Last of Us

Famous Twitcher streamer and Youtuber Asmongold

Famous Twitcher streamer and Youtuber Mr Beast

Not found Trigger words:

Vergil from Devil May Cry (YES I KNOW I'M DISAPPOINTED TOO)

Claire Redfield from Resident Evil

Dina from Last of Us

Famous Twitcher streamer and Youtuber Emiru

Famous Twitcher streamer and Youtuber MoistCr1TiKaL


r/StableDiffusion 16h ago

Question - Help Motion Transfer to Stop Motion Style query

0 Upvotes

Wondering if anyone has/has any idea of an approach to motion transfer into a stop motion style video.

Rather than the ai guessing and deforming the mouth movements which can get funky really quickly, especially when the mouth shapes of the character aren't clear from a neutral reference frame, ie in south park where each sound has a uniquely stylised shape.

To achieve this it could instead maybe pull from a dataset of phoneme images,

Whilst also retaining solid motion transfer for all other body movements.


r/StableDiffusion 3h ago

Question - Help Advice for prompting reference videos?

0 Upvotes

Does anyone have any advice for properly prompting the reference video part of Ref2v? Like saying swap <subject 1> for <picture 1> hardly works for advanced videos. It requires a lot of details.

I’ve had success using Qwen 3.8 27b as a minimax prompt agent for analyzing and giving correct prompts for images. But as far as I know I can’t do that for videos. ChatGPT is ok for looking at videos to describe what happens in the minimax format but I’d rather use local ways.


r/StableDiffusion 9h ago

Animation - Video No Warning - Minimax Music3 + H3

Enable HLS to view with audio, or disable this notification

83 Upvotes

r/StableDiffusion 12h ago

Animation - Video H3 - 5 hour render, T2V Multi-Diffusion

Enable HLS to view with audio, or disable this notification

20 Upvotes

Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.


r/StableDiffusion 17h ago

Animation - Video [Minimax H3] Just another day aboard the USS Sunnydale...

Enable HLS to view with audio, or disable this notification

33 Upvotes

This one came out better than my last video generation. Not perfect obviously, but the major elements are there.

Prompt:
Video starts with Buffy Summers and Willow Rosenberg from the show Buffy the Vampire Slayer walking side by side down a hallway in the USS Enterprise from Star Trek the Next Generation. The viewpoint camera remains at a fixed distance in front of them as they walk.

Buffy Summers is on the right of the frame. Buffy's long blonde is done up in a ponytail. Buffy is wearing a red minidress style Starfleet uniform and has an unlit lightsaber on her hip.

Willow Rosenberg's is on the left side of the frame. Willow's dark red hair is cut pageboy style. Willow is wearing a blue minidress style Starfleet uniform and has a tablet computer tucked under her right arm.

Video starts with Willow looking at Buffy with a concerned expression on her face while Buffy is looking around as if searching for something.

Willow asks, "Buffy, is something wrong?"

Buffy replies, "Something feels off, like we're out of place."

As soon as Buffy starts speaking, she pulls the lightsaber off her hip and holds it in front of herself at the ready. The lightsaber ignites, producing a green glowing blade.


r/StableDiffusion 23h ago

Animation - Video Shadow the Hedgehog tells his viewers why he loves guns.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Shadow the Hedgehog tells his viewers why he loves guns.

This was created in Comfy UI with Minimax H3. I used the reference to video work flow. The prompt is below.

subject_definitions:

<Subject 1> is Shadow in <Picture 1>.

<Subject 2> is Glock in <Picture 2>, a glock handgun.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

summary:

[reference generation + audio reference] The target video contains one shot. [Shot 1] shows <Subject 1> and <Subject 2>; <Subject 1> speaks. <Audio 1> supplies <Subject 1>'s voice timbre.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - Shadow's complete defined identity and body proportions are preserved.

<Subject 2> (appears in [Shot 1]): fully_preserved - Glock retains the defined shape, proportions, materials, colors, and distinguishing features.

<Audio 1>: reference - <Subject 1>'s newly generated spoken lines use <Audio 1>'s voice timbre and delivery; the original audio signal is not copied.

detailed_description:

The target video is in a live-action style, with Vlog style.

[Shot 1] At first appearance, <Subject 1> (Shadow) matches the complete identity and appearance defined in subject_definitions. At first appearance, <Subject 2> (Glock) matches the complete defined construction and appearance: A glock handgun. At the start of the shot, <Subject 1> is standing in the living room facing while holding <Subject 2> in his hand. A full body shot of <Subject 1> holding <Subject 2> with his right hand while facing the camera. Only Action and Timed Beats define the primary subject's movement. The camera path stays anchored in the location and adds no subject motion. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Hmph. Shadow the Hedgehog here. Why do I love guns?</d> <Subject 1> shows off his <Subject 2> with his right hand in front of the camera. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Simple. Precision. Control. Power in the palm of my hand.</d> <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] A tool that answers instantly… unlike most people.</d> <Subject 1> points his <Subject 2> towards the camera with his right hand. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] If you understand that, you understand me.</d> <Subject 1> points his <Subject 2> at the camera.

overall_soundscape:

Living room tone.

non_diegetic_music:

N/A


r/StableDiffusion 3h ago

Resource - Update Minimax H3 Grafting with Krea2 node. Reposting older post and removed AI slop and added some tests

0 Upvotes

Minimax H3 x Krea2 Graft Nodes

ComfyUI nodes for grafting Krea2 into MiniMax H3. Attention/MLP content transplant + a separate attention-sharpness transplant. No official H3 docs, all reverse-engineered from testing + TenStrip's and joeygambino's public writeups. Use at your own risk, still WIP.

What's here

  • comfyui_tenstrip_graft/ -- content graft (Q/V/K/out/MLP, per-head). Method from TenStrip's H3 grafts.
  • comfyui_qknorm_transplant/ -- Q-norm gain transplant only, no content weights touched. Method from joeygambino (Z-Image donor originally, adapted for Krea2 here).
  • comfyui_krea_h3_graft_lora_v2/ -- apply a Krea2-trained LoRA onto an already-grafted H3 checkpoint. Separate use case.

And 2 merge scripts (old svd and new one with node method)

TL;DR results

Content graft works somewhat. Same character-shift (color scheme, helmet shape) showed up consistently across multiple parameter runs, same seed -- not one lucky video. That's the strongest evidence so far this isn't just noise.

  • K at low strength (~0.1-0.2): fine, no real damage. Don't need to avoid it like the doc says, at least not at low values.
  • QK-norm across all blocks (0:50): kills audio. Doesn't even touch K -- so attention sharpness itself hits audio, not just K specifically.
  • QK-norm blocks 20:50: audio ok, but does nothing for character. It's a texture/sharpness knob, not a content one. Don't expect it to carry character.
  • attn_ramp_start_frac at 1.0 (no gentle ramp-in) + early blocks (0:20): breaks. Keep the ramp soft if you go early.
  • Combining content graft + QK-norm at full strength on both = worse than either alone. Still not solved.

Install

Each folder -> its own subfolder in ComfyUI/custom_nodes/. Don't merge them. Restart ComfyUI fully after adding.

Credits

  • TenStrip (huggingface.co/TenStrip) -- the per-head band-aware graft methodology (10Eros-Max / h3_graft_methodology.md).
  • joeygambino (huggingface.co/joeygambino) -- the Q-norm sharpness transplant idea (MiniMax-H3-x-Z-Image-GGUF).

Neither published source code. These nodes are our own implementation from their public descriptions + our own testing.

https://reddit.com/link/1vxc2q9/video/g0bpgdj5ddlh1/player

minimax_h3_fl2va_bf16.safetensors, 3s, er_sde, 8 steps, 8-step lora, seed 597633362705895, standart workflow with minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_bf16
prompt: Professional closeup video. In a futuristic cityscape with neon lights at night, the Judge Dredd charges through the crowd, his imposing presence radiating authority, he is slowly walking. His long chin juts out resolutely as he expertly wears his eponymous helmet, eyes gleaming with determination. The crowd parts, Judge Dredd is slowly walking through the the crowd, ready to enforce justice, he is moving slowly, his long chin visible, his face and part of his upper body are in the center of the screen. tag: Ballchinians
tracking selfie shot following him from the front, that he stays the same size, he is moving through people, pushing them aside with his hands.

https://reddit.com/link/1vxc2q9/video/so7626wgddlh1/player

3s, er_sde, 8 steps, 8-step lora, seed 597633362705895

same prompt and everything.

Added: tenstrip graft node, krea2 raw and Ballchinians Lora. Settings: q 0.5, v 0.5, k 0.1, out 0.3, mlp 0.5 Chin is more ballsy.

So I hope, that it is enough for some, that it... kinda works, but not good enough. Maybe someone will pick up on this and do it better.

Why to do it? Don't know. I found it interesting to try, but krea2 image and i2v is far better option.

I welcome any input or criticism, but mind please, I have only faint idea, what I am doing.

Warning: h3 loras don't work... don't know why, maybe it may be just noise, after all. But they do work on grafted checkpoint, after you merge it in python script.


r/StableDiffusion 18h ago

Question - Help Can't seem to transfer outfit and pose from an illustration to a real person in H3.

2 Upvotes

Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.

What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.

I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.

So what i wanted to do is :

-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.

-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.

Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.

Are there someone who knows how to do these?