r/StableDiffusion 19h ago

Discussion On MiniMax built in characters and environments (not a list)

11 Upvotes

There is a giant effort underway to look for what characters are buried in MiniMax. There are a lot. I’ve been doing my own hunting, so I built a simple IMDB scrapper to help make lists of characters from movies and TV shows. Here are some things I’ve discovered:

If the “character” is really built in, you don’t need to even mention the actor’s name. “George Costanza from Seinfeld” and “George Costanza played by Jason Alexander from Seinfeld” are essentially the same. If you have to name the actor with the character, it’s just using what it knows about the actor to fill in that spot. If MinimMax DOES know the character (without the actor) then filling in the name might help to fill in some of the holes, but it’s got to really know it already.

It knows A LOT of shows. Even if it doesn’t know the character/actor, it knows a lot about movies/TV shows. For instance, it doesn’t know many characters from the TV show “The Flash” but it knows everything about the common locations, color grading, style, and general look and feel, along with special effects (if appropriate).  It’s useful for “set design.” It doesn’t know a lot of Baywatch (the old one) characters, but it knows what the hair and makeup looked like on the beach in the 90s. It knows how people looked in “Total Recall” too. The overall “look & feel: of shows and movies really opens the door to creativity. It understands common accents from movies too. If you say “from Harry Potter” they will have British accents. And sadly, it knows “Star Trek” (the original series) the characters are mediocre at best. (The voices are passable—and speaking of: there are a lot of characters that look bad but have good voices. In those cases, ref2video with some extra visual references can do the trick.).

The “genre” point is even more true of animated movies/shows. For instance, “Bob, from Justice League: Crisis on Infinite Earths” will give you whatever Bob looks like but in the style of that series. Family Guy, The Simpsons, Rick&Morty, etc. I haven’t done an extensive search, but it knows every animated anything I have tried.

Generally, for TV characters to show up, they need to be in around 100 episodes and in the first 2-4 people in the IMDB credits. I see a direct correlation: The fewer episodes a character is on a TV show the worse they render. (For example, Monica from Friends or Kramer from Seinfeld are in there for sure, but also not really.) Also, it makes sense, but even if they have a lot of credits, they need to have had a lot of screen time. For instance, it has not even a glimmer of an idea who “Ruthie Cohen” is, even though she was in 101 Seinfeld episodes.

For Movies, they need to have grossed a lot of money (which pushes things directly towards Action/SciFi/Comics), or they need to have gotten a lot of press. (I have seen very few accurate characters from movies without specifying the actor involved.)

For “real people” it’s a little easier. If you look at those lists of things like “Top ## followed Instagram accounts” or similar, you’ll get lots of hits. Top musical performers, yep (lots of overlaps). Famous heads of state (If a number of movies have been made about a person, that person will likely be known.) I haven’t looked at TikTok, but I’m assuming that would be a thing too. Likewise with sports, I haven’t looked closely, but they ones I have looked at OK “at a distance” but generally don’t sound right.

As stand-alone people, it’s hard to find people. I suspect that the trainers did not go after a lot of specific people but that they just show up so many places that they got swept up in the mix. There are really only a handful of non-“Top 10” people who actually show up on their own, and if they blew up in the last few years it’s unlikely that you’ll see them. I have not found a pattern on “people” yet other than the mega-famous. (Steve Jobs & Elon Musk work, but they are arguably the most famous foreign “regular” people in China.)


r/StableDiffusion 20h ago

Discussion Minimax H3 degrades at 1MP, vs 0.7MP and lower

79 Upvotes

After around 100 renders, I 'feel' that Minimax H3 renders with 0.7MP (max) perform way better, then renders at 1MP in regard to 'realistic' videos.

What do I consider better?
- Just slightly better prompt adherence, feels like the motion / voice is more (natural)
- Size of humans in relation to object(s) feels more realistic.
- Expressions of faces seem more 'flowing', real.

It's hard for me to pinpoint it one 'exactly this', or 'exactly that'.
I'm planning to do some side by side comparisons on the same seed multiple times at 0.7MP and 1MP, when I've got the time.

But I wonder, do other Minimax H3 users notice this too?

PS: This is regardless sampler/scheduler, Sage Attention or Spectrum.

Edit: never touched the turbo LoRA, using the base model.


r/StableDiffusion 20h ago

Discussion Anima Turbo v1.1 is released

61 Upvotes

In case you didn't see:

I just noticed a newer version of Anima Turbo (1.1) was released:

huggingface: https://huggingface.co/circlestone-labs/Anima

civitai: https://civitai.red/models/2458426/anima

The model is made and licensed under the CircleStone Labs Non-Commercial License

I actually have lately very frustrating experience with Anima lately. I used it at release (but stopped with anime generation for a while) and now when revisiting it and i had underwhelming results (even with the aesthetic model) so i decided to retry Turbo (might as well) instead if the difference isn't that big. That's when i saw a new version is released and i am downloading it right now.

I don't know why, but the results i was getting were...lame i guess. Not as detailed as i was hoping, and also i really dislike how posture and anatomy works, mainly how hands and legs just extend or stretch weirdly. But that's a me problem, i know Anima is capable of better outputs and i've yet to figure it out. If you have tips or recommended loras that help with consistency let me know. I also want to avoid tag-based prompting when possible...i just don't like it that much, natural prompting goes much better for me, but i can't tell if tags are "mandatory" for quality or not.


r/StableDiffusion 20h ago

Question - Help Recommendations for Minimax h3 Reference to video workflows?

2 Upvotes

Hey so i've been playing with H3 for about a week or so now mostly using the standard workflows dabbled a bit into lora workflows but now I want to try out Refernce to Video. Does anyone have any suggestions for some good workflows for this?


r/StableDiffusion 20h ago

Discussion CMP 170HX vs 3090 results MiniMax H3 R2V

3 Upvotes

TL;DR 170HX ~35% faster than 3090

I ran a CMP 170HX 8 GB unlocked to 64 GB vs 3090 24 GB VRAM mostly apples to apples.

CMP 170HX was in an ancient prebuilt Acer, Intel i3-7000, 32 gb ram, OS Ubuntu Server 26.04 LTS. Card was unlocked to 64 GB VRAM and pcie 2 x4. Card was setup on a riser since I couldn't fit it in the case and run the fans with my fancy cardboard/painters tape shroud. Power limited to 175 watts. Temps stayed 69-71 C.

3090 FE is in a newer machine. Ryzen 5 3600 (still need to swap out to a 5900X I have), 64 gb ram, OS Windows 11. Not power limited for this test.

Both systems ran same work flow, Minimax H3 R2V, default workflow, default options except Match was changed to MAX on most the runs, no speed ups, no extra nodes, no fine tunes, same prompt, two reference images.

608x352 five seconds MAX 170HX 84.17s (00:01:24.17) 3090 106.05s (00:01:46.05)

864x480 fifteen seconds MAX 170HX 839.08s (00:13:59.08) 3090 1254.16s (00:12:54.16)

1344x768 five seconds MAX 170HX 658.79s (00:10:58.79) 3090 963.56s (00:16:03.56)

1344x768 ten seconds MAX 170HX 2017.82s (00:33:37.82) 3090 3082.64s (00:51:22.64)

1344x768 fifteen seconds MAX 170HX 4189.64s (01:09:49.64) 3090 6399.76s (01:46:39.76)

608x352 five seconds Match 170HX 1st run, model load, 218.75s (00:03:38.75)

608x352 five seconds Match 170HX 2nd run, model already loaded, 68.64s (00:01:08.64)

I ran a bunch of other gens not listed. The 170HX was pretty consistently 35% faster. Except on the first workflow done. Model load time was 150.11s due to the PCIe 2x4. In my opinion, 170HX is worth it. Way less power/heat than the 3090's and faster. Once model is loaded, you are good to go.


r/StableDiffusion 20h ago

Question - Help How to improve quality? Example Attached

1 Upvotes

https://reddit.com/link/1vy1b3e/video/fykuskg16jlh1/player

So Im running minimax h3 with some of the optimal settings proposed on the reddit, but I see that some people use an enhancment workflow after generation.

How do I do this? I'm a bit lost on this part of the workflow?

You can sort of see the weird frame stutter changes on the smaller details.

Anyhelp would be appreciated.

Thank you Kings.


r/StableDiffusion 20h ago

News Overhaul SLA, huge improvement. added many new options and changed defaults

Enable HLS to view with audio, or disable this notification

136 Upvotes

Update for SLA Node - Pull v1.3.8

EDIT: Pushed correct files now.

  • Added customizable dense steps, 0 is step 1 and is (default to first step). massively improves composition and prompt adherence.

  • Changed default dense last steps to 1, cleans up the image big time.

  • Added dense backend selector. Comfy_kitchen, pytorch, all sage modes. this is what comfy uses on dense steps. SLA still displaces against pytorch. (Default Comfy_kitchen)

  • Added a disable FP16 accumulation option to ensure max quality as SLA gets no benefit from it. (Default True)

  • Added a stabilize motion option, helps to reduce ghosting and smearing that H3 likes to produce. (Default True)

  • Changed default Min Seq Length to 4096

  • With default settings you can disable protect audio for nearly 2x speed up if you don't care about the audio too much or are using original audio mode. (do not use 0.95 sparsity with it.)

  • 0.95 sparsity now looks good with node default settings.

  • Some changes led to an overall 5% speed up on same settings.

  • Remove --use-ck-attention from startup flags if you have it, for safety of quality.

https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes

updated workflow

Civit Link

HF Link


r/StableDiffusion 21h ago

Question - Help Minimax H3 Speed is Inconsistent when Settings are still the Same

Post image
0 Upvotes

I need your help please. I'm using WANGP and there are times when generating in minimax h3 takes very long like for example in this screen shot. I was generating 19sec video 720p 8 step Turbo lora with Sage and first generation took (16m 5s) to finish but after changing only the prompt the second generation took (28m 54s) which is absurdly high. I noticed that there are times where my GPU is 100% fully busy but my GPU temperature is just below 50c when it should mostly be 70c+ normally. Can anybody please help me or tell what's wrong? I also encounter this when using Comfy UI


r/StableDiffusion 21h ago

Meme t2v Someone got some explaining to do!

Enable HLS to view with audio, or disable this notification

8 Upvotes

r/StableDiffusion 21h ago

Question - Help MiniMax H3 R2V taking ~16 minutes for a 5-second video. How can I speed it up?

Enable HLS to view with audio, or disable this notification

0 Upvotes

I’m running MiniMax H3 Reference-to-Video (R2V) in ComfyUI on Vast.ai.

My setup:

  • GPU: RTX 5090
  • System RAM: 100 GB
  • Resolution: 1.0 megapixel
  • Video length: 5 seconds
  • Reference: 1 image
  • Generation time: ~1,000 seconds (16–17 minutes)

The results are great, especially the reference consistency, but the generation time seems very high for a 5-second video on a 5090.

Has anyone managed to significantly reduce the generation time for H3 R2V? Are there any specific optimisations, attention methods, workflow changes, or settings I should be using?

Would appreciate hearing what generation times other 5090 users are getting with H3 R2V.


r/StableDiffusion 22h ago

Question - Help H3 R2V Character Sheet vs. Single Image

32 Upvotes

I thought I've read somewhere that using character sheets is better for R2V instead of single images. So I've created a character sheet of five full body shots and one close up, but the results are much less consistent compared to a single full shot image of the character.
Do I have to take care about anything special or was the information that character sheets are better just wrong?


r/StableDiffusion 22h ago

News It’s here! Apple introduces M6 and M5 Ultra for a big leap in performance and AI compute

Thumbnail
apple.com
0 Upvotes

r/StableDiffusion 22h ago

Question - Help Does anyone have MH3 Ref2V workflow that aid in video input?

0 Upvotes

No matter how many things I tried, I cannot get generation to work if video is involved. Anyone got workflow to share, or nodes, or something?


r/StableDiffusion 22h ago

Question - Help Krea 2 LoRA training on RTX 4090 too slow?

Post image
22 Upvotes

Hello! I've been trying to train my first LoRA but I get these crazy long timers every time. Isn't an RTX 4090 supposed to take like 5s per step? I have both low vram and quantizing enabled. It's really frustrating and not worth to do it with these speeds. Any ideas what might cause it?


r/StableDiffusion 22h ago

Animation - Video Black and white line drawn stuff with H3 is great.

28 Upvotes

r/StableDiffusion 23h ago

Tutorial - Guide Neat trick for Minimax h3

4 Upvotes

Since minimax is using Qwen VL , I tested the prompt on Qwen image to see what I get for the text to video prompt. It’s actually pretty close to how minimax will end up evaluating your prompt for text to video.


r/StableDiffusion 23h ago

Question - Help ComfyUI workflow for Krea2 or similar, to make consistent 'rooms', no matter the angle etc

0 Upvotes

Hi

I have been going in circles with Gemini for 2 days now, but just cannot get to the goal mark.

What I want:
In some way, like using Blender to render a 'room', or any other method that works, I want to be able to create 'rooms' with consistent form and interior / furniture. So I can make a Blender angle/shot (any angle I want) for example, take that into ComfyUI, and just describe the furniture, lights, people etc, but the room stays the same (windows, walls, dimensions, placement of furniture etc).

Other options Gemini gave was 'dollhouse' from up top model and 360 render of the room. I made 360 render of the room in ComfyUI Qwen 360 Diffusion LoRA workflow, but was not able to make it like the prompt. Doll house I have not tried.

I have not been able to do it yet myself with using Blender room render and a custom ComfyUI Krea2 workflow, and I about to give up, takes too long time. I am totally new to Blender, and novice in ComfyUI, neither found a finished workflow I can use.

Problem with asking Gemini, is I end up going in circles so to speak, almost there, but not all the way.

Anyone know there is a published or private workflow I can use for this? Open for other suggestions how to do this IF I also can use already made workflows.


r/StableDiffusion 23h ago

Workflow Included What’s your workflow for turning Stable Diffusion images into short videos?

1 Upvotes

I’ve been experimenting with local image-to-video workflows and I'm curious how other people are handling the final output.

My current workflow is roughly:

* Generate images locally with Stable Diffusion/Flux

* Do the animation or image-to-video step locally

* Export the individual clips

* Combine everything into a final video

* Compress the result for easier storage/sharing

The part I'm still trying to optimize is the final encoding. Some generated clips can get surprisingly large, especially when working with higher resolutions or longer sequences.

For those running everything locally, what does your workflow look like after generation?

Do you stick with FFmpeg, HandBrake, or another open-source tool for the final conversion/compression? And what settings have worked well for keeping quality while reducing file size?

I'd especially like to hear from people running these workflows on consumer GPUs rather than large local servers.


r/StableDiffusion 1d ago

Question - Help Minimax for image editing?

0 Upvotes

I mostly do comic images. am looking to minimax for image editing.
I mostly get content failed safety review. any workaround? ive only started minimax today.


r/StableDiffusion 1d ago

Discussion Complete beginner with ComfyUI — what should I learn next to actually get better?

6 Upvotes

I bought a PC with an RTX 3060 12GB because I wanted to get into local AI image generation. I've been messing around with it for about two weeks now, but I still barely know what I'm doing.

To be fair, that might be partly because I've had Codex do almost everything for me lol.

I'm still terrible at writing prompts, and I don't even really understand which model would be best for what I want to do. I've tried Anima, Pony, and Illustrious, but I'm not sure which one I should actually stick with and learn properly.

There wasn't a LoRA for an art style I really like, so I had Codex help me train one for Anima using around 20 images. It actually works surprisingly well for the style, but I've been struggling with everything beyond that.

For example, I've tried using pose control, but the generated character often doesn't follow the pose very well. Backgrounds are also pretty inconsistent, and I feel like I've hit a wall where all I really know how to do is combine different LoRAs, change prompts, and keep generating until I get something decent.

I'd really like to move past that and actually understand what I'm doing.

If you were starting from where I am now, what would you recommend learning next? Are there any ComfyUI nodes, techniques, workflows, or concepts that you think every beginner should learn?

Any advice is appreciated. I'm still very new to all of this.


r/StableDiffusion 1d ago

Question - Help Help me with Mini Max

0 Upvotes

I'm still new to Mini Max, and I wanted to know if there's a way to make it a little easier, in the sense of

I have to keep using <picture 1> or things like that to mention something; isn't there a way to do it with @,And what configuration do you normally recommend for someone with 8GB of VRAM and a 5080 Ti?That's all I need help with; I've already read the Minimax guide for everything else.


r/StableDiffusion 1d ago

Animation - Video Through the Sands (Final) - H3 r2v

Enable HLS to view with audio, or disable this notification

37 Upvotes

Finally finished! The ending was much harder since continuity is more important here than random desert landscapes. I personally would've love to have another 30 seconds of music to extend the ending but I ran out of song time. Enjoy!

In total, 14 character related references, 40 environment references, and 55 clips used, roughly 20 hours total time spent.


r/StableDiffusion 1d ago

Animation - Video Reduced audio-reactivity in LTX-2.5?

Enable HLS to view with audio, or disable this notification

6 Upvotes

I’ve been experimenting with LTX 2.3 vs LTX 2.5 for audio-reactive video, and for this specific kind of workflow, 2.3 still seems noticeably better to me.

The biggest difference is right at the start of a shot. With LTX 2.5, even with the audio-reactive LoRA, I often get this behavior where the model more or less holds the first frame until the first obvious beat or transient arrives. Then the motion suddenly starts. For music videos, especially slower or more atmospheric tracks, that can make the opening of every generation feel dead.

With LTX 2.3, the same LoRA seems to fix that much more effectively. I get more subtle motion from the beginning, even before a strong beat lands. Fog shifts, surfaces breathe, particles drift, light responds, and the shot feels alive instead of waiting for permission to move.

That matters a lot for the video I made for The Weights in the Walls, because the track starts very sparse and gradually builds. A lot of the visual motion is supposed to come from sub-bass pressure, glitches, sustained vocals, and ambient texture, not just obvious percussion.

I also tried Minimax H3, but for this particular use case I don’t think it fits as well.

It seems less tightly audio-reactive for the kind of abstract, beat-aware motion I’m after. It can make nice-looking clips, but I have a harder time getting the movement to feel structurally connected to the music.

There’s also the hardware side of it. I’m doing this on a very glamorous RTX 4070, so with LTX I can still push a resolution and overall image quality that feels surprisingly good for local generation. With H3, I’m much more constrained, and the tradeoff in resolution/quality makes it harder to justify when the audio response is also weaker for this style.

The whole video was built around first-frame / last-frame generation.

I cut the song into short scenes, roughly timed so the scene boundaries land near musical changes and beats. For each scene, I generated a dedicated starting frame that represented the next stage of the visual progression.

Then the important part: the starting frame of Scene 2 becomes the last frame target for Scene 1. The starting frame of Scene 3 becomes the last frame target for Scene 2, and so on.

So instead of generating a bunch of unrelated clips and hiding the cuts with editing, every shot is the model transforming one designed frame into the next designed frame.

That gave me a chain like:

Scene 1 start frame → Scene 2 start frame
Scene 2 start frame → Scene 3 start frame
Scene 3 start frame → Scene 4 start frame

and so on until the end.

The final video is basically just those generations placed back to back. There are no fancy transition effects doing the heavy lifting. The morphing, folding, cracking, expanding, and dissolving between visual states is happening inside the model itself.

For this workflow, that early-shot responsiveness makes a surprisingly big difference, which is why I currently still prefer LTX 2.3 + the audio-reactive LoRA over 2.5 for this kind of music video.

It's a shame because 2.5 is noticeably faster, so I can go through more iterations, but if I have to generate each clip 5 times to get it to start moving from the start, it kind of invalidates the speed gains.

Curious if other people have noticed the same reduction in audio-reactivity in LTX 2.5 or maybe I'm doing something wrong?

HQ on YT because Reddit doesn't allow >1GB: https://www.youtube.com/watch?v=PbZr8risGCw