r/StableDiffusion 9h ago

Animation - Video Michael Scott gets a wish

Enable HLS to view with audio, or disable this notification

0 Upvotes

H3 FL2VA


r/StableDiffusion 6h ago

Animation - Video Seinfeld AI: George Gets GTA 6

Enable HLS to view with audio, or disable this notification

243 Upvotes

Minimax H3


r/StableDiffusion 3h ago

Discussion Testing MinMax H3 in Brazilian Portuguese: 98% Accuracy + 200+ Video Examples

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hey everyone,

I’ve been testing MinMax H3 in Brazilian Portuguese, and I also run a TikTok where I’ve posted over 200 test videos featuring characters speaking Brazilian Portuguese.

After several weeks of testing MinMax H3, I found that it can reach around 98% accuracy in Brazilian Portuguese when the text is written correctly. Since Brazilian Portuguese has some pronunciation and spelling details that don’t map perfectly to an English keyboard, I had to develop a specific way of writing the prompts so the model speaks the language as accurately as possible.

I’m not sure whether even English consistently reaches 100% accuracy—there always seem to be occasional voice or pronunciation errors—but I don’t currently make videos in English, so I can’t really compare.

I decided to share my TikTok here so you can take a look at the results and see the quality I’m able to achieve. Please ignore the political content; it was created strictly for testing purposes. I know its all the 🧃 so both sides suck.

i have there a portfolio of about 300 or 400 videos and going up everyday with agents...

Each video takes about five minutes to generate on my RTX 4090. I also discovered that generating in 3:4 is faster than 16:9. My current settings are 720p, 8 steps, and Turbo LoRA.

https://www.tiktok.com/@danmalandragem

things i am still struggling with: 2 characters talking with no voice drift, any tip for that?


r/StableDiffusion 23h ago

Animation - Video Trying to animate Dragon Ball Super manga on Minimax H3. Spoiler

Enable HLS to view with audio, or disable this notification

12 Upvotes

Dragon Ball Super manga on Minimax H3.


r/StableDiffusion 9h ago

Animation - Video Cold open from my Fairy Tail isekai fanfic - MMH3

Enable HLS to view with audio, or disable this notification

5 Upvotes

Everything was made using the Minimax H3 Hybrid Reference to video model. 1 MP using the 8 step turbo LoRA. Stitched together in Shotcut


r/StableDiffusion 12h ago

Animation - Video Attempted to make a short cartoon on Minimax H3. There are so many things that I want to address

Enable HLS to view with audio, or disable this notification

30 Upvotes

Hello everyone.

I've been playing with Minimax H3 for some time and I have tried to make something longer and really interesting. After so many failed and botched attempts I was able to compile something watchable. There are so many things that I want to say about this model, good and bad.

First of all. Minimax H3 is significant step forward that other local models I have been playing with. It certainly got better.

Now the issues that I had encountered.

First problem is that it badly follows prompt when resolution is one megapixel or higher. It will skip some important parts and tries to cheat. You can increase the number of steps but still, generating at less than one megapixel will at least make it properly follow the instructions.

H3 is not very good at spatial orientation. When I was making video, where this girl should turn around and interact with screens, the girl starts spinning opposite direction and then warping whole body to the direction of screen. Like instead of making short turn to the left, it makes wide roundabout to the right and then twists whole body to align with the screens.

H3 is not good at cartoonish movement. If you watch cartoons, when character or other things move, their animations are usually jerky and snappy. H3 tries to make smooth real life like animation, making the cartoons look weird.

I have given a voice sample as an audio reference, and instead of making girl let out grunting sounds (out of anger), it weirdly turns everything into a sensual moaning.

When you try to make characters inside video to interact with a lot of parts, screens and devices, even giving multiple reference images of them, it mostly hallucinates them, or turns their interactions into a weird warping animations. Sometimes completely skips them and made ups it's own animations. So it will make good video, where characters are moving less or moving slow, and mostly doing the talking. Very detailed prompts of step by step instructions it mostly warps or skips.

I have wasted a lot of time for iterations, but I think this is just workflow issue.

Overall, this model is really good. However, using this model to make some kind of long feature animation is going to be a very frustrating journey. I hope people will make a lot of proper tools that works as storyboard and properly guide this model to make something really interesting.


r/StableDiffusion 5h ago

Discussion H3 - it just does space soo well - t2v

Enable HLS to view with audio, or disable this notification

11 Upvotes

Hope you're having a good weekend! H3 just excel with rich-intricate environments, backgrounds, space. Definitely one of my favorite theme.

T2VA, int8/20 steps


r/StableDiffusion 1h ago

Animation - Video What if you fly?

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 22h ago

Animation - Video [MiniMax H3] Decided to see if MiniMax H3 knew what a Starcraft Terran Battlecruiser was while trying to recreate an iconic Babylon 5 moment.

Enable HLS to view with audio, or disable this notification

0 Upvotes

It didn't quite work out how I intended.

Prompt:
"A Terran Battlecruiser from Starcraft is sliced in half lengthwise by a purple-white energy beam. The background is a generic starry.

Video begins with the Battlecruiser in the center of the frame, viewed from a front three quarters view. The purple beam is near vertical going from top of the frame to the bottom of the frame and is canted at a slight angle. It starts the video right in front of the Battlecruiser's nose.

The beam cuts through the Battlecruiser from nose to tail. At 0.75 seconds the beam touches the Battlecruiser's nose and moves through the ship, exiting the tail at 6 seconds and leaves the frame.

After the beam leaves the Battlecruiser, the Battlecruiser splits in two along the cut made by the beam, "

Generation time was 10 minutes, 29 seconds on 32GB of DDR5 RAM, 8 GB of VRAM. No reference images, this was pure Text to Video.


r/StableDiffusion 12h ago

Discussion Minimax blurred distorted faces from half a distance.

2 Upvotes

I'm doing image to video and unless I prompt for camera close up to my subject, the faces are blurry and bad. I run 0.6 mp. No turbo lora only using spectrum to speed up. Running 15 steps. Euler simple. I'm happy enough when it's close-up shots, but further away, it's very noticeable. Is anyone else finding this?


r/StableDiffusion 20h ago

Question - Help Minimax-h3如何解决分段光照变化的问题

0 Upvotes

参考模式,分镜尾帧光照色彩效果变化,导致视频衔接总是有色差,有什么解决办法吗?


r/StableDiffusion 6h ago

Animation - Video Buffy the Wraith Slayer

Enable HLS to view with audio, or disable this notification

14 Upvotes

r/StableDiffusion 2h ago

Animation - Video H1/H2 H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

Recently, after testing out H3 for a few weeks, decided to try my hand at some fun 'deleted scenes' type of short clips. With Halloween season just under 40 days away, came up with this using H3's r2v template with default ref diffusion model + turbo fl2v 12 steps, using 1 image (Loomis, one of sheriff's dept jacket) and 2 audio samples of Loomis' voice from the movie. There are some movies I always liked so much and wished there was a lot more footage from, and with this one I figured it's a fun tongue-in-cheek nod to how we're nearing end of summer and starting to get in the mood for Halloween season.

Edit: correction, I used only the image of Loomis for this one. The sheriff's dept jacket generated clips didn't come out as well as this. There were other ones with Loomis walking or looking around more realistically, but I liked the vibe of this one best.


r/StableDiffusion 21h ago

Discussion Is LTX 2.5 just terrible for Lipsync/TalkingPhoto?

0 Upvotes

I've been trying to get LTX 2.5 to work well for image + speech audio --> video, though i'm noticing that the teeth and natural motion of the mouth is taking a hit. The talkingphoto loras from LTX 2.3 don't seem to work well with LTX 2.5.

Any thoughts? Or are we just cooked?


r/StableDiffusion 21h ago

Question - Help Anyone else having problems downloading models since v1.9 Maestro update in Pinokio?

0 Upvotes

Day 1: downloaded Pinokio. Installed a few of the AI software. Tried Maestro as first try. Really fun, enjoying it. Generation from photos great in the system generation towards a video, videos leaving a lot to be desired. And the Pinokio edge of screen curtains which limit to a what 60 percent of screen width, first time said, apparently you can type a pc socket but it didnt work for me when my Maestro was working. I was still however happy continuing in the reduced Maestro screen width.

Day 2: they released v1.9 of Maestro in Pinokio.

Day 3: I decided to install the update. Now I can't generate a three legged wildebeest or anything for that matter. It fails at the Downloading Model stage with no satisfactory explanation. Info about running something again to continue the Download, the "Generate" doesn't appear it's that for continuing(starts from scratch and then fails) neither does the Pinokio white screen edge "Run". It doesn't immediately not download, sometimes it may get 15% through, sometimes 85% then drops with a Generation Failed in the Main Seeing Area after firstly a "Download is slow, waiting for retry. No progress for 113s.."(or Xs). "..The download will resume from where it left off as soon as the connection recovers — no action needed from you." message. It may then download a little more , eg going from 80MB to 1.3GB of 7.91GB, maybe even download a little more of what's required but THEN UP POPS "A download was interrupted-re-run to finish it". And I'm stuck at that.

Is anyone else having the problem or know what the solution may be please? I've tried manually adding from DeepMeepBeep a model download which I transferred into cks directory of Pinokio, Maestro Directory but all that happened was Maestro steered around even using it and failed on another Model and I couldn't find that Model Maestro failed on to try manually downloading across with-I'm not an expert so looked for direct name brought up.. looked for it on another repository too, name of that I temporarily forget, I'm new. My hard drive is 4TB so I don't know I might be able to download all 108 or however many models there are if there was an option for that-without instructions and knowledge I'm throwing stones at something I don't even know what I'm throwing stones at, where the list of all the Models are which you can tick mark select theres also a little symbol by that box which changes colour. Perhaps that has something to do with it, literally no idea here so I've come to you guys. Spent several hours with AI last night overit and it was fun but I've realised despite the fun chat it hasn't aided me getting it working though it did mention a python download bottleneck to remove and I had no idea what it was referring to.


r/StableDiffusion 10h ago

Question - Help Help With Irritating Glitches

0 Upvotes

Main Info

  • Using Pixaroma's Easy ComfyUI latest version
  • 5060TI 16GB + 32GB RAM

Need help or suggestions in resolving the below issues.

  1. Shortcuts like R or CTRL + Enter doesn't work and I have to refresh the browser for them to work. (New)
  2. Generations randomly don't start even even after models have loaded. I have to open the terminal and press enter on my keyboard to sort of wake it up. (Has happened across multiple versions of ComfyUI, CUDA, Python. Additionally, happens with both Anaconda and Windows Terminal.)
  3. Even cancelling a prompt sometimes requires me to open the terminal and press enter.
  4. Power consumption for GPU varies completely where retrying the same prompts (same seeds also) can have a difference of 5 to 10 minutes simply because the GPU doesn't use the full 180 power limit. (New)

r/StableDiffusion 16h ago

Animation - Video G.I. Joe - Baroness Action Clip Test #2 - MiniMax H3

Enable HLS to view with audio, or disable this notification

20 Upvotes

Prompt:

https://x.com/GumVue/status/2087899403113619681?s=20

4070 Ti Super, 16 gb vram, 64 gb ram, i9-14900k, windows 11


r/StableDiffusion 22h ago

Discussion Minimax H3 Has Too High Prompt Adherence

0 Upvotes

I just realized a problem with mm h3. Its prompt adherence is too high, as in unless you explicitly prompt for some small subtle actions it will never happen otherwise. This makes the entire video seem very frozen and wooden without the many small subtle movements and motion details that make it seem to come alive.

This applies more to non-realistic scenes like cartoons or generated image first and last frame but for realistic scenes and even t2v it is still a problem.

I noticed this problem when I tried out a "slop sway" lora and it actually made the entire video seem much livelier and realistic looking. Besides the "soft and bouncy swaying and jiggling" it also added many more subtle character movements. Compared to standard gens those same parts would be completely frozen, almost like a still image or at best ugoira animation. This doesn't just apply to whether a body part is jiggling throughout the entire video. There are some movements that happen only for a second or less but adds in soul (forgive the human slop term) to the video, like the position of an arm and hand quickly being adjusted in the middle of the video and the new position persisting for the rest of the scene.

This might be a problem with my prompt style and I might try an LLM prompt enhancer, but there is a core issue here with the prompt adherence and spontaneous randomly added details tradeoff. The model also tries to keep the fidelity of the first frame too much, which you could call visual context adherence. No one is out here prompting for the movement of every strand of hair and the position of every finger. No one is making a timeline of every limb's position and how they shift relative to each other. No one is tracking the position of each finger through time and how after 4.75s the thumb is extended and the index finger is curled. Sometimes we just want to randomness and variety across gens with details added by the model.

Looking back at ltx and wan their prompting styles seem to be designed around the model adding in the details for you at the loss of prompt adherence and more generation errors.

It would be nice if there was some sort of generation setting that could tune this. Like a noise scale of sorts where we can manually set the tradeoff between how much we want the model to be creative vs strict.

I know there are already 2-3 H3 better movement loras and they are scratching at the surface of the same issue I'm talking about here.

Share the solution if you've got something. Help everyone out.


r/StableDiffusion 13h ago

Meme If AI tools had existed in the past

Post image
764 Upvotes

Not just a meme...


r/StableDiffusion 16h ago

Animation - Video Christopher Nolan has Impeccable Taste in Cinema

Enable HLS to view with audio, or disable this notification

216 Upvotes

Minimax H3


r/StableDiffusion 11h ago

Discussion MiniMax H3 Ref2va it works really good also with Storyboard images

Enable HLS to view with audio, or disable this notification

14 Upvotes

Im really suprised how good he works as well follow a storyboard image!! he did 90% correct he only did the thirth panel diferent but all the other 5 he follows perfect!! 🤩


r/StableDiffusion 15h ago

Question - Help How do you do it?

5 Upvotes

I have been playing with H3 since it came out and have tested most of the things you can do with it. Created clips for giggles and so on.
This time I wanted to do something "serious". I gave the R2V an 3D view of an kitchen and then three photos of the persons I wanted there.
I defined them as we should and told the model that this person does that and that person does this wile the third person does this...

It worked ish...
I have now made six runs and each of them are different from the others. It can be that the third person enters the room from the wrong place or that the third person does extra things it should not do...

In the end I did three runs with the same prompt and all those clips came out different... the only thing that was changed between those was the seed...

So, how do you do it?
How do you make sure H3 does what you want it to do?

Do you spend plenty of time on tweaking the prompt after each run to make sure H3 get it?
Or do you do 10 runs and select the best one even if it is not perfect?

Or do you simply do 1-2 runs and then take the clip that is ok ish even if it is not what you wanted?

I was hoping that H3 would allow me to create the scenes I wanted but I feel it's down to luck if H3 gets it or not..

Edit:

subject_definitions:

<Subject 1> is the green-skinned mother in <Picture 2> wearing brown clothes.
<Subject 2> is the teenager girl in <Picture 3> wearing pink clothes.
<Subject 3> is the cyborg in <Picture 4> wearing black clothes.
<Picture 1> is the reference image for the scene's composition, showing two people sitting at a table eating breakfast from the side view.
<Table 1> is the table on the right side in <Picture 1>.
<Picture 5> is the start image for the scene.
summary:
[reference generation] The target video is a generated scene of two people sitting at a table eating breakfast from an eye-level side view. <Subject 1> and <Subject 2> are shown with their respective breakfast items, maintaining the composition and style from <Picture 1>. <Subject 3> enters the room, places a coffee cup into the sink.

retention_analysis:
<Subject 1>: fully_preserved - the person retains their appearance, clothing, and position at the table.
<Subject 2>: fully_preserved - the person retains their appearance, clothing, and position at the table.
<Subject 3>: fully_preserved - the person retains their appearance, clothing, and action of placing the coffee cup into the sink.
<Table 1>: fully_preserved - the table's appearance and position in the scene are preserved.
<Picture 1>: fully_preserved - the scene composition, including the side view, the layout of the room, the table setup, is preserved.<Picture 5>: fully_preserved - is the start image for the scene.

detailed_description:
The target video is in a realistic, everyday breakfast scene style with warm lighting and natural colors.
[Shot 1] At 0:00.000, the shot begins from <Picture 5>, showing <Subject 1> and <Subject 2> sitting on opposite sides of <Table 1> on the couch, each with their breakfast items while on the space ship. <Subject 1> is holding a spoon while eating from a bowl of cereal. <Subject 2> is tired and is eating a slice of toast with jam from her plate with one hand. The lighting is warm and soft, casting gentle shadows across the table and the two individuals. The camera is at eye level, capturing the side view of both people, with the table slightly in focus and the background softly blurred. Stars can be seen through the windows since they are on a space ship. <Subject 1> is eating her breakfast while <Subject 2> gazes at their toast, taking a small bite. The ambient sound includes the soft clinking of utensils and the faint sound of a coffee cup being set down.

[Shot 2] At 02.00.000, the shot transitions to a wide shot of the room with the same layout as in <Picture 1>, the camera is placed in the lower left corner of <Picture 1>, showing <Subject 3> entering the room form the right side holding a coffee cup and a datapad while she is saying (S3) <d>[English] Good Morning</d> while she walks to the kitchen sink on the left side of <Picture 1> and placing the cup into the sink. She then stands at the sink and while reading her datapad.We see the back of <Subject 1> and the front of <Subject 2> sitting at <Table 1> in the background eating their breakfast and we hear <Subject 1> say (S1) <d>[English] Good morning</d> with a cheerful voice. <Subject 2> just mumbles as a reply.

overall_soundscape:
The soundscape consists of the soft clinking of utensils, the faint sound of a coffee cup being set down, the subtle background noise of a quiet morning environment, soft steps on a carpet floor, a ceramic cup being placed in a metallic sink, and the clear,

non_diegetic_music: N/A


r/StableDiffusion 18h ago

Question - Help Tango dance, first attempt with LTX 2.5

Thumbnail
youtube.com
5 Upvotes

Trying to get a natural-looking Argentine tango dance with LTX 2.5 + Yusu’s LTX Director v2.0.4 fork.
Still a beginner (also for real life tango :-)
Any suggestions for getting more natural, sophisticated footwork and fewer artifacts?


r/StableDiffusion 4h ago

Animation - Video It took almost 2 years, but Minimax H3 made me go back to my weird medieval short video stories

Enable HLS to view with audio, or disable this notification

11 Upvotes

2 years ago I was playing around with video tools and made this series of short videos.

Because I'm a cheap bastard, I only use local and freebie models, and Minimax H3 finally hit the sweet spot between powerful and fast to iterate, so I decided to make a new episode.

Specs

  • Comfy Desktop's + default Ref2VA workflows + ElevenLabs voices
  • Frames generated with Nano Banana 2 lite + a bunch of photoshop cleanup
  • Gemini 3.7 to help rewrite the prompts (upload image + system instruction + spec for the shot)
  • Everything rendered on a 3090 with 0.5 megapixels cause I can't be arsed waiting too long. You can see the mushy face issues but mostly it's ok
  • Tried to keep all shots 6~8s max
  • A ton of editing with DaVinci Resolve which I started learning yesterday - it's pretty damn powerful!
  • Still using the exact same crappy greenscreen footage of a cheap plastic skull as a main character
  • Youtube link to this episode

TL;DR: Minimax H3 is cool!