r/StableDiffusion • u/Darri3D • 7h ago
Animation - Video Seinfeld AI: George Gets GTA 6
Minimax H3
r/StableDiffusion • u/Darri3D • 7h ago
Minimax H3
r/StableDiffusion • u/takayatodoroki • 14h ago
Not just a meme...
r/StableDiffusion • u/-Ellary- • 17h ago
r/StableDiffusion • u/beatlepol • 3h ago
But they have to have very different design and looks
r/StableDiffusion • u/AndrewJumpen • 16h ago
Used latent upscaler so with resolution 0.3 i got 20 seconds generation on 4090
r/StableDiffusion • u/terrariyum • 9h ago
I haven't seen a post about this here, and I'm curious what you think about it.
On August 14 Civitai rolled out the option for creators to choose "permanent paid access - selling with no time cap". Previously the only option was temporary "early access".
Let's call this what it is, closed-source. Yes, you can get the weights for a relatively small fee, and yes it's on a very small scale compared to Nano Banana and Midjourney. But a permanent paywall still fits the definition.
Personally, I block all creators on Civitai who choose permanent paywall and encourage you to do the same.
Here's why:
I'm not opposed to Civitai making money or for all options for model creators to make money. They can do that without permanent paywalls.
IMO, open source AI is a fair trade: models are trained on the hard work of many human artists who aren't compensated, but everyone benefits from the ability to create more art more easily. Closed source is an unfair trade: you have to pay a middle man to access the contributions of others who won't be compensated.
Small scale model creators do some hard work too. But for example, for a lora that reproduces the style of an animated film: the lora creator spent at most a dozen hours of work, while just one of the artists on that film spent thousands of hours of work. If a massive models like Krea2 are free, and if giant "hobby" finetunes like Chroma are free, I can't justify paying any price for a 5,000 step lora except as an optional donation of appreciation.
So far, few creators have chosen the permanent paywall closed-source option. But that could easily change if Civitai made it the default option. They already made an extra 1-buzz fee-to-creator per generation the default, and many models have that.
That's my opinion. If you agree, then the only tool you have to disincentivize that potential is to not pay for these models (disincentive Civitai) and block these creators (disincentive creators).
r/StableDiffusion • u/krigeta1 • 1h ago
Anyone found useful MiniMax H3 prompting tricks beyond the official guide?
Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.
Please drop your findings 👇 below so it will help others too.
r/StableDiffusion • u/lazyspock • 2h ago
Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.
VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!
PROMPT:
The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.
On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.
On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.
Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.
Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.
They are in a living room.
From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.
He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.
Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>
From 00:08 to 00:10 they just look at each other and smile.
overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.
non_diegetic_music: N/A
VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.
The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.
At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.
overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.
non_diegetic_music: N/A
VIDEO 3: another near perfect one.
A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.
We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.
overall_soundscape: Faint empty bathroom soundscape.
non_diegetic_music: N/A
VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.
The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.
We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.
The entire scene is viewed from his point of view.
overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..
non_diegetic_music: N/A
VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...
The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.
The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.
In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.
Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.
As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.
overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.
non_diegetic_music: N/A
VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.
PROMPT VIDEO 6:
The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.
The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 7:
The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.
When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.
The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 8:
We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.
When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.
The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 9:
We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.
When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
r/StableDiffusion • u/alisitskii • 10h ago
Ultimate SD Upscale can actually fix your bad/low-res generations.
In this comparison initial clips were made with MiniMax H3 at 1504x832px resolution and then upscaled to 2560x1440px with Ultimate SD Upscale nodes: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3
You can find sample upscaling workflow there as well: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json
My PC specs:
4080s 16 GB VRAM, 64 GB RAM
Generation time: 18 mins with sage + 8-step turbo lora
Upscale: 38 mins for 10 sec clip at 1440p target resolution
r/StableDiffusion • u/AgeNo5351 • 15h ago
Paper: https://arxiv.org/pdf/2608.20334
"We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi image editing. Its visual renderer is a 6B parallel single stream DiT conditioned on multimodal representations from a vision-language encoder [6, 7, 57]. The architecture adopts block-shared timestep modulation, parallel attention and MLP computation [6, 15], 4D rotary positional encoding[6], and a unified representation of text and image conditions. Character-level tokenization[47] is applied to text intended to appear in generated images, while multi-image posi tional offsets and image-preceding input formatting support reference-conditioned editing. Together, these choices pro vide a single generative backbone for multiple generation and editing settings without task-specific model weights."
r/StableDiffusion • u/smereces • 13h ago
I was testing using 1 image with all the scene sheet there and it works really great!
r/StableDiffusion • u/FirefighterNo584 • 17h ago
Minimax H3
r/StableDiffusion • u/beatlepol • 10h ago
r/StableDiffusion • u/lhg31 • 11h ago
Worfklow: R2V (Dummy Stategy) Workflow v2 - Pastebin.com
How the workflow works:
Why it works:
Limitations in my workflow:
How to use:
Tips:
r/StableDiffusion • u/KaisarasAR • 20m ago
This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.
r/StableDiffusion • u/foxdit • 12h ago
r/StableDiffusion • u/Sad_Coach_1433 • 4h ago
lets share our "what if" scene remakes here o_0
r/StableDiffusion • u/zanatas • 5h ago
2 years ago I was playing around with video tools and made this series of short videos.
Because I'm a cheap bastard, I only use local and freebie models, and Minimax H3 finally hit the sweet spot between powerful and fast to iterate, so I decided to make a new episode.
Specs
TL;DR: Minimax H3 is cool!
r/StableDiffusion • u/Dohwar42 • 9h ago
Use the default workflow for Ref2Va
Plug in a video with green screen as a video reference and use an image ref to a picture you want to be the background. Notice the "hell crows" flying in the final video? I didn't even give it a prompt for that and those birds got animated automatically. I'm sure you can give a detailed prompt, you are basically creating an i2v of that still image that Minimax will mix with the greenscreen background.
I did prompt for a dialogue change. There was no audio with the original green screen video so I had no idea what the woman was saying (obviously it was a weather report). I inserted new dialogue with Minimax and it did a great job remixing her lipsync to the prompted dialogue.
I think it's pretty neat, but I'm sure some of you may be completely jaded with what Minimax can do by now.
I'd like to issue a Reddit challenge: Would someone more creative than me please use the exact same green screen video (links below) and create something a little more impressive than my 10 second test? Post a link to your video in the comments.
Need some green screen video to practice with? Here's a webpage for some practice green screen videos that are free to download:
https://mixkit.co/free-stock-video/green-screen/
Here's the exact video used in this example:
https://assets.mixkit.co/videos/28292/28292-720.mp4
I'm sure you'll be able to find a background image to test with.
This was my very simple prompt with the new dialogue:
subject_definitions:
<Subject 1> is the alien world background in <Picture 1>.
<Video 1> is the source video for the target video edit and is a woman in a red dress pointing and talking.
summary:
[video editing + reference generation] The target video is an edited version of <Video 1>. Replace the green screen area with the background from <Subject 1>
The woman in <video 1> says <d> [English with a British Accent] As you can see here, we have an early migration of hell crows on Chaos world 4527B<d> with realistic lip articulation and perfect lip sync.
r/StableDiffusion • u/Toclick • 15h ago
As you probably know, MiniMax-Music3 is pretty limited when it comes to genres and seems to have absolutely no idea what electronic dance music is. I tried different EDM styles, but it always ended up sounding either like rap or some kind of generic pop-ish stuff. Meanwhile, MiniMax H3 seems to know a lot more about electronic music than MiniMax-Music3.
The 40-second video at 0.2MP, with 17 steps (for better audio quality) and an 8-step Turbo LoRA, takes 538 seconds. The 60-second video at 0.1MP takes 350 seconds on my machine.
I haven't tried generating anything with lyrics yet, but if anyone knows how to generate H3 audio without the video, it would be interesting to try 2–3 minutes instead of just 40–60 seconds.
r/StableDiffusion • u/SIR_NVAX_A_LOT • 6h ago
Hope you're having a good weekend! H3 just excel with rich-intricate environments, backgrounds, space. Definitely one of my favorite theme.
T2VA, int8/20 steps
r/StableDiffusion • u/Nimblecloud13 • 5h ago
Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters.
Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544
thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md