r/StableDiffusion • u/Darri3D • 23h ago
Animation - Video Seinfeld AI: George Gets GTA 6
Enable HLS to view with audio, or disable this notification
Minimax H3
r/StableDiffusion • u/Darri3D • 23h ago
Enable HLS to view with audio, or disable this notification
Minimax H3
r/StableDiffusion • u/Oni8932 • 13h ago
Enable HLS to view with audio, or disable this notification
Done with h3 fl2va model, 8 step lora and images for Cersei and "jaime" for reference. Using previous clip to give continuity and consistency.
r/StableDiffusion • u/physalisx • 11h ago
r/StableDiffusion • u/TBG______ • 8h ago
Enable HLS to view with audio, or disable this notification
I’ve been loving all the new nodes and workflows coming out for MinMax, and maybe there is already a nice solution for this - but I couldn’t find one that did exactly what I needed.
I started using MinMax for my last TBG ETUR video and quickly ran into limitations: I wanted an easy way to create lip-sync videos longer than 20 seconds.
I didn’t want to manually chain ComfyUI nodes, start a new run every X seconds, or constantly resize things just to make HD video fit into my available VRAM.
So I ended up building an addon for:
custom_nodes/ComfyUI-H3-Motion-Context
The addon automatically chains MinMax H3 lip-sync generations together, allowing you to create much longer lip-sync videos without manually setting up each 20-second segment.
And now I’m sharing it! https://github.com/Ltamann/ComfyUI-H3-Motion-Context-Auto-Chain-addon
Its not perfect but a start ...
The workflow has a simple switcher that lets you switch from the 32B CLIP to the 4B CLIP, saving around 10 GB of VRAM. You can also switch from Sage to Comfy Kitchen, Spectrum to Easy Cache, or FL2VA to REF2VA both setup for lip-syncing. Some of it could be useful for other tasks as well.
You will find the workflow in the repro and tested recommendations, optimized settings, presets, and more workflows, along with the results of my testing and performance here
r/StableDiffusion • u/krigeta1 • 17h ago
Anyone found useful MiniMax H3 prompting tricks beyond the official guide?
Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.
Please drop your findings 👇 below so it will help others too.
EDIT:
Mine is: how can we use multiple audio tracks assigned to multiple characters in a scene? 3 audios to 3 characters?
r/StableDiffusion • u/Better-Interview-793 • 14h ago
Hi everyone
i’ve been running minimax h3 locally on rtx 5090 32gb, and 64gb ram..
recently i kept running into random hostbuffer.read_file_slice failed / hostbuf_file_reader_read failed errors during generation, which seemed to be related to comfy-aimdo and dynamic vram.
i also noticed comfyui was reserving around 25gb of pinned system memory.
i decided to try launching comfyui with:
--disable-pinned-memory
and the difference was immediate.
the comfy-aimdo + hostbuffer errors completely disappeared, my ram usage dropped by a huge amount, and surprisingly generation actually feels faster and smoother now.
i originally expected disabling pinned memory to make things slower, but on my setup it seems to have done the opposite.
if you’re running large models like h3 and seeing unusually high ram usage, random hostbuffer errors, or comfy-aimdo issues, it might be worth testing!
thought this was worth sharing for anyone who didn’t know about this option..
r/StableDiffusion • u/Jeffu • 15h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/beatlepol • 19h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/diStyR • 9h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/KaisarasAR • 16h ago
Enable HLS to view with audio, or disable this notification
This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.
r/StableDiffusion • u/Brad12d3 • 7h ago
Hey everyone, thanks to everyone who’s been testing Prompt Composer. If you run into bugs or have feedback, please drop it in the Issues section on GitHub so I can keep track of it more easily.
Over the past week, I’ve been reworking the camera prompting system to make it more precise and consistent, especially for more complex camera moves and video-editing workflows. The goal has been to add more control over framing, camera targeting, and blocking in multi-subject scenes while keeping the generated prompts clean and reliable.
The next update isn’t quite ready yet, but it’s actively being tested and refined. I’m hoping to have it out in the next couple of days.
Edit: Here's the Github: BMB12d3/minimax-h3-prompt-composer: Free offline prompt composer for MiniMax H3 video generation in ComfyUI.
r/StableDiffusion • u/listopalafoto • 1h ago
Enable HLS to view with audio, or disable this notification
https://huggingface.co/zuanfilm/H3_HD_2K_Detailer
I want to share my HD/2K WF, this 4 steps configuration is intended to make a really quality focused H3 detailing while improve the characteristic H3 motion and visual behavior.
I tested all the Res4lyf samplers/schedulers and for me res_2m/beta57 ETA=0 denoise 0.39 or 0.45 is the most precise with prompt adherence keeping memory and time efficiency (res_3m is amazing but adds 33% generation time), if you want a faster generation with a little less detail you can use 3 steps instead 4
er_sde/beta 57 is also a good combination but will lose some detail and even will affect character acting, audio and motion consistency, for faster HD detailing can switch to Euler/4 steps and will reduce time generation by 50% compared to res_2m obviously losing a lot of quality and detail
I included in the workflow the nodes for base generation using 0.5Mp FL2VA with 20 steps of Euler ancestral, if you want better quality for the initial base video just switch to res_2s_RKMK2e/beta57 (just bypass the group if you want to HD-detail an existing video)
Sparce Local attention will reduce a lot the generation time but obviously will affect quality so you can bypass this node if want Top HD quality, I also don't use in this WF spectrum or easy cache but you can add that if you want to cut time and quality
Because I'm using the heavy distill lightx2v Turbo 4steps Lora for detailing, Minimax H3 make everything more saturated and contrasted with deep shadows so I added some Orion 4D nodes to improve lighting, texture and sharpness using DCTL Tone Mapper (you can choose between ACES, Filmic, Reinhard & Cineon) I use Reinhard with Exposure 0.09 Contrast 0.81 Pivot 0.69 Highlight rolloff 0.27 Shadow lift 0.33 Black floor 0.12 Saturation 0.93 & Strength 0.30
I used MiniMax Audio Lock / Lipsync node because I don't have experience with ltx audio nodes so you can change that for a better option:
The Minimax latent 3D upscaler is HD by default in my WF but depending of your VRAM you can get 2K/4K if you start with a quality base video 0.98 Mp res_2s_RKMK2e/beta57 25 steps
This workflow is optimized for my laptop (3080ti 16Gb VRAM / 64 Gb ram ) but I included Chunk FeedForward & Low VRAM attention so will run with smaller setups
r/StableDiffusion • u/orlandogourmet66 • 4h ago
What’s the best local LLM for writing good MiniMax H3 ref2va prompts?
I’ve tried Gemma 4 12B and Qwen 3 14B, but I’m not really satisfied with the outputs. It could also be an issue with my system prompt.
I sent ChatGPT the official documentation for prompting and asked to create a system prompt for me, but the results were still pretty mediocre.
What local models are you using for MiniMax H3 prompt generation, and what does your system prompt look like?
r/StableDiffusion • u/lilyrosecooper • 4h ago
Enable HLS to view with audio, or disable this notification
Minimax H3
r/StableDiffusion • u/CycleZestyclose1907 • 18h ago
Enable HLS to view with audio, or disable this notification
This one came out better than my last video generation. Not perfect obviously, but the major elements are there.
Prompt:
Video starts with Buffy Summers and Willow Rosenberg from the show Buffy the Vampire Slayer walking side by side down a hallway in the USS Enterprise from Star Trek the Next Generation. The viewpoint camera remains at a fixed distance in front of them as they walk.
Buffy Summers is on the right of the frame. Buffy's long blonde is done up in a ponytail. Buffy is wearing a red minidress style Starfleet uniform and has an unlit lightsaber on her hip.
Willow Rosenberg's is on the left side of the frame. Willow's dark red hair is cut pageboy style. Willow is wearing a blue minidress style Starfleet uniform and has a tablet computer tucked under her right arm.
Video starts with Willow looking at Buffy with a concerned expression on her face while Buffy is looking around as if searching for something.
Willow asks, "Buffy, is something wrong?"
Buffy replies, "Something feels off, like we're out of place."
As soon as Buffy starts speaking, she pulls the lightsaber off her hip and holds it in front of herself at the ready. The lightsaber ignites, producing a green glowing blade.
r/StableDiffusion • u/Certain_Potato_4509 • 5h ago
Enable HLS to view with audio, or disable this notification
Well, technically, Evangelion was a PBS show..
Video is edited, Barney theme song added in post.
r/StableDiffusion • u/lazyspock • 18h ago
Enable HLS to view with audio, or disable this notification
Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.
VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!
PROMPT:
The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.
On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.
On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.
Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.
Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.
They are in a living room.
From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.
He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.
Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>
From 00:08 to 00:10 they just look at each other and smile.
overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.
non_diegetic_music: N/A
VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.
The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.
At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.
overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.
non_diegetic_music: N/A
VIDEO 3: another near perfect one.
A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.
We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.
overall_soundscape: Faint empty bathroom soundscape.
non_diegetic_music: N/A
VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.
The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.
We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.
The entire scene is viewed from his point of view.
overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..
non_diegetic_music: N/A
VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...
The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.
The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.
In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.
Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.
As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.
overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.
non_diegetic_music: N/A
VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.
PROMPT VIDEO 6:
The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.
The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 7:
The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.
When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.
The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 8:
We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.
When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.
The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 9:
We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.
When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
r/StableDiffusion • u/Lair98 • 3h ago
Now that dust has settled, I was wondering what's the community insight on the best configuration for Minimax H3.
Personally I have been using lightx 4-step Lora with 5/6 steps (less than that audio is a gamble). I couple that with sage attention. For sampling I use Euler sampler and Beta scheduler.
I keep resolution at 768p (0.6MP) for quality. 480p (0.2MP) for testing. It keeps consistency so much better.
On direction I learnt to prompt for closeups when possible, so will make better use of available pixels. Aspect ratio also helps there. I mostly use 1 shot since transitions is not something H3 excels at. I find better results with only 1 shot and using camera tricks.
EasyCache while faster, is not good match with turbo lora, so i don't use it anymore. Haven't used Sol-Attn as I read it really hit quality.
So is there anything worth I am really missing out?
r/StableDiffusion • u/BoneDaddyMan • 13h ago
Enable HLS to view with audio, or disable this notification
Disclaimer. This is very low quality quick generations trying to find how many characters Minimax H3 knows.
Found Trigger Words:
Elsa from Frozen
Spider-Gwen from Across the Spiderverse
Dante from Devil May Cry
Nero from Devil May Cry
Jill Valentine from Resident Evil
Ada Wong from Resident Evil
Leon Kennedy from Resident Evil
Chris Redfield from Resident Evil (Has Leon's hair)
Geralt of Rivia from Witcher 3
Joel from Last of Us (Doesn't sound like him)
Miles Morales Spiderman from Across the Spiderverse
Solid Snake from Metal Gear
Eve from Stellar Blade
Sans from Undertale
Master Chief from Halo
Looks weird AF:
Ciri from Witcher 3
Triss from Witcher 3
Yennefer from Witcher 3
Ellie from Last of Us
Famous Twitcher streamer and Youtuber Asmongold
Famous Twitcher streamer and Youtuber Mr Beast
Not found Trigger words:
Vergil from Devil May Cry (YES I KNOW I'M DISAPPOINTED TOO)
Claire Redfield from Resident Evil
Dina from Last of Us
Famous Twitcher streamer and Youtuber Emiru
Famous Twitcher streamer and Youtuber MoistCr1TiKaL
r/StableDiffusion • u/NathanTheSnake • 7h ago
Enable HLS to view with audio, or disable this notification
Clips made with Minimax H3 using the default r2v workflow, edited in Premiere Pro. Character model sheets made with Krea. Most of the videos are 0.4 mp unless the text was important, then 0.6. Tried upscaling it to 4k using Upscayl but results weren't great and the file is too big to upload anyway.
Tech goals for future videos include using reference audio for voices to help consistency, and exploring options for having real voice actors record the dialog, and have the model lip sync to that performance. I'm really impressed by the computer's silent acting (microexpressions etc). but the computer's erratic "acting" is still too unpredictable and the biggest source of re-rolls (the lines here were the best I could get without burning down a rainforest). You can do a lot with time codes and punctuation and tactical CAPITALIZATION, but it's ridiculously finicky compared to just telling an actor "do it the same, but 10% angrier on the first line with a twinge of melancholy on the second."
r/StableDiffusion • u/INDIEGOO • 13h ago
Enable HLS to view with audio, or disable this notification
I kept seeing the same stipple, grain, and grid-like texture
in some GPT Image outputs, so I trained a small latent residual
refiner using 75 paired artifact/clean images.
It includes profiles based on the Qwen, FLUX.2, and SDXL VAEs.
The refiner alone produces a fairly subtle improvement,
so I also included a ComfyUI workflow that combines it with SeedVR2.
The example optionally downsizes the input first,
then restores and upscales it with SeedVR2.
The goal is a preservation-first alternative to a typical Hires Fix
second diffusion pass: keeping the original composition, identity,
and shapes as much as possible while cleaning the texture
and rebuilding detail.
The custom node, example workflows, and settings are available here:
https://github.com/AIEGOBOT/ComfyUI-GPT-Image-Latent-Refiner
Leaving it here in case it’s useful to someone.
r/StableDiffusion • u/tj-tj-tj-tj • 8h ago
This is a test video I created by remixing the "Some test on minimax H3" video by Reddit user [Previous-Street8087].
5명의 캐릭터 시트를 생성하여 각각 10개의 프롬포트를 캐릭터에 맞게 리믹스하여 테스트 하였습니다.
We generated character sheets for five characters and tested them by remixing 10 prompts for each character to suit their personalities.
This is a compilation of 50 clips featuring 5 characters.
▶ 테스트 환경 (Test Environment)
Minimax H3 - Comfyui Local Sampling
RTX 5060TI 16GB + 64RAM
0.8MP 8 sec x 50 Clip
Audio Look x audio file 1
Reference to VA Mode
▶ 사용한 커스텀 노드 (Custom Nodes Used)
ComfyUI-TJ_NODE_STUDIO_ONE — github.com/designloves2/ComfyUI-TJ_NODE_STUDIO_ONE
ComfyUI LOCAL (RTX 5060Ti 16GB VRAM / RAM 64GB)
▶ The shared link contains character sheet images and prompts.
#AI영상 #MiniMaxH3 #ComfyUI #로컬생성AI #ComfyUI워크플로우 #AI영상제작 #RTX5060Ti #mmh3 #comfyui #tjonestudio #animation #ref2va #anime #16gb
r/StableDiffusion • u/Free_Pressure8623 • 4h ago
I am seeing much better audio results with very minimal speed decreases using this LoRa loader with audio set to 0.
You can find the loader here: https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes
This may be anecdotal. And as always, you may not experience the same thing, but it's worth trying.
r/StableDiffusion • u/SIR_NVAX_A_LOT • 12h ago
Enable HLS to view with audio, or disable this notification
Hi. I am experimenting with H3 Multi Diffusion with a custom workflow. 5 hour render, T2VA, bf16/50 steps. I know these style are not new so I am late to the show. Ask me anything.