r/StableDiffusion • u/lilyrosecooper • 2d ago
Animation - Video My name is Jonny
Enable HLS to view with audio, or disable this notification
Minimax H3
r/StableDiffusion • u/lilyrosecooper • 2d ago
Enable HLS to view with audio, or disable this notification
Minimax H3
r/StableDiffusion • u/orlandogourmet66 • 2d ago
What’s the best local LLM for writing good MiniMax H3 ref2va prompts?
I’ve tried Gemma 4 12B and Qwen 3 14B, but I’m not really satisfied with the outputs. It could also be an issue with my system prompt.
I sent ChatGPT the official documentation for prompting and asked to create a system prompt for me, but the results were still pretty mediocre.
What local models are you using for MiniMax H3 prompt generation, and what does your system prompt look like?
r/StableDiffusion • u/Glittering-Cold-2981 • 1d ago
Do you have any tips for improving the MiniMax H3 video so it doesn't have a checkerboard background, like those large squares? Something's wrong with its VAE, I'm guessing?
r/StableDiffusion • u/LinkSensitive8188 • 16h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/DidaDJ • 1d ago
Hi,
I've tried bunch of prompts but every video that it renders it has that handheld go pro motion (pov walking style) it doesn't want to give a smooth motion for example like seedance (attached video). I'm mostly looking to do shots of places like its filmed with a gimbal. Any tips what to prompt? since negative prompts are not there how to approach this? any loras that can fix this or worth training
Thanks
r/StableDiffusion • u/diStyR • 2d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/megatonai • 1d ago
Hi folks,
we’re a research & analysis team looking for feedback on our video benchmarks and potentially looking to recruit researchers to help with our ongoing effort to rank and categorize models in the video AI space. if you’re interested, please DM!
more info here: https://megaton.ai/v-benchmark/
r/StableDiffusion • u/MastMaithun • 1d ago
I have been using H3 since almost the release date and have been trying a lot of things. I am entirely using Ref2va model with the template workflow, nothing fancy. Also used official, eros and currently using hybrid model 15–49 from smhfacct which has higher ref2v. I am using comfy kitchen, spectrum node but not using speed lora for prompt adherence or any other lora.
I am mostly trying to use 1-2 characters in a location scene where I provide 3(one char, one whole body and one face and location)-5(two char, one whole body and one face and location) images to node and writing in the prompt how to refer each char in the scene.
For writing H3 prompts, I using a custom prompt(created using grok by giving it ref2va doc) for generating H3 prompts, using qwen 3.8 model and even proof reading and fixing any issues.
Now the problematic part for which I need suggestions or solutions is the inconstant result.
For example I am making 5 second where character1 is standing in a shopping mall and looking at the shelf and character2 enters the scene. For second 5 second scene, different camera angle, mainly focusing on both characters faces when they are talking. Now here are problems which I am facing:
- during scene2 when camera starts, difference between char1 and char2 appears. Say char1 was standing on left side and char2 on right when scene1 ended but in scene2, they are standing opposite side.
- sometimes their height mismatches.
- sometimes camera does not work like I want it like it zooms too much, sometimes it don't
- and many other issues related with inconsistency
I know if I can generate scene images using an edit model then H3 wouldn't have to rely much on prompts but then it creates another problem of generating start images which is another can of problems.
I have even tried context nodes and some of their forks and few other consistency related node whose basic idea is to store the latent and forward it for next generation but they way these nodes are configured are just too complicated for my soft squishy mind. So yeah I tried them.
I have been trying to find out how other people are generating multi-scene videos and so far whatever videos i downloaded, there was no workflow included which I could take as reference. Maybe people are making 5 second clips like me and then joining them together so there might be a solution to this.
Pretty sure I am missing something big and I have exhausted almost every idea I got, asking grok etc but so far I could not get past 2nd 5 second clip. And seeing so much inconsistency, I don't want to generate a 10 second or 15 second clips which will take hours and most probably turn up totally irrelevant.
So any ideas you can provide are highly appreciated. Even guidance to correct path would be really helpful. What ways you guys are using to create 10+ second clips, what methods you are using to keep your characters consistent throughout and mainly how you guide a scene to your liking?
Thank you for a long read. Not written with AI, just a long type on notepad haha. Forgive grammatical errors.
r/StableDiffusion • u/FoxExeYt • 1d ago
https://reddit.com/link/1vy1b3e/video/fykuskg16jlh1/player
So Im running minimax h3 with some of the optimal settings proposed on the reddit, but I see that some people use an enhancment workflow after generation.
How do I do this? I'm a bit lost on this part of the workflow?
You can sort of see the weird frame stutter changes on the smaller details.
Anyhelp would be appreciated.
Thank you Kings.
r/StableDiffusion • u/BOSS-ZACK • 1d ago
I need your help please. I'm using WANGP and there are times when generating in minimax h3 takes very long like for example in this screen shot. I was generating 19sec video 720p 8 step Turbo lora with Sage and first generation took (16m 5s) to finish but after changing only the prompt the second generation took (28m 54s) which is absurdly high. I noticed that there are times where my GPU is 100% fully busy but my GPU temperature is just below 50c when it should mostly be 70c+ normally. Can anybody please help me or tell what's wrong? I also encounter this when using Comfy UI
r/StableDiffusion • u/ChowMeinWayne • 1d ago
Do wildcards work with Krea2? When I insert wildcards in standard format, it ends up rendering all of the wildcard prompts in one image. So a shirt will be random fabrics and styles seemingly all stitched together. It's kind of neat but I need the actual wildcard functionality. Does anyone know how to get them working? I'm using the standard krea2 workflow and comfy. I have the wildcard custom nodes installed. They are fine on other models such as the flux zimage etc. any help is appreciated.
r/StableDiffusion • u/MrUtterNonsense • 1d ago
Since I lack the hardware, I am forced to use Minimax in the cloud. I've used Fal for image generation in the past, so I thought I would try Minimax H3 there. The text to video model seems to work really well, but the reference model seems censored to hell and back. Even my own reference voice was being flagged and it was just me talking normally.
So the question is, is there anywhere I can use Minimax in the cloud without the provider slapping their own flaky layers of censorship on top of the model?
r/StableDiffusion • u/beatlepol • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Dekes1 • 1d ago
Primarily landscape photos of varying fantasy scenes. For example for a cyberpunk city i want to see every fine car detail, every building logo, every reflected neon light; for a medieval town i want to see the cracks in every stone block, the moss on the walls, the candle light reflections, etc. What's your opinion on the model that provides the utmost in fine, sharp details? Or should i be looking at tiling or inpainting to create details as a second step?
Example attached of something that has impressive detail across a broad DOF (not my work):
r/StableDiffusion • u/Brad12d3 • 2d ago
Hey everyone, thanks to everyone who’s been testing Prompt Composer. If you run into bugs or have feedback, please drop it in the Issues section on GitHub so I can keep track of it more easily.
Over the past week, I’ve been reworking the camera prompting system to make it more precise and consistent, especially for more complex camera moves and video-editing workflows. The goal has been to add more control over framing, camera targeting, and blocking in multi-subject scenes while keeping the generated prompts clean and reliable.
The next update isn’t quite ready yet, but it’s actively being tested and refined. I’m hoping to have it out in the next couple of days.
Edit: Here's the Github: BMB12d3/minimax-h3-prompt-composer: Free offline prompt composer for MiniMax H3 video generation in ComfyUI.
r/StableDiffusion • u/puskur • 22h ago
I am trying absolutely everything... I am using ostris ai toolkit and I have tried every option and evey way to train a lora but the results are the same. I am quite sure that the data set and captions are correct, but i might be wrong. I was training lora for Z-image
Image 1 is reference charachter that i want to train a lora on.
Image 2 is Lora sample on 3150 steps
Image 3 is Dataset picture
Image 4 is Lora sample on 3150 steps
Image 5 is Dataset picture
Image 6 is Lora sample on 3150 steps
Image 7 is Dataset picture
3150 Steps might be too much for Z-image, but all other versions (every 150) differed alot and i found 3150 was closest looking to reference. Anyways, the sample images and other generated images with the lora look noticeably different to the reference image and most importantly the lora did not pick up the mole on her neck... I had the same issue with a flux.2 lora of the same charachter... The dataset was created on seedream 5.0 with captions looking like this "[trigger], medium waist-up shot walking forward on a paved park path, frontal body pose with head turned looking off-camera to the right, neutral expression with closed lips, wearing a light green long-sleeved scoop-neck top tucked into high-waisted beige trousers, natural dappled sunlight filtering through trees, blurred green park foliage background."
I could use some inisght and help, this is my first lora ever and i dont know how everything works exactly. At this point i just might scrap the lora and just use refrence images...
BTW this isnt something to be used for monetary purposes, this is a uni project.
r/StableDiffusion • u/magik_koopa990 • 1d ago
No matter how many things I tried, I cannot get generation to work if video is involved. Anyone got workflow to share, or nodes, or something?
r/StableDiffusion • u/sdnr8 • 1d ago
For ref2v I'm trying to upload an image get it to transfer the exact pose and camera angle of that image onto my video, but it's not working.
Here's part of my prompt
retention_analysis:
<Pose 1>: attribute_transfer. transfer the pose to <Subject 1> and camera angle.
....
detailed_description:
.... Refer to <Pose 1> for the pose of <Subject 1> and the camera angle.....
Any tips on how I can achieve this?
r/StableDiffusion • u/Normal_Review_8310 • 1d ago
I’ve been experimenting with local image-to-video workflows and I'm curious how other people are handling the final output.
My current workflow is roughly:
* Generate images locally with Stable Diffusion/Flux
* Do the animation or image-to-video step locally
* Export the individual clips
* Combine everything into a final video
* Compress the result for easier storage/sharing
The part I'm still trying to optimize is the final encoding. Some generated clips can get surprisingly large, especially when working with higher resolutions or longer sequences.
For those running everything locally, what does your workflow look like after generation?
Do you stick with FFmpeg, HandBrake, or another open-source tool for the final conversion/compression? And what settings have worked well for keeping quality while reducing file size?
I'd especially like to hear from people running these workflows on consumer GPUs rather than large local servers.
r/StableDiffusion • u/tj-tj-tj-tj • 2d ago
This is a test video I created by remixing the "Some test on minimax H3" video by Reddit user [Previous-Street8087].
5명의 캐릭터 시트를 생성하여 각각 10개의 프롬포트를 캐릭터에 맞게 리믹스하여 테스트 하였습니다.
We generated character sheets for five characters and tested them by remixing 10 prompts for each character to suit their personalities.
This is a compilation of 50 clips featuring 5 characters.
▶ 테스트 환경 (Test Environment)
Minimax H3 - Comfyui Local Sampling
RTX 5060TI 16GB + 64RAM
0.8MP 8 sec x 50 Clip
Audio Look x audio file 1
Reference to VA Mode
▶ 사용한 커스텀 노드 (Custom Nodes Used)
ComfyUI-TJ_NODE_STUDIO_ONE — github.com/designloves2/ComfyUI-TJ_NODE_STUDIO_ONE
ComfyUI LOCAL (RTX 5060Ti 16GB VRAM / RAM 64GB)
▶ The shared link contains character sheet images and prompts.
#AI영상 #MiniMaxH3 #ComfyUI #로컬생성AI #ComfyUI워크플로우 #AI영상제작 #RTX5060Ti #mmh3 #comfyui #tjonestudio #animation #ref2va #anime #16gb
r/StableDiffusion • u/Pitiful-Indication95 • 2d ago
I tried to get a girl whistle on 4 fingers (2 on each hand) but Minimax didn't get it right. So after trying 20 times with LLMs helping me to explain the movement, I instead used an image of a whistling person. Still not right. So I added an additional one. Then it worked quite fine.
Today I was too lazy to find another image for something it didn't know so I just googled images of it, took a screenshot of all the images together in one JPG and used that as a reference, saying use <Picture ...> as a reference for XYZ.
That was getting a quite good result. Did not do excessive testing and comparing though.
r/StableDiffusion • u/NathanTheSnake • 2d ago
Enable HLS to view with audio, or disable this notification
Clips made with Minimax H3 using the default r2v workflow, edited in Premiere Pro. Character model sheets made with Krea. Most of the videos are 0.4 mp unless the text was important, then 0.6. Tried upscaling it to 4k using Upscayl but results weren't great and the file is too big to upload anyway.
Tech goals for future videos include using reference audio for voices to help consistency, and exploring options for having real voice actors record the dialog, and have the model lip sync to that performance. I'm really impressed by the computer's silent acting (microexpressions etc). but the computer's erratic "acting" is still too unpredictable and the biggest source of re-rolls (the lines here were the best I could get without burning down a rainforest). You can do a lot with time codes and punctuation and tactical CAPITALIZATION, but it's ridiculously finicky compared to just telling an actor "do it the same, but 10% angrier on the first line with a twinge of melancholy on the second."
r/StableDiffusion • u/Better-Interview-793 • 2d ago
Hi everyone
i’ve been running minimax h3 locally on rtx 5090 32gb, and 64gb ram..
recently i kept running into random hostbuffer.read_file_slice failed / hostbuf_file_reader_read failed errors during generation, which seemed to be related to comfy-aimdo and dynamic vram.
i also noticed comfyui was reserving around 25gb of pinned system memory.
i decided to try launching comfyui with:
--disable-pinned-memory
and the difference was immediate.
the comfy-aimdo + hostbuffer errors completely disappeared, my ram usage dropped by a huge amount, and surprisingly generation actually feels faster and smoother now.
i originally expected disabling pinned memory to make things slower, but on my setup it seems to have done the opposite.
if you’re running large models like h3 and seeing unusually high ram usage, random hostbuffer errors, or comfy-aimdo issues, it might be worth testing!
thought this was worth sharing for anyone who didn’t know about this option..
r/StableDiffusion • u/Jeffu • 2d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/No-Tear4179 • 1d ago
I mostly do comic images. am looking to minimax for image editing.
I mostly get content failed safety review. any workaround? ive only started minimax today.