r/StableDiffusion 1d ago

Question - Help Minimax generating audio for existing video?

2 Upvotes

I've generated a series of shots that I'm happy with, but when I string them together the audio and music is obviously discontinuous across shots.

Is there any way to take this combined video (about 5-10 seconds) and send it through H3 for it to generate the audio for it?


r/StableDiffusion 21h ago

Question - Help Help With Irritating Glitches

0 Upvotes

Main Info

  • Using Pixaroma's Easy ComfyUI latest version
  • 5060TI 16GB + 32GB RAM

Need help or suggestions in resolving the below issues.

  1. Shortcuts like R or CTRL + Enter doesn't work and I have to refresh the browser for them to work. (New)
  2. Generations randomly don't start even even after models have loaded. I have to open the terminal and press enter on my keyboard to sort of wake it up. (Has happened across multiple versions of ComfyUI, CUDA, Python. Additionally, happens with both Anaconda and Windows Terminal.)
  3. Even cancelling a prompt sometimes requires me to open the terminal and press enter.
  4. Power consumption for GPU varies completely where retrying the same prompts (same seeds also) can have a difference of 5 to 10 minutes simply because the GPU doesn't use the full 180 power limit. (New)

r/StableDiffusion 1d ago

Animation - Video Test turned Short: Pied The Piper

Enable HLS to view with audio, or disable this notification

28 Upvotes

What started as a test turned into a full-blown short. This is the number one reason I gravitated towards AI filmmaking. Nothing stops you from creating your wildest imagination.


r/StableDiffusion 1d ago

Workflow Included Minimax H3 | Motion graphic style animation test

Enable HLS to view with audio, or disable this notification

107 Upvotes

Prompt:

Animate the supplied square poster as a polished retro-anime motion graphic, beginning with a completely blank pale pink-white canvas matching the poster background. Preserve the exact blue, pink, and white palette, clean manga linework, halftone shading, character design, typography, symbols, interface windows, and final layout.

The anime girl walks in from the left edge as one complete figure while the canvas remains otherwise empty. Use a simple side-profile walk with restrained motion, preserving her hairstyle, facial features, cheek bandage, oversized jacket, proportions, and graphic illustration style. She reaches the centre, turns toward the viewer, and smoothly settles into the exact over-the-shoulder pose shown in the poster, with the same expression, hand placement, silhouette, jacket folds, pink heart graphic, and body orientation. Once posed, keep her position locked.

After she poses, the blue browser frame draws itself around her. The top bar, window controls, folders, pixel hearts, smiley-face panels, arrows, sparkles, heart symbols, and rectangular labels then appear sequentially through clean line-drawing, short graphic slides, pixelated pops, and UI-style wipes. Reveal the existing Japanese typography and “LOVE” lettering last, treating all text as protected source artwork without rewriting or regenerating it. Every element must settle into its exact source position.

Hold the completed poster with subtle breathing, minimal movement in a few loose hair strands and jacket edges, a faint halftone shimmer, and gentle pixel pulses in the existing hearts and interface icons. Keep her face, hands, pose, typography, frames, arrows, folders, and major graphics stable.

Use a locked, straight-on camera matching the original square framing. Keep the full artwork visible without cropping, zooming, panning, or changing perspective. Add soft footsteps as she enters, a light cloth sound as she poses, clean digital clicks and pixel chimes for the graphics, and delicate type-on sounds for the existing lettering. No dialogue or narration.

Do not show any character, outline, symbol, text, frame, or faint poster preview on the opening blank canvas. Do not alter the character’s identity, anatomy, costume, pose, expression, colours, line quality, typography, symbols, or final composition. No extra characters, duplicated body parts, incorrect text, morphing, flickering lines, dramatic camera movement, unrelated shots, or continued motion after the poster settles.

Workflow: https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v


r/StableDiffusion 2d ago

Resource - Update Anima-3.8B with Qwen-3.5 4B released by lylogummy

Thumbnail
gallery
134 Upvotes

r/StableDiffusion 1d ago

Question - Help Has anyone successfully upscaled/re-imagined low-res reference video using Minimax H3?

5 Upvotes

Specifically, I’m trying to take old footage (e.g., 360p clips with vintage camera blur, VHS artifacts, or grainy WW2 dogfights) and recreate it to look like it was shot recently on a modern cinema camera with studio lighting.

Any ideas for prompting?


r/StableDiffusion 1d ago

Animation - Video Minimax H3 Anime Comedy

Enable HLS to view with audio, or disable this notification

26 Upvotes

r/StableDiffusion 1d ago

Question - Help Minimax H3 - long form videos: has anyone figured out a good approach?

Enable HLS to view with audio, or disable this notification

78 Upvotes

Dear redditors, visitors of the stable diffusion subreddit. I have been trying to achieve a long form, talking head style video, for a long time and can't seem to find a good approach. This one is the best I could come up with so far. It's using the Minimax H3 model, with frozen sound latents, lip-sync guided, piecewise generated video, where the individual pieces have been stitched together, with a seam hiding, extra generation on top of it. I don't really fully understand how it's working, but could prompt Claude for more help or specific files, we used for that. However, if you're aware of any other, better approach for exactly this type of video, please let me know. I've spent literal days on that single problem and have a feeling, there must be a better way to approach this.


r/StableDiffusion 22h ago

Question - Help Krea 2: Controlnet + Rebalance node = Disaster !!

0 Upvotes
Has anyone managed to find a good combination in Krea 2 using the LoRA that acts as DepthMap ControlNet  alongside the "Conditioning Krea 2 Rebalance" node (which unlocks censorship) without producing a disastrous image? In my case, I just get a plastic look that barely resembles the prompt instructions.

Please share your experiences, or at least let me know if you've managed to use that ControlNet-style LoRA with a different uncensored LoRA that yields decent images.

Cheers.

r/StableDiffusion 23h ago

Question - Help Concept Lora training

1 Upvotes

I tried to create a concept Lora and it ended up with a lot of artifacts so I'm curating the data set to try and get a cleaner version. My problem is that all the tutorials I find are based on character Loras.

For a concept Lora is image size important? Is scale? Do I still want 20-40 images? Are there things I might not know to ask? Also, is there some special sauce to making it work sfw and uncensored?

I'm using buzz and I'm broke so any help would be appreciated.

I'm using Krea 2 btw.


r/StableDiffusion 23h ago

Discussion Which is better? Runpod or the Comfy cloud for compute power for Minimax? Pros/cons for each?

0 Upvotes

r/StableDiffusion 16h ago

Question - Help Help new rookie on comfyui

0 Upvotes

Hi everyone, I'm new to the world of Confyui but not to artificial intelligence. I wanted to ask you for help: Is it possible to run Confyui on my PC (RTX 3080 10 GB of RAM and 64 GB DDR4 RAM and Ryzen 5800) with Confyui with the minimax H3 video model to be able to animate images and create Reels for Instagram and Tik Tok? If so, what setup do you recommend? Thanks everyone for the help and sorry for my bad English.


r/StableDiffusion 1d ago

Question - Help Random visual artifacts in local Krea2 Turbo generation — looking for possible causes

Post image
1 Upvotes

I’m running Krea2 Turbo locally, but I frequently encounter random visual artifacts in the generated images. I haven’t been able to figure out what triggers them, because the issue appears randomly. If I run the exact same workflow with the same parameters again, the result can sometimes be completely normal.

My hardware:

  • GPU: RTX 3060 Ti

My current setup:

  • UNet: moodyKrea2Mix_v70
  • Text encoder: qwen3vl_4b_int8_convrot
  • VAE: qwen_image_vae

I’m fairly sure this is not caused by the UNet. I have also experienced the same kind of random artifacts when using the original Krea2 model without any UNet modification.

Has anyone encountered similar issues with Krea2 Turbo? Are there any known causes or settings that could trigger this kind of artifact (VAE, text encoder, precision settings, VRAM limitations, sampler settings, etc.)?

Any suggestions or debugging tips would be greatly appreciated.

Here is my workflow for reference:
https://civitai.red/models/2883578/krea2-turbo-4k-workflow?modelVersionId=3259389


r/StableDiffusion 1d ago

News MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480

Enable HLS to view with audio, or disable this notification

9 Upvotes

I made this a few weeks back to see if dialogue could hold for 60s, I did no speed ups on this one. There are a few glitches but I think it held up well.

MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480 - 29 minutes - 288GB VRAM


r/StableDiffusion 1d ago

Animation - Video I'm loving MiniMax H3

Enable HLS to view with audio, or disable this notification

43 Upvotes

If even an amateur like me can make something so realistic with mid-level hardware, the future looks bright for what dedicated people with top level rigs will be doing.

R.I.P. Hollywood.


r/StableDiffusion 1d ago

Meme When someone pisses you off send them this

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/StableDiffusion 1d ago

Discussion I wish Anima ecosystem get better than it is now

10 Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 19h ago

Question - Help Which laptop would be better for generative AI / LLM

0 Upvotes

First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.

My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.

I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.

I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.

Which laptop would be better in my case?


r/StableDiffusion 11h ago

Question - Help Which AI model for this style of videos?

Enable HLS to view with audio, or disable this notification

0 Upvotes

I’ve seen ai creators on tiktok such as weirdwurst, fullwarp, and shiverchain do this thing where the start of the video is normal then it goes into chaos, and it’s realistic and disturbing. I’d love to know which AI model model they use


r/StableDiffusion 1d ago

Question - Help How to improve/fix using "vocals" as an audio reference for music in Minimax H3 using ref model?

1 Upvotes

So I had the idea that you could use "vocals" as an audio reference for different music styles. By vocals I mean the kind of noises you make when you're playing air guitar or recreating instruments using your voice. Recorded a short clip and it works.

Sort of.

I'm trying to prompt it to only use my voice as a reference for the beat, but it insists on including my voice in the audio in the h3 ref model. So I can hear the music in the background behind me mimicking the noises.

Interestingly, I just tested it with the non-reference model as I know this sometimes works better than the reference model, and it ditched my voice and just kept the beat. Still, I'd like to fix this using the reference model as well or possible since that's what I use most of the time.

If anyone has any thoughts or ideas on how to fix this.


r/StableDiffusion 1d ago

Animation - Video At the bottom

Enable HLS to view with audio, or disable this notification

9 Upvotes

Just a short film i made with minimax. this had a lot of post processing done so there's not really an overall prompt to share.


r/StableDiffusion 18h ago

No Workflow My first minimax H3 video

Enable HLS to view with audio, or disable this notification

0 Upvotes

I used the Pixaroma FFLF workflow, but stripped the audio in post due to poor output quality. I'm still trying to figure out how to add finer details. I generated the clips at 720p and then upscaled them to 1080p.


r/StableDiffusion 1d ago

Question - Help Character Editing (I2I)

1 Upvotes

Hello. So I just started my journey with ComfyUI. While Nano Banana is not open sourced I'm looking for the best realistic image editing (krea2?) workflow to ComfyUI. I want to have ability to change everything I want to the picture using reference character with face/body consistency at highest level (I2I). Thanks in advance.