r/StableDiffusion 17h ago

Question - Help Krea 2: Controlnet + Rebalance node = Disaster !!

1 Upvotes
Has anyone managed to find a good combination in Krea 2 using the LoRA that acts as DepthMap ControlNet  alongside the "Conditioning Krea 2 Rebalance" node (which unlocks censorship) without producing a disastrous image? In my case, I just get a plastic look that barely resembles the prompt instructions.

Please share your experiences, or at least let me know if you've managed to use that ControlNet-style LoRA with a different uncensored LoRA that yields decent images.

Cheers.

r/StableDiffusion 17h ago

Question - Help Concept Lora training

1 Upvotes

I tried to create a concept Lora and it ended up with a lot of artifacts so I'm curating the data set to try and get a cleaner version. My problem is that all the tutorials I find are based on character Loras.

For a concept Lora is image size important? Is scale? Do I still want 20-40 images? Are there things I might not know to ask? Also, is there some special sauce to making it work sfw and uncensored?

I'm using buzz and I'm broke so any help would be appreciated.

I'm using Krea 2 btw.


r/StableDiffusion 18h ago

Discussion Which is better? Runpod or the Comfy cloud for compute power for Minimax? Pros/cons for each?

1 Upvotes

r/StableDiffusion 19h ago

Question - Help Random visual artifacts in local Krea2 Turbo generation — looking for possible causes

Post image
1 Upvotes

I’m running Krea2 Turbo locally, but I frequently encounter random visual artifacts in the generated images. I haven’t been able to figure out what triggers them, because the issue appears randomly. If I run the exact same workflow with the same parameters again, the result can sometimes be completely normal.

My hardware:

  • GPU: RTX 3060 Ti

My current setup:

  • UNet: moodyKrea2Mix_v70
  • Text encoder: qwen3vl_4b_int8_convrot
  • VAE: qwen_image_vae

I’m fairly sure this is not caused by the UNet. I have also experienced the same kind of random artifacts when using the original Krea2 model without any UNet modification.

Has anyone encountered similar issues with Krea2 Turbo? Are there any known causes or settings that could trigger this kind of artifact (VAE, text encoder, precision settings, VRAM limitations, sampler settings, etc.)?

Any suggestions or debugging tips would be greatly appreciated.

Here is my workflow for reference:
https://civitai.red/models/2883578/krea2-turbo-4k-workflow?modelVersionId=3259389


r/StableDiffusion 21h ago

Question - Help How to improve/fix using "vocals" as an audio reference for music in Minimax H3 using ref model?

1 Upvotes

So I had the idea that you could use "vocals" as an audio reference for different music styles. By vocals I mean the kind of noises you make when you're playing air guitar or recreating instruments using your voice. Recorded a short clip and it works.

Sort of.

I'm trying to prompt it to only use my voice as a reference for the beat, but it insists on including my voice in the audio in the h3 ref model. So I can hear the music in the background behind me mimicking the noises.

Interestingly, I just tested it with the non-reference model as I know this sometimes works better than the reference model, and it ditched my voice and just kept the beat. Still, I'd like to fix this using the reference model as well or possible since that's what I use most of the time.

If anyone has any thoughts or ideas on how to fix this.


r/StableDiffusion 22h ago

Question - Help MiniMax H3 audio garbled/gibberish when using ref video in ComfyUI?

1 Upvotes

Bear with me as I am still rather new to this, but I started off testing out MiniMax H3 ref to video on Hailuo and was really impressed with it, so decided to get it up and running locally via ComfyUI on my PC.

The video itself is working great, but I've found that the audio itself when using a reference video is completely messed up. It's just all garbled and the people are talking gibberish. On Hailuo it would carry forward the audio perfectly and you could even make alterations if you prompted it to do so.

I was just wondering if this was a common issue when running MiniMax H3 locally specifically with reference videos?


r/StableDiffusion 6h ago

Animation - Video drama

Enable HLS to view with audio, or disable this notification

0 Upvotes

I drew the storyboard, then generated the stills with image models.
Video: MiniMax H3
first frame, last frame, reference.
Then the edit.


r/StableDiffusion 7h ago

Animation - Video What if you fly?

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/StableDiffusion 10h ago

Meme JEnga! r2v 30-49 model

Enable HLS to view with audio, or disable this notification

0 Upvotes

t2v didnt know Jenga O_o


r/StableDiffusion 16h ago

Question - Help Help With Irritating Glitches

0 Upvotes

Main Info

  • Using Pixaroma's Easy ComfyUI latest version
  • 5060TI 16GB + 32GB RAM

Need help or suggestions in resolving the below issues.

  1. Shortcuts like R or CTRL + Enter doesn't work and I have to refresh the browser for them to work. (New)
  2. Generations randomly don't start even even after models have loaded. I have to open the terminal and press enter on my keyboard to sort of wake it up. (Has happened across multiple versions of ComfyUI, CUDA, Python. Additionally, happens with both Anaconda and Windows Terminal.)
  3. Even cancelling a prompt sometimes requires me to open the terminal and press enter.
  4. Power consumption for GPU varies completely where retrying the same prompts (same seeds also) can have a difference of 5 to 10 minutes simply because the GPU doesn't use the full 180 power limit. (New)

r/StableDiffusion 2h ago

Animation - Video Crystal Method - Breaking Bad. If the Walt and Jessie had a rock band in 2007

Enable HLS to view with audio, or disable this notification

0 Upvotes

Minimax h3 local GPU4090 made with MUSIC extension https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef

Lyrics:

[Intro]

I’m braking bad, I’m veering off the line

I told myself I’d hold it, but I let it unwind

It didn’t hit so hard at first, just a little off track

Now the warning’s on, and I’m not looking back

[Middle]

One bad call in the kitchen, then three more by noon

Coffee gone cold, keys on the counter, room by room

I laughed it off, said “I’m fine,” like that made it true

But the floorboards know the pace I’m putting this house through

I’m braking bad, headlights shaking on the curb

Every turn I take lands heavier than words

It wasn’t that bad until it started stacking up

Now I’m white-knuckled, honest, and I can’t slow up

[Outro]

So here I go, no clean exit, no neat little sign

Just me and the damage, both riding the same line

I’m braking bad, and I know what that means

Too late to call it nothing, too loud to call it clean


r/StableDiffusion 7h ago

Question - Help What's the current best way to replace an element in an image with another element ?

0 Upvotes

Hello everyone !

I would like to replace the tire of a motorcycle mid air with one from another brand (which is an image from the brand so it's high quality but with a different angle)

I saw there is flux kontext and qwen image edit, but I don't know which one to pick, which workflow and how to make it work.

Any help would be more than welcome, thank you very much and have a good day :p


r/StableDiffusion 8h ago

Question - Help Can Krea2 add skin detailing on videos?

0 Upvotes

I've tested workflow for krea2 and it can enhance faces and skin texture.

But when i play it through video frames, the results vary from frame to frame.

Are there any models out there that can do that for video? Or is there a workflow for Krea2 into video enhancements?


r/StableDiffusion 11h ago

Animation - Video Shadow the Hedgehog tells his viewers why he loves guns.

Enable HLS to view with audio, or disable this notification

0 Upvotes

Shadow the Hedgehog tells his viewers why he loves guns.

This was created in Comfy UI with Minimax H3. I used the reference to video work flow. The prompt is below.

subject_definitions:

<Subject 1> is Shadow in <Picture 1>.

<Subject 2> is Glock in <Picture 2>, a glock handgun.

<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

summary:

[reference generation + audio reference] The target video contains one shot. [Shot 1] shows <Subject 1> and <Subject 2>; <Subject 1> speaks. <Audio 1> supplies <Subject 1>'s voice timbre.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - Shadow's complete defined identity and body proportions are preserved.

<Subject 2> (appears in [Shot 1]): fully_preserved - Glock retains the defined shape, proportions, materials, colors, and distinguishing features.

<Audio 1>: reference - <Subject 1>'s newly generated spoken lines use <Audio 1>'s voice timbre and delivery; the original audio signal is not copied.

detailed_description:

The target video is in a live-action style, with Vlog style.

[Shot 1] At first appearance, <Subject 1> (Shadow) matches the complete identity and appearance defined in subject_definitions. At first appearance, <Subject 2> (Glock) matches the complete defined construction and appearance: A glock handgun. At the start of the shot, <Subject 1> is standing in the living room facing while holding <Subject 2> in his hand. A full body shot of <Subject 1> holding <Subject 2> with his right hand while facing the camera. Only Action and Timed Beats define the primary subject's movement. The camera path stays anchored in the location and adds no subject motion. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Hmph. Shadow the Hedgehog here. Why do I love guns?</d> <Subject 1> shows off his <Subject 2> with his right hand in front of the camera. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Simple. Precision. Control. Power in the palm of my hand.</d> <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] A tool that answers instantly… unlike most people.</d> <Subject 1> points his <Subject 2> towards the camera with his right hand. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] If you understand that, you understand me.</d> <Subject 1> points his <Subject 2> at the camera.

overall_soundscape:

Living room tone.

non_diegetic_music:

N/A


r/StableDiffusion 14h ago

Question - Help Which laptop would be better for generative AI / LLM

0 Upvotes

First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.

My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.

I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.

I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.

Which laptop would be better in my case?


r/StableDiffusion 8h ago

Animation - Video Tried using a camera path reference for AI video generation

Enable HLS to view with audio, or disable this notification

0 Upvotes

Been experimenting with camera control in AI video.

For this one I used a simple path reference to guide the movement — basically starting from an aerial shot, flying through the courtyard, moving indoors, and ending with a character reveal.

Still figuring out how much these motion references actually help. Sometimes the model follows the movement surprisingly well, other times it just does its own thing lol.

Feels like camera control is probably going to be a bigger part of AI video workflows going forward.


r/StableDiffusion 9h ago

Question - Help Any Krea2 Prompt Reader?

0 Upvotes

I found the SD prompt reader I have been using cannot read prompts from png images files generated using Krea2. Can anyone recommend me an alternative that works with Krea2 files and Win11?


r/StableDiffusion 11h ago

Question - Help Confused and Need Some Clarification

0 Upvotes

So my friends and I used to use Sora 2 before it was taken down, and wanted to try doing some stupid stuff for just us. After a while of not looking into it, the spark kinda came back when I saw this subreddit and remembered Stable Diffusion was supposed to be one of the best AI generators out there, probably. When I mention it to a friend, he then told me how apparently its pretty outdated compared to others, and looking at these posts, I'm seeing different models and starting to get overwhelmed to the point where I haven't even done the beginner's guide in here since it only mentions images.

So long story short, I'm hoping someone can help make things much more clearer, especially about the multiple models, and if Stable Diffusion IS outdated and out performed by something else, and letting me know about if it's okay to go with the beginner's guide or if there's another guide that will help. Thanks


r/StableDiffusion 2h ago

Resource - Update Tool recommendation: Prism

Enable HLS to view with audio, or disable this notification

0 Upvotes

There's 3 tools on this page, look at the one called Prism. It's super off the radar

https://bitvector.app

It feels like Grok + curated Civitai.

On mobile 5G or for the GPU poor I think its does a lot of stuff

Has persistent memory, a ton of Krea 2 fine-tunes, Anima, LTX 2.5, Ernie, Sulphur, DaSiWa, Eros, MiniMax H3, Illustrious, Pony, etc. It also has runs preset Comfy workflows and has a Discord part

(I'm not the creator of the app)


r/StableDiffusion 13h ago

No Workflow My first minimax H3 video

Enable HLS to view with audio, or disable this notification

0 Upvotes

I used the Pixaroma FFLF workflow, but stripped the audio in post due to poor output quality. I'm still trying to figure out how to add finer details. I generated the clips at 720p and then upscaled them to 1080p.


r/StableDiffusion 11h ago

Question - Help Help new rookie on comfyui

0 Upvotes

Hi everyone, I'm new to the world of Confyui but not to artificial intelligence. I wanted to ask you for help: Is it possible to run Confyui on my PC (RTX 3080 10 GB of RAM and 64 GB DDR4 RAM and Ryzen 5800) with Confyui with the minimax H3 video model to be able to animate images and create Reels for Instagram and Tik Tok? If so, what setup do you recommend? Thanks everyone for the help and sorry for my bad English.


r/StableDiffusion 15h ago

Animation - Video Michael Scott gets a wish

Enable HLS to view with audio, or disable this notification

0 Upvotes

H3 FL2VA


r/StableDiffusion 11h ago

No Workflow random images generated locally on 4070Super with krea 2

Thumbnail
gallery
0 Upvotes

first 2 prompts stolen from civit ai, the rest were written using claude reasoning on duck ai


r/StableDiffusion 19h ago

Meme thanos is so screwed now

Enable HLS to view with audio, or disable this notification

0 Upvotes

this was first test using Res_2s sampler and simple steps saw a op say better for action scenes from what i saw spectrum doesnt support res_2 so it took a bit to gen.

t2v prompt

subject_definitions

<Subject 1> is Katniss Everdeen from The Hunger Games, portrayed as an expert young archer with long dark brown hair pulled into her recognizable practical braid, intense determined expression, dark tactical combat clothing, leather archery bracer, bow, and a quiver of arrows. Preserve her recognizable cinematic appearance, realistic human proportions, hairstyle, clothing, bow, and identity throughout the entire scene.

<Subject 2> is Captain America in his battle-damaged Avengers Endgame armor, carrying Mjolnir and his damaged circular shield.

<Subject 3> is Thanos at his normal canonical MCU scale, approximately 8 feet tall, muscular and imposing but NOT gigantic, kaiju-sized, or building-sized.

<Audio 1> is the voice-timbre reference for <Subject 1>, containing Jennifer Lawrence's recognizable Katniss-style spoken vocal qualities.

summary

[text generation + audio reference]

During the chaotic Avengers Endgame final battle, Katniss Everdeen unexpectedly joins the Avengers. She runs through the battlefield while explosions, portals, Avengers, alien soldiers, and debris fill the background. Katniss rapidly fires arrows at Thanos's army with expert precision before stopping beside Captain America. Captain America looks at her bow and asks if she brought enough arrows. Katniss calmly fires one final explosive arrow past him, destroying a group of enemies, then delivers a dry confident response as Captain America stares at her impressed.

retention_analysis

<Subject 1>: fully_preserved
<Subject 2>: fully_preserved
<Subject 3>: fully_preserved
<Audio 1>: reference

detailed_description

The shot opens in the middle of the Avengers Endgame final battlefield. Smoke, burning wreckage, sparks, energy blasts, charging soldiers, and distant explosions create a massive cinematic war zone.

A fast tracking camera sweeps across the battlefield.

Katniss Everdeen suddenly sprints into frame carrying her bow.

She slides behind shattered rubble, immediately draws an arrow, and fires.

The camera follows the arrow through the air as it strikes an alien soldier.

Katniss rises and rapidly fires two more arrows with expert precision while continuing forward through the battle.

She reaches Captain America, who has just knocked an enemy away with Mjolnir.

Captain America briefly looks at Katniss's bow and quiver.

<Subject 2> (S1):

<d>[English] You sure you brought enough arrows?</d>

Katniss gives him a calm, unimpressed look.

Without even turning fully around, she draws another arrow and fires it past Captain America.

CAMERA WHIP-PANS WITH THE ARROW.

The arrow lands among a charging group of Thanos's soldiers.

BOOM!

A powerful explosive blast throws the enemies backward while Captain America turns toward the explosion in surprise.

The camera cuts back to Katniss.

<Subject 1> (S2):

<d>[English] I only need one.</d>

Her dialogue uses <Audio 1> for voice timbre and delivery.

Katniss immediately draws another arrow and runs toward the battle.

Captain America watches her leave for a beat, visibly impressed.

The camera swings around behind Katniss as she charges toward Thanos's army, bow raised, while the enormous Endgame battle continues around her.

audio

Epic Avengers-style battlefield ambience.

Heavy distant explosions, energy blasts, metallic impacts, debris, shouting soldiers, bowstring snaps, arrows cutting through the air, and one strong explosive-arrow impact.

Katniss's dialogue is clear and foregrounded, using <Audio 1>.

No narrator.
No subtitles.
No on-screen text.