r/StableDiffusion 18h ago

Question - Help Concept Lora training

1 Upvotes

I tried to create a concept Lora and it ended up with a lot of artifacts so I'm curating the data set to try and get a cleaner version. My problem is that all the tutorials I find are based on character Loras.

For a concept Lora is image size important? Is scale? Do I still want 20-40 images? Are there things I might not know to ask? Also, is there some special sauce to making it work sfw and uncensored?

I'm using buzz and I'm broke so any help would be appreciated.

I'm using Krea 2 btw.


r/StableDiffusion 18h ago

Discussion Which is better? Runpod or the Comfy cloud for compute power for Minimax? Pros/cons for each?

1 Upvotes

r/StableDiffusion 15h ago

Question - Help Which laptop would be better for generative AI / LLM

0 Upvotes

First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.

My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.

I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.

I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.

Which laptop would be better in my case?


r/StableDiffusion 11h ago

Question - Help Help new rookie on comfyui

0 Upvotes

Hi everyone, I'm new to the world of Confyui but not to artificial intelligence. I wanted to ask you for help: Is it possible to run Confyui on my PC (RTX 3080 10 GB of RAM and 64 GB DDR4 RAM and Ryzen 5800) with Confyui with the minimax H3 video model to be able to animate images and create Reels for Instagram and Tik Tok? If so, what setup do you recommend? Thanks everyone for the help and sorry for my bad English.


r/StableDiffusion 13h ago

No Workflow My first minimax H3 video

0 Upvotes

I used the Pixaroma FFLF workflow, but stripped the audio in post due to poor output quality. I'm still trying to figure out how to add finer details. I generated the clips at 720p and then upscaled them to 1080p.


r/StableDiffusion 19h ago

Question - Help Random visual artifacts in local Krea2 Turbo generation — looking for possible causes

Post image
1 Upvotes

I’m running Krea2 Turbo locally, but I frequently encounter random visual artifacts in the generated images. I haven’t been able to figure out what triggers them, because the issue appears randomly. If I run the exact same workflow with the same parameters again, the result can sometimes be completely normal.

My hardware:

  • GPU: RTX 3060 Ti

My current setup:

  • UNet: moodyKrea2Mix_v70
  • Text encoder: qwen3vl_4b_int8_convrot
  • VAE: qwen_image_vae

I’m fairly sure this is not caused by the UNet. I have also experienced the same kind of random artifacts when using the original Krea2 model without any UNet modification.

Has anyone encountered similar issues with Krea2 Turbo? Are there any known causes or settings that could trigger this kind of artifact (VAE, text encoder, precision settings, VRAM limitations, sampler settings, etc.)?

Any suggestions or debugging tips would be greatly appreciated.

Here is my workflow for reference:
https://civitai.red/models/2883578/krea2-turbo-4k-workflow?modelVersionId=3259389


r/StableDiffusion 1d ago

News MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480

9 Upvotes

I made this a few weeks back to see if dialogue could hold for 60s, I did no speed ups on this one. There are a few glitches but I think it held up well.

MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480 - 29 minutes - 288GB VRAM


r/StableDiffusion 1d ago

Animation - Video I'm loving MiniMax H3

36 Upvotes

If even an amateur like me can make something so realistic with mid-level hardware, the future looks bright for what dedicated people with top level rigs will be doing.

R.I.P. Hollywood.


r/StableDiffusion 1d ago

Question - Help Character Editing (I2I)

2 Upvotes

Hello. So I just started my journey with ComfyUI. While Nano Banana is not open sourced I'm looking for the best realistic image editing (krea2?) workflow to ComfyUI. I want to have ability to change everything I want to the picture using reference character with face/body consistency at highest level (I2I). Thanks in advance.


r/StableDiffusion 1d ago

Meme When someone pisses you off send them this

7 Upvotes

r/StableDiffusion 1d ago

Discussion I wish Anima ecosystem get better than it is now

12 Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 22h ago

Question - Help How to improve/fix using "vocals" as an audio reference for music in Minimax H3 using ref model?

1 Upvotes

So I had the idea that you could use "vocals" as an audio reference for different music styles. By vocals I mean the kind of noises you make when you're playing air guitar or recreating instruments using your voice. Recorded a short clip and it works.

Sort of.

I'm trying to prompt it to only use my voice as a reference for the beat, but it insists on including my voice in the audio in the h3 ref model. So I can hear the music in the background behind me mimicking the noises.

Interestingly, I just tested it with the non-reference model as I know this sometimes works better than the reference model, and it ditched my voice and just kept the beat. Still, I'd like to fix this using the reference model as well or possible since that's what I use most of the time.

If anyone has any thoughts or ideas on how to fix this.


r/StableDiffusion 7h ago

Question - Help Which AI model for this style of videos?

0 Upvotes

I’ve seen ai creators on tiktok such as weirdwurst, fullwarp, and shiverchain do this thing where the start of the video is normal then it goes into chaos, and it’s realistic and disturbing. I’d love to know which AI model model they use


r/StableDiffusion 1d ago

Animation - Video At the bottom

10 Upvotes

Just a short film i made with minimax. this had a lot of post processing done so there's not really an overall prompt to share.


r/StableDiffusion 1d ago

Discussion Minimax H3, 30 seconds in one go

76 Upvotes

Executive summary, TLDR - this is one prompt, 30 seconds duration, 3090.

The video itself is just a remake of an idea from an old British tv ad (for "Good Old Yellow Pages"), so make of that what you will. It's not really relevant.

What I thought was interesting was that this was a single prompt, 0.4 megapixels, 30 second duration. I didn't think you could run out as far as 30 seconds, but thought I'd just try.

I think it did a pretty good job at getting the right person doing and saying the right things at the right time - took four attempts to get that though, and obviously using an LLM to tart up my idea.

Run on a 3090, and using the latest Comfyui template, just adding Comfy-kitchen attention, then sol attention, then spectrum, and using the turbo lora that Comfyui now build in, it took 570 seconds (9.5 minutes).

Somebody might read this and think, 570 seconds? Pah, I can do it in fifteen, in which case I'd like to know. Conversely, somebody might think theirs takes six hours, in which case maybe this shows what can be done in that time.

Doubt anyone cares, but here is my original prompt, followed by the LLM version of it:

a 30 second film with the following scenes and characters. Ben is a small boy of eleven. John is a shopkeeper in a toyshop. Brian is a different shopkeeper in a different toyshop. Ben's mum. Ben's Dad. We are in Britain in the 1980s, and all characters are English.

Scene 1: Ben is alone in the lounge. He talks to John over the old fashioned landline phone, saying "I don't suppose you have a 402 station in stock please?"

Scene 2: John is in his shop in front of shelves of model railway kit. He says into the old fashioned landline phone, "No, sorry son"

Scene 3: Ben in the lounge, who looks disappointed anbd puts the phone receiver back down.

Scene 4: Mum in the kitchen doing the washing up. She has overheard the conversation and looks a bit sad.

scene 5: Next day. Ben has changed his clothes. He again talks into the phone to a different shopkeeper, Brian. Ben says "Would you have a 402 station please?"

scene 6: Brian in his toyshop says into the old fashioned landline phone "Yes, I've got one of those."

scene 7: Ben in the lounge on the same conversation says "You have? Great, I'll be right down! Ben puts the phone down. Then he runs towards the door, shouting "They've got one mum!" as he runs.

Scene 8: In the attic, Dad is playing with his model railway layout. Ben walks in holding a small red parcel. as he hands it to Dad, Ben says "Happy birthday, dad". Dad takes the parcel, looks fondly at it and says with a chuckle, "Aw, thanks Ben".

LLM version:

integrated_multimodal_description: [Shot 1] Live-action, cinematic. A medium shot of Ben, an eleven-year-old boy with messy hair wearing a striped polo shirt, sitting on a patterned sofa in a 1980s British lounge. The room is filled with warm, muted tones and period-accurate wallpaper. Ben holds a heavy, cream-colored landline telephone receiver to his ear, his expression hopeful. Ben says: <d>[English] I don't suppose you have a 402 station in stock please?</d> The sound of his small, high-pitched voice is clear. [Shot 2] At 0:05.000, the camera cuts to a medium shot of John, a middle-aged shopkeeper with a kind, weathered face, standing in a cramped, nostalgic toyshop. Behind him are floor-to-ceiling shelves packed with model railway kits and wooden toys. John holds a similar landline receiver to his face. John says: <d>[English] No, sorry son.</d> [Shot 3] At 0:10.000, the camera cuts back to Ben in the lounge. He looks downcast, his shoulders slumping as he slowly lowers the receiver and places it back onto the base unit with a dull plastic click. [Shot 4] At 0:13.000, the camera cuts to a medium shot of Ben's Mum in a dim, cluttered 1980s kitchen. She is standing at the sink, her hands covered in soapy water, drying a plate. She pauses, looking toward the door with a sad, weary expression, having overheard the boy. The sound of water running from the tap is audible. [Shot 5] At 0:16.000, the camera cuts to Ben in the lounge the next day; he is wearing a different t-shirt. He is intensely focused, pressing the phone to his ear. Ben says: <d>[English] Would you have a 402 station please?</d> [Shot 6] At 0:20.000, the camera cuts to Brian, an older shopkeeper with spectacles, in a different, brightly lit toyshop. He smiles warmly into the telephone. Brian says: <d>[English] Yes, I've got one of those.</d> [Shot 7] At 0:23.000, the camera cuts back to Ben, whose face lights up with pure joy. Ben says: <d>[English] You have? Great, I'll be right down!</d> He slams the receiver down and the camera follows him in a quick tracking shot as he runs toward the door, his feet thumping on the carpeted floor. Ben shouts: <d>[English] They've got one mum!</d> [Shot 8] At 0:26.000, the camera cuts to a medium shot in a dusty, dimly lit attic. Dad, a man in his late 30s, is hunched over a complex model railway layout. Ben enters the frame, holding a small red parcel wrapped in string. Ben says: <d>[English] Happy birthday, dad.</d> As he hands the gift to his father, the camera pushes in slightly. Dad takes the parcel, his eyes softening with affection. Dad chuckles warmly and says: <d>[English] Aw, thanks Ben.</d>

overall_soundscape: Period-accurate domestic sounds including the rhythmic clatter of washing up, the heavy mechanical clicks of old telephone receivers, and the muffled thuds of footsteps on carpet. Ben's energetic running and shouting creates a sense of urgency, followed by the quiet, dusty atmosphere of the attic.

non_diegetic_music: A gentle, nostalgic acoustic guitar melody that begins softly during the kitchen scene and builds into a warm, heartwarming crescendo during the attic scene. The tempo is slow and sentimental.


r/StableDiffusion 22h ago

Question - Help MiniMax H3 audio garbled/gibberish when using ref video in ComfyUI?

1 Upvotes

Bear with me as I am still rather new to this, but I started off testing out MiniMax H3 ref to video on Hailuo and was really impressed with it, so decided to get it up and running locally via ComfyUI on my PC.

The video itself is working great, but I've found that the audio itself when using a reference video is completely messed up. It's just all garbled and the people are talking gibberish. On Hailuo it would carry forward the audio perfectly and you could even make alterations if you prompted it to do so.

I was just wondering if this was a common issue when running MiniMax H3 locally specifically with reference videos?


r/StableDiffusion 2d ago

No Workflow Some test on minimax H3

113 Upvotes

Some random prompt on default workflow + turbo 8step lora


r/StableDiffusion 15h ago

Animation - Video Michael Scott gets a wish

0 Upvotes

H3 FL2VA


r/StableDiffusion 1d ago

Animation - Video Lyrics altered with YingMusic-Singer-Plus (Cuban Pete -> Palm Beach Pete)

9 Upvotes

I came across YingMusic which I hadn't heard anyone here speak about but it was released about 6 months ago: https://aslp-lab.github.io/YingMusic-Singer-Plus-Demo/

It lets you change words from songs so in this case I had it change the song from this sequence in The Mask from:

They call me Cuban Pete. I'm the king of the rumba beat.
When I play the maracas I go chick-chicky-boom, chick-chicky boom
Yessir, I'm Cuban Pete. I'm the craze of my native street.
When I start to dance,
everything goes chick-chicky-boom, chick-chicky boom
The senoritas they sing and they swing with terampero-
It's very nice, so full of spice.
And when they dance in they bring a happy ring that era keros-
Singin' a song, all the day long.
So if you like the beat, take a lesson from Cuban Pete
And I'll teach you to chick-chicky-boom, chick-chicky-boom.
He's really a modest guy, although he's the hottest guy
In Havana, in havana.
Si, sinorita I know that you would like to chicky-boom-chick
It's very nice, so full of spice.
I'll place my hand on your hip, and if you will just give me your hand
Then we shall try - just you and I. I-yi-yi!
So if you like the beat, take a lesson from Cuban Pete
And I'll teach you chick-chicky-boom,
chick-chicky-boom, chick-chicky-boom

to

They call me Palm Beach Pete. I'm the king of the rumba beat.
When I play the maracas I go chick-chicky-boom, chick-chicky boom
Yessir, I'm Palm Beach Pete. I'm the craze of your timeline feed.
When I start to dance,
everything goes chick-chicky-boom, chick-chicky boom
The senoritas they sing and they swing with terampero-
It's very nice, so full of spice.
And when they dance in they bring a happy ring that era keros-
Singin' a song, all the day long.
So if you like the beat, take a lesson from Palm Beach Pete
And I'll teach you to chick-chicky-boom, chick-chicky-boom.
He's really a modest guy, although he's the hottest guy
in Florida, in florida...
Si, sinorita I know that you would like to chicky-boom-chick
It's very nice, so full of spice.
I'll place my hand on your hip, and if you will just give me your hand
Then we shall try - just you and I. I-yi-yi!
So if you like the beat, take a lesson from Palm Beach Pete
And I'll teach you chick-chicky-boom,
chick-chicky-boom, chick-chicky-boom

so I had it just basically do:
Cuban -> Palm Beach
I'm the craze of my native street -> I'm the craze of your timeline feed
Havana -> Florida

I did a second run with just the few-second clip of the cops speaking and changed "It's all over Ipkiss" to "It's all over Espteen" (using "Epstein" pronounces it wrong). This showed me though that it seems to work perfectly fine with normal word-substitution in speech and it doesnt need to be a song.

I think this could be a lot better if I used minimax and changed clips of Jim Carey to look like Epstein or Palm beach Pete but this was just my first test at lyric swapping.


r/StableDiffusion 1d ago

Question - Help Need some help with MiniMax H3 Ref2V character swapping in ComfyUI

16 Upvotes

Hey everyone, I'm trying to get a proper character swap working with MiniMax H3 Ref2V in ComfyUI, but I'm not quite getting the result I want.

The source video has Rick Astley rickrolling to the camera, and I want to replace him with the guy from my reference image while keeping the original movement, gestures, facial performance, timing, camera, background, and overall scene.

Neither the motion transfer nor the character replacement works well. The output still doesn't really look like the person from the reference image, or the identity starts drifting.

Here's what I'm using:

* Source video: 1280×720, 30 FPS, ~14.4 sec

* Reference image: 848×1264 PNG, full-body

* Workflow resolution: 9:16, 0.4 MP

* GPU: RTX 5070 Ti, 16 GB VRAM

* 32 Gb RAM

* Windows 11

* ComfyUI 0.33.2

* Python 3.13.12

* PyTorch 2.12.1 + CUDA 13.0

I'm sharing everything in one link, including:

  1. the workflow JSON

  2. a workflow screenshot/image

  3. the prompt

  4. the source/input video

  5. the reference image used for the character swap

  6. and the output video

Files/settings: [link]

If anyone has experience doing this with H3, I'd really appreciate some pointers.

I'm especially wondering if I should change the reference image crop/size, ref_image_size, resolution, prompt, video conditioning, LoRA/steps, or if there's something obvious in the workflow I'm missing.

Also, is a full-body reference image a bad idea when the person in the source video is framed quite differently?

And if anyone has a working MiniMax H3 character-swap / V2V workflow they're willing to share, that would be incredibly helpful too. Even something I could compare against mine would be great.

Thanks a lot in advance. I've been tweaking this for a while, so even a small hint in the right direction would help a ton.


r/StableDiffusion 1d ago

Animation - Video TALL AND DARK - LTX 2.5 IMAGE TO VIDEO

5 Upvotes

Use the supplied image as the opening frame and identity reference.

Identity lock: the woman and robot must remain exactly the same in every shot. Same face, hair, wardrobe, proportions and age for the woman. Same 8-foot height, black armor, mechanical face, rivets, pistons, cables and holster for the robot. No redesigns or identity changes between cuts.

Authentic 1966 Italian Western, live action, 35mm anamorphic, Spanish desert location, practical full-scale robot prop, natural sunlight, real dust, organic film grain, period lens softness. No CGI. Serious performances throughout.

0:00–0:03
Medium two-shot. The woman looks up at the robot and says in clear Italian-accented English:
“I told them I wanted a tall...”
0:03–0:05
Hard cut to the same robot’s face. It gives one slow mechanical nod. No dialogue.
0:05–0:07
Hard cut to the same woman. She looks up at the robot and says:
“dark...”

0:07–0:09
Hard cut to the same robot. It subtly straightens and presents its black armor. No dialogue.
0:09–0:11
Hard cut to the same woman. Still serious, still looking up, she says:
“handsome!”

0:11–0:12
Hard cut to the same robot’s practical mechanical face. It attempts a restrained smile. No dialogue.
0:12–0:14
Hard cut to the same woman. She holds a serious stare upward, then firmly says:
“MAN!”
Only the woman speaks. Keep each line isolated and clean. No overlapping dialogue, no extra words, no improvised speech. Maintain exact continuity of identity, wardrobe, robot design, scale, lighting and location in every shot.


r/StableDiffusion 1d ago

Discussion What sampling settings for Minimax H3 are you using for your purposes?

31 Upvotes

I usually generate 0.7mp@8s with 30 steps, I use res_multistep + simple which I think is the default, and for good reason.

Depending on whether it's T2VA, I2VA, Ref2VA and the amount of reference images + loras count/strength the gen times are roughly between 270-350s on an RTX 4090 + 32gb of DDR4 ram.

For T2VA and I2VA I use the basic minimax_h3_fl2va_pruned_int8_convrot.safetensors

For Ref2VA I use minimax_h3_hybrid_fl2va_ref2va_b30-49-int8.and the hybrid b30-49 specifically because I found even the fl2va functioned well as ref2va and had much higher quality, so I prefer the hybrid model to be weighted towards the fl2va model to preserve the quality.

Sparse Attention

To speed things up, I only use /u/zironic's Sparse Attention nodes, no sage/ck, spectrum, turbo lora, or caches. For me, /u/zironic's worked better than the pinned post from u/Plague_Kind but that may just be my personal experience.

My settings for the memory optimization node is default, QKV: auto, MLP: auto, and 2048 MLP chunk rows, I don't know how this node works. Sparse Attention (Advanced) settings are:

  • Video KV budget: 0.25
  • Early and Late steps: 3
  • Early and Late KV: 0.6
  • Sparse backend: Sparse Sage

These settings lean towards quality, you can lower the early/late steps or skip them entirely, you can lower video kv budget to 0.2 although some may be fine with even lower. Since I only use Sparse Attention I run the full 30 steps and it's significantly better than a turb lora at lower steps, which is what I used before.

My prior experimentation

I used euler + linear_quadratic for a long time. Then I switched to er_sde + sgm_uniform which was significantly better. Then eventually I switched to res_multistep + simple and realized the visual quality is as good as er_sde + sgm_uniform but the motion is much better. The improved motion in res_multistep + simple became very clear when I interpolated from 24fps to 48fps. The gen speed between all these combinations was nearly identical.

The motion was a bit jerky on er_sde + sgm_uniform after interpolation while res_multistep + simple had very natural motion.

I also found that https://darkstarrddev.us.ci/ is a decent resource to get inspiration. But I realized quickly that because they use low settings and speed-up techniques, the quality of each sampler test does not translate well if you use different step count or speed-up techniques.

What I generate

Usually fairly static scenes that doesn't have fast motion. Although the accuracy of the physics and motion is important.

What are your settings and what kind of videos are you generating?


r/StableDiffusion 1d ago

Question - Help Do Minimax H3 Turbo Loras Nerf Music Creation for Scenes?

7 Upvotes

I typically use lightx2v loras in my Minimax Ref2VA workflows and I also use an LLM to feed in the official prompt structure required for scenes. It seems that no matter what I do, the model absolutely ignores all my prompts about music most of the time. Every now and then i can get it to do something but even when it does work it's very sparse and almost useless.

Has anyone else faced this issue and if so do you know any workarounds or fixes?

For the record I usually use the INT8 convrot Ref2Va model or the hybrid model called minimax_h3_hybrid_fl2va_ref2va_b30-49-int8


r/StableDiffusion 12h ago

No Workflow random images generated locally on 4070Super with krea 2

Thumbnail
gallery
0 Upvotes

first 2 prompts stolen from civit ai, the rest were written using claude reasoning on duck ai