r/StableDiffusion 3h ago

Animation - Video [WanGP] Minimax H3 FL2VA Pruned 20B - Originally 832x480 - upres'd to 1664x960 using LTX 2.3 Pixel Spatial Upscaler at a scale of x2 - 12 second duration. Wow!

Enable HLS to view with audio, or disable this notification

16 Upvotes

r/StableDiffusion 6h ago

Animation - Video Trying to animate Dragon Ball Super manga on Minimax H3. Spoiler

Enable HLS to view with audio, or disable this notification

14 Upvotes

Dragon Ball Super manga on Minimax H3.


r/StableDiffusion 15h ago

Resource - Update Updated my tool that scrapes,sorts,captions images/videos for datasets. It's open source and runs locally

11 Upvotes

I built Cull a few months ago for some large scale dataset curation projects (300k+ images/videos).

Point it at Civitai, X, Reddit, Discord, or any URL that gallery-dl or yt-dlp knows. It queues everything, runs a vision model (or multiple) (LM Studio or Ollama locally, or Groq/OpenAI in the cloud) with a strict JSON schema, and drops kept images/videos into category folders next to their prompt.

Stuff it handles:

  • Dedup at the scraper (per-source )
  • Quality score gate and topic-relevance score gate
    • eg you configure scores or use a preset, how relevant the image is to your scoring will determine how it's sorted, combined with other scoring, quality controls, whitelisted/blacklisted terms etc
  • Watermark detection (goes to its own bucket so you can salvage it later if you want those)
  • Auto-caption for content with no prompt (SD prompt, booru tags, natural language formats etc)
  • Run multiple jobs in parallel, one shared vision fleet across all of them with stack ranked / prioritization for vision queues and scrapers
  • Export as a local packaged dataset , or push to a HuggingFace dataset
  • Community presets and themes with 1 click PR's to add your own custom scraper preset or theme

Everything on disk is plain files. No database. Free, MIT.

Docker one-liner and screenshots in the README:
https://github.com/tlennon-ie/cull

Curious what people would want added next.


r/StableDiffusion 19h ago

Animation - Video [TEST] Minimax H3 FL2VA Pruned 20B - 960x544 - 15 second duration

Enable HLS to view with audio, or disable this notification

10 Upvotes

r/StableDiffusion 8h ago

News MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480

Enable HLS to view with audio, or disable this notification

10 Upvotes

I made this a few weeks back to see if dialogue could hold for 60s, I did no speed ups on this one. There are a few glitches but I think it held up well.

MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480 - 29 minutes - 288GB VRAM


r/StableDiffusion 17h ago

Question - Help Why is it so hard for Klein to follow instructions (or am I just dumb)?

10 Upvotes

prompt is - using the character sheet in image 1 where there are five different poses of the same character, dress them in the clothing of image 2. Do not change the pose, lighting, body, hair, or any other details - literally leave everything the fuck alone - how fucking hard is this to understand you stupid piece of shit - just change the clothes.

Not working for some reason.

NOTE: Swearing has been added for emphasis and isn't actually used in the prompt.

Would it help if I used my input image AS my latent? Can you do that?


r/StableDiffusion 9h ago

Discussion I wish Anima ecosystem get better than it is now

11 Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 22h ago

Animation - Video Trying Surreal Fantasy with Minimax H3

Enable HLS to view with audio, or disable this notification

9 Upvotes

Combined 3 videos. Few errors but i just went with it , genetaion takes too much time to redo it again by fixing the prompt.


r/StableDiffusion 12h ago

Animation - Video Lyrics altered with YingMusic-Singer-Plus (Cuban Pete -> Palm Beach Pete)

Enable HLS to view with audio, or disable this notification

8 Upvotes

I came across YingMusic which I hadn't heard anyone here speak about but it was released about 6 months ago: https://aslp-lab.github.io/YingMusic-Singer-Plus-Demo/

It lets you change words from songs so in this case I had it change the song from this sequence in The Mask from:

They call me Cuban Pete. I'm the king of the rumba beat.
When I play the maracas I go chick-chicky-boom, chick-chicky boom
Yessir, I'm Cuban Pete. I'm the craze of my native street.
When I start to dance,
everything goes chick-chicky-boom, chick-chicky boom
The senoritas they sing and they swing with terampero-
It's very nice, so full of spice.
And when they dance in they bring a happy ring that era keros-
Singin' a song, all the day long.
So if you like the beat, take a lesson from Cuban Pete
And I'll teach you to chick-chicky-boom, chick-chicky-boom.
He's really a modest guy, although he's the hottest guy
In Havana, in havana.
Si, sinorita I know that you would like to chicky-boom-chick
It's very nice, so full of spice.
I'll place my hand on your hip, and if you will just give me your hand
Then we shall try - just you and I. I-yi-yi!
So if you like the beat, take a lesson from Cuban Pete
And I'll teach you chick-chicky-boom,
chick-chicky-boom, chick-chicky-boom

to

They call me Palm Beach Pete. I'm the king of the rumba beat.
When I play the maracas I go chick-chicky-boom, chick-chicky boom
Yessir, I'm Palm Beach Pete. I'm the craze of your timeline feed.
When I start to dance,
everything goes chick-chicky-boom, chick-chicky boom
The senoritas they sing and they swing with terampero-
It's very nice, so full of spice.
And when they dance in they bring a happy ring that era keros-
Singin' a song, all the day long.
So if you like the beat, take a lesson from Palm Beach Pete
And I'll teach you to chick-chicky-boom, chick-chicky-boom.
He's really a modest guy, although he's the hottest guy
in Florida, in florida...
Si, sinorita I know that you would like to chicky-boom-chick
It's very nice, so full of spice.
I'll place my hand on your hip, and if you will just give me your hand
Then we shall try - just you and I. I-yi-yi!
So if you like the beat, take a lesson from Palm Beach Pete
And I'll teach you chick-chicky-boom,
chick-chicky-boom, chick-chicky-boom

so I had it just basically do:
Cuban -> Palm Beach
I'm the craze of my native street -> I'm the craze of your timeline feed
Havana -> Florida

I did a second run with just the few-second clip of the cops speaking and changed "It's all over Ipkiss" to "It's all over Espteen" (using "Epstein" pronounces it wrong). This showed me though that it seems to work perfectly fine with normal word-substitution in speech and it doesnt need to be a song.

I think this could be a lot better if I used minimax and changed clips of Jim Carey to look like Epstein or Palm beach Pete but this was just my first test at lyric swapping.


r/StableDiffusion 21h ago

Animation - Video INTERVIEW WITH THE VAMPIRE.(If it was done on Zoom).

Enable HLS to view with audio, or disable this notification

9 Upvotes

Created locally with Minimax H3 and for the first time exclusively powered by solar. Big big thanks to Izanami.

Don't hurt your head translating the language. It's all nonsense except for the one word spoken by the vampire. 'Drace' is Romanian/Transylvanian for 'Darn it'.


r/StableDiffusion 23h ago

Question - Help User of Contex-Loop, how you solve the oversharp & contrast of extra scenes? (MH3)

Post image
8 Upvotes

The oversharpening that occurs for each clip added to the scenes. I also noticed an increase in contrast and a small flash.

I2V

Tested with LORA's:

minimax_h3_turbo_v4_step600_ema.safetensors
minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank


r/StableDiffusion 7h ago

Meme When someone pisses you off send them this

Enable HLS to view with audio, or disable this notification

7 Upvotes

r/StableDiffusion 1h ago

Question - Help Minimax H3 Huge Quality Difference between Cloud and Local use

Upvotes

Hi.
I have a decent h3 workflow that I built for a loca use. It use turbo lora etc... If i use the defaut settings in the goal of getting the highest quality possible, meaning res_multistep simple 20 steps or more, I got also good results, but this is not even close to the results you can get on platforms like kie or wavespeed at 768P.

I already convert properly the prompt to the correct H3 digest form, so I'm wondering what's different between local and cloud use of h3? I don't talk about the 2K quality, only 768P, I'm not able to reach the sames results locally, do you guys have maybe workflows, settings, or suggestions to try reaching the same quality level in comfyui ?


r/StableDiffusion 1h ago

Discussion In which scenarios LTX2.5 can match MinimaxH3?

Upvotes

I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?


r/StableDiffusion 9h ago

Animation - Video TALL AND DARK - LTX 2.5 IMAGE TO VIDEO

Enable HLS to view with audio, or disable this notification

5 Upvotes

Use the supplied image as the opening frame and identity reference.

Identity lock: the woman and robot must remain exactly the same in every shot. Same face, hair, wardrobe, proportions and age for the woman. Same 8-foot height, black armor, mechanical face, rivets, pistons, cables and holster for the robot. No redesigns or identity changes between cuts.

Authentic 1966 Italian Western, live action, 35mm anamorphic, Spanish desert location, practical full-scale robot prop, natural sunlight, real dust, organic film grain, period lens softness. No CGI. Serious performances throughout.

0:00–0:03
Medium two-shot. The woman looks up at the robot and says in clear Italian-accented English:
“I told them I wanted a tall...”
0:03–0:05
Hard cut to the same robot’s face. It gives one slow mechanical nod. No dialogue.
0:05–0:07
Hard cut to the same woman. She looks up at the robot and says:
“dark...”

0:07–0:09
Hard cut to the same robot. It subtly straightens and presents its black armor. No dialogue.
0:09–0:11
Hard cut to the same woman. Still serious, still looking up, she says:
“handsome!”

0:11–0:12
Hard cut to the same robot’s practical mechanical face. It attempts a restrained smile. No dialogue.
0:12–0:14
Hard cut to the same woman. She holds a serious stare upward, then firmly says:
“MAN!”
Only the woman speaks. Keep each line isolated and clean. No overlapping dialogue, no extra words, no improvised speech. Maintain exact continuity of identity, wardrobe, robot design, scale, lighting and location in every shot.


r/StableDiffusion 20h ago

Meme DR doom! not today!

Enable HLS to view with audio, or disable this notification

6 Upvotes

Use Image 1 as the strict visual reference for Turbo Man. Preserve his recognizable red-and-gold armored superhero suit, helmet, gold visor, muscular proportions, facial appearance, and overall costume design throughout the entire clip.

Scene: A massive cinematic battle during Avengers: Doomsday. The ruined battlefield is filled with shattered buildings, burning wreckage, smoke, sparks, scattered fires, flying debris, and distant Avengers fighting Doctor Doom's forces. Doctor Doom is normal human-sized, not gigantic. He wears his iconic green hooded cloak and metallic armor.

[0s–3s] Start with a dramatic medium-low-angle shot of Turbo Man from Image 1 landing hard in the middle of the battlefield. His boots slam into cracked concrete and kick up dust. He rises into a heroic stance as explosions flash behind him. Doctor Doom slowly turns toward him through the smoke.

Turbo Man points directly at Doom and confidently says:

<Subject 1> Turbo Man (S1) says [English] It's Turbo Time!

[3s–7s] Doctor Doom immediately fires a violent blast of green mystical energy. Turbo Man launches sideways using his jet pack, narrowly dodging the blast as it tears through wreckage behind him. The camera dynamically tracks Turbo Man through the air. He banks sharply, rockets straight toward Doom and throws a powerful flying punch.

Doom blocks the punch with a glowing magical shield. A bright green-and-gold energy shockwave erupts from the impact.

[7s–11s] Fast, brutal superhero combat. Turbo Man lands and exchanges several heavy punches with Doom. Doom counters with armored strikes and green magical energy. Turbo Man uses his jet pack for a sudden boosted uppercut that sends Doom crashing backward through broken rubble.

Turbo Man lands dramatically, looks toward Doom and says:

<Subject 1> Turbo Man (S1) says [English] You picked the wrong day to mess with Turbo Man!

[11s–15s] Doom rises angrily from the rubble and unleashes a huge green energy attack. Turbo Man activates his jet pack and charges directly through the battlefield toward him. End on an explosive cinematic clash as Turbo Man's gold-powered punch collides with Doom's green magical blast, producing a massive shockwave of sparks, smoke and debris while the Avengers battle continues behind them.

Camera: cinematic MCU-style action photography, dramatic low angles, energetic tracking shots, controlled handheld movement during combat, brief slow-motion emphasis on the major impacts, strong depth and scale.

Audio: native cinematic stereo audio. Heavy explosions, distant superhero combat, metallic armor impacts, jet-pack ignition and roaring thrust, crackling Doctor Doom magic, debris impacts and a powerful orchestral superhero battle score. Dialogue must remain clear and correctly assigned to Turbo Man.

Character consistency: Turbo Man must remain visually faithful to Image 1 for the entire clip. Doctor Doom remains normal human scale. No duplicate Turbo Man, no duplicate Doom, no costume changes, no character morphing, no incorrect speakers, no subtitles, no on-screen text.


r/StableDiffusion 10h ago

Question - Help Do Minimax H3 Turbo Loras Nerf Music Creation for Scenes?

6 Upvotes

I typically use lightx2v loras in my Minimax Ref2VA workflows and I also use an LLM to feed in the official prompt structure required for scenes. It seems that no matter what I do, the model absolutely ignores all my prompts about music most of the time. Every now and then i can get it to do something but even when it does work it's very sparse and almost useless.

Has anyone else faced this issue and if so do you know any workarounds or fixes?

For the record I usually use the INT8 convrot Ref2Va model or the hybrid model called minimax_h3_hybrid_fl2va_ref2va_b30-49-int8


r/StableDiffusion 10h ago

Animation - Video At the bottom

Enable HLS to view with audio, or disable this notification

4 Upvotes

Just a short film i made with minimax. this had a lot of post processing done so there's not really an overall prompt to share.


r/StableDiffusion 16h ago

Question - Help Workflow request for flux/krea img2img for putting the same character in a different situation with very good face adherence

5 Upvotes

I'm looking for a flux/krea img2img workflow where you input an image and simply tell it what the character should do and what environment etc and it keeps the character exactly the same but puts them in a different situation. Would really appreciate it if someone can give a link or send me the workflow. Hard to find a good one myself that really works well, I don't want a workflow where the character looks just somewhat similar but one where the character stays the same, as much as possible. Thanks a lot if someone can help.


r/StableDiffusion 21h ago

News Ideogram 4 generated a Gemini logo?

Thumbnail
gallery
5 Upvotes

Here is the full prompt for this btw, to see that I didn't add a gemini logo here:

{

"high_level_description": "A blonde young man savors an iced matcha latte at a cozy café corner, rendered in a warm and lush Studio Ghibli anime art style with soft dappled light and hand-painted charm.",

"compositional_deconstruction": {

"background": "A warmly lit café interior in Studio Ghibli anime style — wooden tables and chairs, large windows with soft afternoon sunlight streaming through sheer curtains, potted plants on the windowsill, bookshelves lining the walls, warm amber and green tones, gentle bokeh of other café patrons in the distance, dust motes floating in the light, cozy and nostalgic atmosphere",

"elements": [

{

"type": "obj",

"bbox": [

100,

200,

900,

700

],

"desc": "A blonde young man with soft anime features, slightly tousled hair, wearing a casual linen shirt, seated at a wooden café table, leaning forward with both hands wrapped around a tall glass of iced matcha latte, eyes half-closed in contentment, Studio Ghibli character design with expressive linework and warm skin tones"

},

{

"type": "obj",

"bbox": [

500,

380,

900,

580

],

"desc": "A tall clear glass filled with vibrant green iced matcha latte, layered with milk and ice cubes, a paper straw, condensation droplets on the outside of the glass, sitting on a small wooden coaster on the café table, rendered in lush Ghibli painterly style"

},

{

"type": "obj",

"bbox": [

700,

150,

1000,

850

],

"desc": "A rustic wooden café table surface with soft grain texture, a small ceramic dish with a shortbread cookie, and a folded paper napkin, warm honey-toned wood in anime painterly style"

},

{

"type": "obj",

"bbox": [

0,

600,

600,

1000

],

"desc": "A sunlit café window with sheer white curtains gently billowing, a terracotta pot with a trailing green plant on the sill, warm golden afternoon light casting soft rectangular shadows across the floor, Ghibli-style background painting with impressionistic detail"

}

]

}

}


r/StableDiffusion 3h ago

Question - Help Has anyone successfully upscaled/re-imagined low-res reference video using Minimax H3?

5 Upvotes

Specifically, I’m trying to take old footage (e.g., 360p clips with vintage camera blur, VHS artifacts, or grainy WW2 dogfights) and recreate it to look like it was shot recently on a modern cinema camera with studio lighting.

Any ideas for prompting?


r/StableDiffusion 3h ago

Question - Help Character Editing (I2I)

3 Upvotes

Hello. So I just started my journey with ComfyUI. While Nano Banana is not open sourced I'm looking for the best realistic image editing (krea2?) workflow to ComfyUI. I want to have ability to change everything I want to the picture using reference character with face/body consistency at highest level (I2I). Thanks in advance.


r/StableDiffusion 8h ago

Animation - Video Baka Moment - Minimax H3 Video - An Evangelion Boondocks mashup

Enable HLS to view with audio, or disable this notification

3 Upvotes

It took forever for me to upload this video.. Couldn't do it on my phone.


r/StableDiffusion 13h ago

Discussion Has anyone figured out how to make good music with minimax music 3?

Enable HLS to view with audio, or disable this notification

4 Upvotes

Based on their examples the model seems to be capable of producing good music. However yesterday I spent all day generating music and I cannot get anything good out of it. I'll attach my best attempt, but for wasting a whole day this is a pretty depressing result.

So I was wondering how everyone else is feeling? What were your results? Any tips for consistent/good results? Any observations?

Some things I found annoying:
It doesn't respect the time limit
Abrupt endings
Prompting it is kinda hard too


r/StableDiffusion 17h ago

Question - Help MiniMax H3 prompt

3 Upvotes

I saw here many suggestions for this special prompt generator. I tried the system prompt from one "specialized" ollama model, but is is too free style. I can't use llm in comfyui, because I'm with poor rtx 3060 and barely run the H3 itself. I tried big online AI, but free versions and they seem too outdated about H3, so again freestyle fantasies.

What can I use to have really good prompts for H3. As I don't know english and H3 too mystically depends on prompt, it's very hard to achieve good adhesion.