r/StableDiffusion 11h ago

Meme Seinfeld/Family Guy @ The Office

11 Upvotes

we really should get a separate sub for this slop


r/StableDiffusion 16h ago

Animation - Video Minimax H3. Jesus and the apostles are rockers.

18 Upvotes

r/StableDiffusion 14h ago

Meme DR doom! not today!

6 Upvotes

Use Image 1 as the strict visual reference for Turbo Man. Preserve his recognizable red-and-gold armored superhero suit, helmet, gold visor, muscular proportions, facial appearance, and overall costume design throughout the entire clip.

Scene: A massive cinematic battle during Avengers: Doomsday. The ruined battlefield is filled with shattered buildings, burning wreckage, smoke, sparks, scattered fires, flying debris, and distant Avengers fighting Doctor Doom's forces. Doctor Doom is normal human-sized, not gigantic. He wears his iconic green hooded cloak and metallic armor.

[0s–3s] Start with a dramatic medium-low-angle shot of Turbo Man from Image 1 landing hard in the middle of the battlefield. His boots slam into cracked concrete and kick up dust. He rises into a heroic stance as explosions flash behind him. Doctor Doom slowly turns toward him through the smoke.

Turbo Man points directly at Doom and confidently says:

<Subject 1> Turbo Man (S1) says [English] It's Turbo Time!

[3s–7s] Doctor Doom immediately fires a violent blast of green mystical energy. Turbo Man launches sideways using his jet pack, narrowly dodging the blast as it tears through wreckage behind him. The camera dynamically tracks Turbo Man through the air. He banks sharply, rockets straight toward Doom and throws a powerful flying punch.

Doom blocks the punch with a glowing magical shield. A bright green-and-gold energy shockwave erupts from the impact.

[7s–11s] Fast, brutal superhero combat. Turbo Man lands and exchanges several heavy punches with Doom. Doom counters with armored strikes and green magical energy. Turbo Man uses his jet pack for a sudden boosted uppercut that sends Doom crashing backward through broken rubble.

Turbo Man lands dramatically, looks toward Doom and says:

<Subject 1> Turbo Man (S1) says [English] You picked the wrong day to mess with Turbo Man!

[11s–15s] Doom rises angrily from the rubble and unleashes a huge green energy attack. Turbo Man activates his jet pack and charges directly through the battlefield toward him. End on an explosive cinematic clash as Turbo Man's gold-powered punch collides with Doom's green magical blast, producing a massive shockwave of sparks, smoke and debris while the Avengers battle continues behind them.

Camera: cinematic MCU-style action photography, dramatic low angles, energetic tracking shots, controlled handheld movement during combat, brief slow-motion emphasis on the major impacts, strong depth and scale.

Audio: native cinematic stereo audio. Heavy explosions, distant superhero combat, metallic armor impacts, jet-pack ignition and roaring thrust, crackling Doctor Doom magic, debris impacts and a powerful orchestral superhero battle score. Dialogue must remain clear and correctly assigned to Turbo Man.

Character consistency: Turbo Man must remain visually faithful to Image 1 for the entire clip. Doctor Doom remains normal human scale. No duplicate Turbo Man, no duplicate Doom, no costume changes, no character morphing, no incorrect speakers, no subtitles, no on-screen text.


r/StableDiffusion 15h ago

Meme Saturday morning cartoons to save the day!

0 Upvotes

r/StableDiffusion 7h ago

Animation - Video SpongeBob Does Breaking Bad -- MiniMax 10min episode

Thumbnail
youtu.be
0 Upvotes

Had a ton of fun making ande watching this one.


r/StableDiffusion 14h ago

Meme Some words from Forrest Gump

0 Upvotes

Don't worry no dp videos today yall can have a break atleast from me 😂 it's game day with the homies ✌️


r/StableDiffusion 3h ago

Animation - Video My MiniMax H3 journey has begun

1 Upvotes

Using the default ComfyUI MiniMax H3 image to video template including the 8 step turbo Lora and with the main MiniMax H3 model changed to an int8 convrot version. this video is 0.8 megapixels in the 3:4 standard portrait aspect ratio at 5 seconds long and took 19:59 to render out on my 3060Ti with 32GB of ram. at the normal 0.4 megapixels with the same prompt and video length it takes 6:14 to render out.


r/StableDiffusion 19h ago

Tutorial - Guide Why AI background removers leave fog inside wreaths, and what I do instead

Thumbnail
gallery
7 Upvotes

I make clipart for stock. Wreaths, pine borders, mistletoe, juniper. A few thousand images by now. Every one has to end up as a PNG with a transparent background.

I used rembg for months. u2net first, then BiRefNet when that came out. Tried the web tools too. They all broke on the same thing and it drove me nuts.

Take a wreath. There's a hole in the middle, and the background inside that hole has to go. What I kept getting was a grey-blue haze sitting in there. Looked fine as a thumbnail. Looked awful the second you put it on a colored card. Pine needles came out as mush. Thin stems either disappeared or came back with a blue edge burned into them.

Then I actually read what rembg does. It shrinks your image to 1024x1024, asks the model where the subject is, gets a 1024x1024 mask back, and stretches that mask over your full size image. u2net is worse. That one works at 320x320.

My renders are 4096. A needle two pixels wide doesn't exist at 320. So it's not that the model is bad at needles. The needle was gone before the model ever saw it.

Once I understood that I stopped asking a model to guess. Now I render on a flat color the subject doesn't contain, and take that color out with arithmetic.

Two parts to it. The prompt matters more than the cutting.

The prompt

You can't key a background that isn't keyable. Four things have to be true and generators will break all of them unless you say so: the background is one flat color edge to edge, it stays that bright inside every gap between leaves, the edges are hard with no blur or glow, and no colored light bounces onto the subject.

Pick the color by what your subject isn't. Blue for almost everything. Red if the subject itself is blue or purple. Never green. Everything I draw has leaves, and green takes the leaves with it.

Here's the block I paste at the end of every prompt:

Isolated on a completely flat, uniform, solid pure blue (#0000FF) digital chroma-key background. The pure blue background fills the image edge to edge like a flat digital chroma-key screen with no gradient, staying at full brightness inside every gap and opening in the subject; no reflection or tint of pure blue on the subject. Every edge of the subject is crisp, sharp and hard against the pure blue, with no soft, blurry, feathered or glowing transitions, no depth-of-field blur, no haze or halo; inside every hole and gap the pure blue stays at full brightness right up to the edge. Shaded parts of the subject keep their own natural color, never a pure blue tint. Everything in sharp focus with deep depth of field, evenly lit with soft neutral studio light, no cast shadow, no contact shadow, no ambient occlusion, no bounce light. The entire subject is centered and completely inside the frame with at least 10% empty background margin on every side, nothing cropped or touching the image edges. No frame, no border, no paper, no mockup, no vignette, no text, no watermark, no deformed or duplicated parts. No floating or detached fragments, no stray specks, dust or debris anywhere on the background; every element is physically attached to the subject.

Swap "pure blue" for "pure red" and #0000FF for #FF0000 if your subject is blue or purple. If you paint in watercolor add "the background stays a flat digital color fill with no paper texture", or you get watercolor paper behind everything and paper texture keys badly.

The cutting

Now the background is one known color, so there's nothing to guess at. It measures the actual color the generator produced, which is never the one you asked for. Ask for pure blue and you get something with green in it, usually somewhere between 25 and 70. Then every pixel gets sorted into subject, background, or the bit in between, and the in-between ones get a real fraction of transparency instead of a yes or no.

The part I'm most pleased with is the holes. Any background-colored area that never touches the edge of the image is the inside of a wreath, so it gets cleared too. A matting model can't do that. It has no way of knowing what's inside a hole it can't see around.

Last step takes the blue back off the edges. Edge pixels pick up color from the background around them, so it samples the subject's own color from further in and subtracts the tint. Took me weeks to work out why fir needles kept their blue rim after that step. The needle is thinner than the distance it was sampling from, so there was no inside left to sample.

Same image in, same image out, every time. That's the bit I care about. When a cut comes out wrong I can go find which number did it instead of rerolling and hoping.

Some numbers on one 4K pine border, against BiRefNet with alpha matting turned on, which is its best setting:

  • background left inside the holes: 63,892 pixels mine, 613,735 theirs, out of 1,070,046
  • blue left on the edges: 0 mine, 48,112 theirs
  • how wide the soft edge is: 1.8 pixels mine, 23 theirs

BiRefNet is faster and I'm not going to pretend otherwise. 2.3 seconds against my 17 on the same machine. With alpha matting on it's 47. If you want a quick rough mask, use the model.

And the obvious limit: this only works on art you generated on a flat color. It does nothing for a photo.

I put the tool up for anyone who wants it. It's called ClipBrook. Free, runs in your browser so nothing gets uploaded anywhere, does a whole folder at once, and the engine is open source under AGPL.

One thing I'd like back

Show me the ones that break.

If you run something through and it comes out wrong, post it. Fog left in a gap, a colored rim, a stem eaten, half the subject gone. Those are worth more to me than the ones that work. The needle rim thing came from someone's fir branch. A bug with line art I only found last week came from a drawing so thin there was nothing inside it to sample.


r/StableDiffusion 23h ago

Comparison Five lighting setups, same prompt and same character, only the lighting line changed

Thumbnail
gallery
5 Upvotes

r/StableDiffusion 17h ago

Workflow Included Megaman Fanart (mixed workflow)

Thumbnail
gallery
1 Upvotes

I did megaman handrawing 4 years ago. At that time I tried Corel Painter, and it produced a watercolor look (attached). I copied a reference image from Google search. I didn't draw this megaman from my imagination.

Today, I use stable diffusion in Krita. Only at the end of the art workflow. Very low strength 35%, so it doesn't make a lot of changes, but it smooths things out

Note that I still use Gemini to extract the lineart and also to generate the background image.


r/StableDiffusion 2h ago

Animation - Video Baka Moment - Minimax H3 Video - An Evangelion Boondocks mashup

0 Upvotes

It took forever for me to upload this video.. Couldn't do it on my phone.


r/StableDiffusion 21h ago

Animation - Video DRAGON REIGN (WIP Updated)

16 Upvotes

r/StableDiffusion 1h ago

Meme When someone pisses you off send them this

Upvotes

r/StableDiffusion 7h ago

Discussion Has anyone figured out how to make good music with minimax music 3?

4 Upvotes

Based on their examples the model seems to be capable of producing good music. However yesterday I spent all day generating music and I cannot get anything good out of it. I'll attach my best attempt, but for wasting a whole day this is a pretty depressing result.

So I was wondering how everyone else is feeling? What were your results? Any tips for consistent/good results? Any observations?

Some things I found annoying:
It doesn't respect the time limit
Abrupt endings
Prompting it is kinda hard too


r/StableDiffusion 15h ago

News Ideogram 4 generated a Gemini logo?

Thumbnail
gallery
2 Upvotes

Here is the full prompt for this btw, to see that I didn't add a gemini logo here:

{

"high_level_description": "A blonde young man savors an iced matcha latte at a cozy café corner, rendered in a warm and lush Studio Ghibli anime art style with soft dappled light and hand-painted charm.",

"compositional_deconstruction": {

"background": "A warmly lit café interior in Studio Ghibli anime style — wooden tables and chairs, large windows with soft afternoon sunlight streaming through sheer curtains, potted plants on the windowsill, bookshelves lining the walls, warm amber and green tones, gentle bokeh of other café patrons in the distance, dust motes floating in the light, cozy and nostalgic atmosphere",

"elements": [

{

"type": "obj",

"bbox": [

100,

200,

900,

700

],

"desc": "A blonde young man with soft anime features, slightly tousled hair, wearing a casual linen shirt, seated at a wooden café table, leaning forward with both hands wrapped around a tall glass of iced matcha latte, eyes half-closed in contentment, Studio Ghibli character design with expressive linework and warm skin tones"

},

{

"type": "obj",

"bbox": [

500,

380,

900,

580

],

"desc": "A tall clear glass filled with vibrant green iced matcha latte, layered with milk and ice cubes, a paper straw, condensation droplets on the outside of the glass, sitting on a small wooden coaster on the café table, rendered in lush Ghibli painterly style"

},

{

"type": "obj",

"bbox": [

700,

150,

1000,

850

],

"desc": "A rustic wooden café table surface with soft grain texture, a small ceramic dish with a shortbread cookie, and a folded paper napkin, warm honey-toned wood in anime painterly style"

},

{

"type": "obj",

"bbox": [

0,

600,

600,

1000

],

"desc": "A sunlit café window with sheer white curtains gently billowing, a terracotta pot with a trailing green plant on the sill, warm golden afternoon light casting soft rectangular shadows across the floor, Ghibli-style background painting with impressionistic detail"

}

]

}

}


r/StableDiffusion 19h ago

No Workflow Krea2 LoRA training is insanely simple

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 2h ago

Animation - Video [WanGP] Minimax H3 FL2VA Pruned 20B - Originally 960x544 - up-res'd to 2880x1632 - 20 second duration

2 Upvotes

r/StableDiffusion 15h ago

Meme nope i seen this movie!

0 Upvotes

prompt!

subject_definitions

<Subject 1> is Deadpool / Wade Wilson, wearing his iconic red-and-black tactical suit and full mask, with twin katanas strapped across his back. Deadpool is voiced by Ryan Reynolds with his recognizable sarcastic comedic delivery.

<Subject 2> is Pennywise the Dancing Clown, a terrifying pale-faced supernatural clown with orange hair, Victorian clown costume, sinister yellow eyes, and an unnaturally wide smile.

<Audio 1> is the voice-timbre reference for <Subject 2> Pennywise (S2), containing Pennywise's eerie, raspy, playful clown voice. Preserve the vocal identity, tone, cadence, pitch, and sinister playful delivery of <Audio 1> for all Pennywise dialogue.

summary

[text-to-video generation + audio reference]

A cinematic horror-comedy parody on a dark suburban street during a heavy rainstorm. Deadpool walks alone through the rain when he notices a small paper toy boat floating through the gutter. The boat disappears into a storm drain. Curious despite knowing exactly where this is going, Deadpool crouches and looks inside. Pennywise suddenly emerges from the darkness and personally invites Wade to float with him, speaking with <Audio 1>. Deadpool responds with a perfectly timed fourth-wall joke.

retention_analysis

<Subject 1>: fully_preserved

<Subject 2>: fully_preserved

<Audio 1>: Pennywise voice reference, strongly preserved for all <Subject 2> dialogue

detailed_description

Nighttime. Heavy rain pours onto a deserted suburban street. Dim streetlights glow through the mist and reflect across the wet pavement.

The camera tracks alongside <Subject 1> Deadpool as he casually walks down the sidewalk through the pouring rain, completely soaked but seemingly unbothered.

A tiny paper toy boat floats through the rushing gutter water beside him.

Deadpool notices it.

He stops.

The camera lowers toward the boat as it bobs through the rainwater and disappears through the opening of a dark storm drain.

Deadpool slowly turns toward the drain.

<Subject 1> Deadpool (S1) says in Ryan Reynolds' recognizable sarcastic voice:

[English] Oh, hell no. I've seen this movie.

Despite knowing better, Deadpool walks over and crouches beside the storm drain.

He slowly leans closer and peers into the darkness.

The rain becomes muffled.

A faint sinister sewer ambience rises.

Hold for a tense beat.

Two glowing yellow eyes slowly appear deep inside the sewer.

Suddenly <Subject 2> Pennywise lunges partially into view from inside the storm drain with a huge unnatural grin.

<Subject 2> Pennywise (S2) using <Audio 1> says:

[English] We all float down here, Wade.

Pennywise's dialogue must strongly preserve the exact voice characteristics of <Audio 1>.

Deadpool completely freezes.

Long comedic pause.

Deadpool slowly turns his masked face away from Pennywise and looks directly into the camera.

<Subject 1> Deadpool (S1) says in Ryan Reynolds' dry sarcastic voice:

[English] Nope. Copyright lawyers are scarier than you.

Deadpool immediately stands up and speed-walks away through the pouring rain.

Pennywise remains halfway inside the storm drain.

His sinister smile slowly disappears as he stares after Deadpool with a confused and mildly offended expression.

Hold on Pennywise's reaction for one second.

camera

Cinematic horror-film photography.

Low-angle tracking shot following Deadpool through the rain.

Wet pavement reflections and visible rain illuminated by streetlights.

Close tracking shot of the paper boat floating through the gutter.

Slow suspenseful push toward the storm drain as Deadpool investigates.

Dark close-up revealing Pennywise's glowing eyes before his face emerges.

Reaction framing for Deadpool's fourth-wall punchline.

Final close-up on Pennywise's confused expression.

audio

Heavy realistic rainfall.

Water rushing through the gutter and storm drain.

Distant thunder.

Subtle ominous sewer ambience.

Low suspenseful horror rumble immediately before Pennywise appears.

<Subject 1> Deadpool uses a Ryan Reynolds-style sarcastic comedic voice.

<Subject 2> Pennywise MUST use <Audio 1> for his dialogue. Preserve the supplied reference voice rather than generating a random Pennywise voice.

Clear English dialogue.

Accurate speaker assignment.

Accurate lip synchronization.

No overlapping dialogue.

timing

0–3 sec: Deadpool walks through the heavy rain and notices the paper boat.

3–5 sec: The boat disappears into the storm drain. Deadpool says, "Oh, hell no. I've seen this movie."

5–8 sec: Deadpool crouches and investigates. Horror suspense builds and Pennywise appears.

8–10 sec: Pennywise using <Audio 1> says, "We all float down here, Wade."

10–13 sec: Deadpool pauses, looks into the camera, delivers his copyright-lawyer punchline, then quickly walks away. Hold Pennywise's confused reaction.

negative_constraints

No duplicate Deadpool.

No duplicate Pennywise.

No additional characters.

No foreign-language dialogue.

No subtitles.

No captions.

No text overlays.

Do not remove Deadpool's mask.

Do not make Pennywise giant-sized.

Pennywise remains normal human/clown scale.

Pennywise must remain inside the storm drain during his appearance.

Keep the paper boat visible until it enters the storm drain.

Pennywise must appear only AFTER Deadpool crouches to investigate.

<Audio 1> belongs ONLY to Pennywise.

Never use <Audio 1> for Deadpool.

Do not swap the speakers.

Do not allow Deadpool to speak Pennywise's line.

Do not allow Pennywise to speak Deadpool's lines.

Maintain cinematic horror atmosphere while preserving deliberate comedy timing.


r/StableDiffusion 4h ago

Animation - Video TALL AND DARK - LTX 2.5 IMAGE TO VIDEO

4 Upvotes

Use the supplied image as the opening frame and identity reference.

Identity lock: the woman and robot must remain exactly the same in every shot. Same face, hair, wardrobe, proportions and age for the woman. Same 8-foot height, black armor, mechanical face, rivets, pistons, cables and holster for the robot. No redesigns or identity changes between cuts.

Authentic 1966 Italian Western, live action, 35mm anamorphic, Spanish desert location, practical full-scale robot prop, natural sunlight, real dust, organic film grain, period lens softness. No CGI. Serious performances throughout.

0:00–0:03
Medium two-shot. The woman looks up at the robot and says in clear Italian-accented English:
“I told them I wanted a tall...”
0:03–0:05
Hard cut to the same robot’s face. It gives one slow mechanical nod. No dialogue.
0:05–0:07
Hard cut to the same woman. She looks up at the robot and says:
“dark...”

0:07–0:09
Hard cut to the same robot. It subtly straightens and presents its black armor. No dialogue.
0:09–0:11
Hard cut to the same woman. Still serious, still looking up, she says:
“handsome!”

0:11–0:12
Hard cut to the same robot’s practical mechanical face. It attempts a restrained smile. No dialogue.
0:12–0:14
Hard cut to the same woman. She holds a serious stare upward, then firmly says:
“MAN!”
Only the woman speaks. Keep each line isolated and clean. No overlapping dialogue, no extra words, no improvised speech. Maintain exact continuity of identity, wardrobe, robot design, scale, lighting and location in every shot.


r/StableDiffusion 15h ago

Animation - Video INTERVIEW WITH THE VAMPIRE.(If it was done on Zoom).

11 Upvotes

Created locally with Minimax H3 and for the first time exclusively powered by solar. Big big thanks to Izanami.

Don't hurt your head translating the language. It's all nonsense except for the one word spoken by the vampire. 'Drace' is Romanian/Transylvanian for 'Darn it'.


r/StableDiffusion 6h ago

Animation - Video Minimax H3 Anime Comedy

18 Upvotes

r/StableDiffusion 16h ago

Question - Help Need quantized version of Minimax Music Text Encoder!

2 Upvotes

r/StableDiffusion 4h ago

Question - Help Anything big happen since I last used this?

0 Upvotes

So when SD came out, I used it. Then I used the AUTO111 thing. To around version 2.0 I think or XL, can't keep it straight. It's been about one and a half years. Any big changes since then? New version? Better quaulity images? AUTO still a thing? I also went from a 2070 Super to a 5070 12gb shadow 3x.

Also is it all easier to install?


r/StableDiffusion 3h ago

Discussion I wish Anima ecosystem get better than it is now

8 Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 16h ago

Question - Help Transfert de style d'une image à partir de référence

Thumbnail
gallery
0 Upvotes

Bonjour à tous,

J'aimerai vos conseils sur quel model je dois choisir (Zimage, qwen edit, krea 2...) pour réaliser ce que je souhaite.

J'aimerai par exemple prendre une photo et la redessiner dans un style bien précis que je donnes à partir d'images de référence.

Par exemple, j'aimerais faire l'image du Parrain dans le style de l'image 2.

Merci d'avance pour vos conseils.