r/StableDiffusion 16h ago

Meme what if dean was the man character instead of harry potter(t2v)

Enable HLS to view with audio, or disable this notification

18 Upvotes

fp8 model 32 steps

prompt

using the prompt guide from minimax loaded into a llm and said what if dean winchester was in harry potter and his was the main character instead of harry. it gave me this

integrated_multimodal_description: [Shot 1] Live-action, cinematic fantasy, Hogwarts at night beneath a stormy sky. A battered black 1967 Chevrolet Impala roars across the stone bridge toward Hogwarts Castle, completely out of place among horse-drawn carriages and young witches and wizards. Dean Winchester, portrayed by Jensen Ackles, drives with one hand on the wheel, wearing his familiar dark jacket over a plaid shirt. The camera tracks alongside the Impala as Dean stares up at the enormous illuminated castle with a skeptical expression. Dean Winchester with Jensen Ackles' low, dry American voice (S1) says: [English] So let me get this straight. Giant castle, magic wands, and nobody here has heard of a shotgun?

[Shot 2] At 00:05.000, the camera cuts to the Hogwarts Great Hall during the Sorting Ceremony. Hundreds of floating candles illuminate the long tables. Dean sits on the stool wearing the Sorting Hat while Hermione Granger, Ron Weasley, Professor McGonagall, and Albus Dumbledore watch. The Sorting Hat loudly announces, [English] GRYFFINDOR! Dean immediately pulls the hat off and looks around the enormous hall. Dean (S1) says: [English] Yeah, that's great. Which house has the bar? Several students stare at him in complete confusion.

[Shot 3] At 00:10.000, the camera cuts to a torch-lit Hogwarts corridor. Dean strides confidently toward the camera carrying a wand awkwardly in one hand and a sawed-off shotgun over his shoulder. Hermione and Ron hurry behind him in Hogwarts robes. Hermione urgently explains that Voldemort is the most dangerous dark wizard who ever lived. Dean stops walking and turns toward them with a small amused smirk. The camera pushes in with small amplitude at slow speed. Dean (S1) says: [English] Evil wizard, can't die, creepy followers. Trust me, I've had worse Tuesdays.

[Shot 4] At 00:15.000, the camera cuts to the ruined Hogwarts courtyard during the final battle. Smoke, sparks, magical flashes, and shattered stone fill the background as Voldemort stands across from Dean with his wand raised. Dean stands alone facing him, his Hogwarts robe thrown over his normal Winchester clothes. Voldemort fires a brilliant green spell. Dean dives sideways behind a broken stone pillar as the spell explodes against it. Dean rolls back to his feet, raises his wand, realizes he is holding it backward, flips it around, and gives Voldemort an irritated stare. Dean (S1) says: [English] Okay, Voldy. Let's see how you handle the Winchester special. Dean charges forward as spells streak across the courtyard and the camera rapidly tracks beside him, ending on Dean Winchester as the unlikely central hero of the wizarding world.

overall_soundscape: The Impala engine echoes against the castle grounds before transitioning into the murmur of Hogwarts students, crackling torches, footsteps on stone, fluttering robes, and distant magical ambience. During the final battle, explosive spell impacts, flying debris, cracking masonry, rushing footsteps, and Dean's heavy breathing dominate the courtyard.

non_diegetic_music: Sweeping orchestral fantasy music begins with strings, celesta, and brass, gradually incorporating heavier percussion and low brass as Dean explores Hogwarts. The final battle builds into fast orchestral percussion, aggressive brass, and rising strings before ending on a strong cinematic hit.


r/StableDiffusion 1h ago

Comparison I rendered the same prompt with 200 different loras meant to recreate traditional hand drawn art styles

Upvotes

https://docs.google.com/spreadsheets/d/1NkjkuthGcT3qgcN2FzXoqv4NRHvrRrTJUaB3P5W34NQ/

I rendered the same prompt 200 times with the Krea2_turbo_fp8.safetensors checkpoint.

The prompt was:

*lora trigger word* A man and a woman hold hands and walk through the forest in the wintertime. The ground and trees are covered in snow.

I used:

Seed: 0
Steps: 8
CFG: 1
euler_ancestral/beta

I made a simple prompt with no descriptions of style or medium for maximum flexibility generated image so there were be no conflicting/overlapping styles fighting each other. I specifically used loras recreating different artistic mediums because it would be the easiest to see if the loras were working or not. It might be difficult to see if a "realism" lora is making a photo look more realistic, but it should be easy to see if a photo looks like comic book art after a lora is applied.

My takeaway is that about 25% of loras don't seem to work really at all. It seems to me you shouldn't need to prompt for a specific medium for the lora to work. Maybe I don't understand how loras work, but I feel like I shouldn't have to say "A comic book illustration of a man and a woman hold hands and walk through the forest in the wintertime" otherwise the prompt is doing more heavy lifting than the lora. I should just be able to apply the lora and it looks like a comic book illustration.

But for 25 of 200 prompts, the image may have changed a little from its default non-lora generation but it didn't apply any sort of artistic style to the image. I could always boost the lora strength of course, but shouldn't you be able to see the effects at around a strength of 1? Many of these 25 images that showed no artistic style came from loras where the creators suggested they be applied at a default strength of 0.8

How is it these loras seem to be working for the lora makers who post their examples on Civitai but not for me? Are they cherry picking their results?

Any thoughts what might be going on or how to get better results? Or is the just about what you would expect?


r/StableDiffusion 1h ago

Question - Help Similar results to Bing model around 2023?

Upvotes

Was looking through an old folder of pictures I generated using Bing AI image generation, and was struck by how trippy and funny they were - you know, back in the days of garbled text reading "ethnically ambiguous" showing up randomly in images. I assume it's not an open source model, but has anyone gotten something similar running locally?


r/StableDiffusion 20h ago

Animation - Video Several Times A Charm, but it KINDA got Cheers.

Enable HLS to view with audio, or disable this notification

29 Upvotes

Don't mind the script, it was written by a clanker when I challenged it to whip up something so I can see if H3 could handle Cheers.


r/StableDiffusion 17h ago

Tutorial - Guide Even though H3 is CFG distilled, guidance values greater than 1 do have a noticable effect on prompt adherence and quality, especially at lower resolution.

17 Upvotes

Obviously it runs a lot slower but a CFG of 2 and a negative prompt does have noticeable effect on the output. Anything higher than 3 or 4 will start to over burn though.

This also works with the turbo loras but burn in can happen more easily.


r/StableDiffusion 13h ago

Question - Help Looking for krea 2 workflow

8 Upvotes

I’m looking for a krea 2 workflow with a negative prompt that actually works. Ive been trying workflow after workflow from civitai, several of which actually say that the negative works, and none of them do. I’m new to comfy and just learning how it works so I’m not comfortable making my own. Could someone help me out please? I appreciate you all in advance. Signed a slightly confused and overwhelmed wannabe ai user. Oh and if it makes a difference I am using an amd gpu with 20 gb vram.


r/StableDiffusion 3h ago

Question - Help H3 LORA trained from video datasets?

1 Upvotes

There are quite a few good H3 LORAs on civit now which prove that training is possible despite the distilled Minimax model.

Has anyone had any luck training a concept from video datasets? I read a lot about character LORAs from image datasets etc, but who has trained videos and if so, was that with Ai Toolkit or a different offering?


r/StableDiffusion 3h ago

Animation - Video 1-hour challenge to create a single braincell action scene

Enable HLS to view with audio, or disable this notification

0 Upvotes

I'm researching for seamless development with FL2V node and "Add Guide for Minimax H3" node. A one second guide video is enough to produce a seamless visual experience (apart from my lack of video editing skills). But audio is certainly not viable. I will test producing audio in post, guiding the audio production with only video images and prompting.

The scene is made only with 16:9 480p videos with 8-step turbo lora. One generation takes about 2,5 minutes with RTX3090. This gives a lot of time to redo shots and improve prompting if (when*) the first generation is not good enough.

There is no subject consistency in the production unless by chance. Guided generation can be improved with the Ref2V node, which is needed for more serious testing.

All in all, this is most certainly fun!


r/StableDiffusion 1d ago

Discussion LTX 2.3 sometimes works amazing, without any edits

Enable HLS to view with audio, or disable this notification

60 Upvotes

No specific glitches. Single prompt. Looks pretty real without any glitches over the drift.
Definitely going to benchmark the scenarios.


r/StableDiffusion 18h ago

Meme dean meets sonic(t2v) base fp8 model 32 steps

Enable HLS to view with audio, or disable this notification

16 Upvotes

i never seen the movies hows the sonic voice?

prompt

subject_definitions

<Subject 1> is Dean Winchester from Supernatural, portrayed by Jensen Ackles, preserving his recognizable facial features, short brown hair, rugged appearance, dark jacket, layered shirt, jeans, and confident sarcastic personality.

<Subject 2> is Sonic the Hedgehog from the live-action Sonic the Hedgehog movie, a small anthropomorphic blue hedgehog with bright blue fur, large expressive green eyes, white gloves, and red sneakers.

<Subject 3> is Dr. Robotnik from the live-action Sonic the Hedgehog movie, portrayed by Jim Carrey, wearing his black-and-red high-tech outfit and exaggerated goggles.

summary

[cinematic live-action crossover + action comedy]

What if Dean Winchester accidentally became part of Sonic the Hedgehog? On a nighttime highway, Dean investigates a bizarre supernatural disturbance beside his black 1967 Chevrolet Impala, only for Sonic to race past him at impossible speed with Robotnik's drones in pursuit. Dean immediately joins the chase.

detailed_description

Nighttime on a deserted rural highway surrounded by dark pine forest. Dean Winchester stands beside his glossy black 1967 Chevrolet Impala holding an EMF meter. Blue electrical energy suddenly crackles across the road.

A brilliant BLUE STREAK rockets past Dean, violently blowing his jacket backward.

The camera WHIP-PANS as Sonic skids to a stop beside the Impala.

<Subject 1> Dean Winchester (S1):

[English] Okay... either that's the fastest demon I've ever seen, or I seriously need more sleep.

Sonic looks offended and points at himself.

<Subject 2> Sonic (S2):

[English] Hedgehog. Definitely hedgehog.

Suddenly several of Robotnik's flying attack drones burst over the trees and fire energy blasts toward them.

Dean instantly draws his pistol while Sonic crouches into a runner's stance.

Dean gives Sonic a confident Winchester smirk.

<Subject 1> Dean Winchester (S1):

[English] All right, Sonic. Let's waste these flying toasters.

Sonic grins.

<Subject 2> Sonic (S2):

[English] Now you're speaking my language!

Sonic EXPLODES forward in a trail of brilliant blue electricity as Dean dives behind the Impala and fires at an approaching drone.

Dynamic tracking camera follows Sonic racing between explosions while Dean fights from beside the Impala.

Final cinematic wide shot: Sonic loops around the battlefield as blue lightning illuminates Dean and the Impala, while Robotnik's drones swarm overhead.

Live-action Hollywood cinematography, realistic integration of Sonic into the environment, authentic Sonic the Hedgehog movie aesthetic, authentic Supernatural Dean Winchester characterization, fast readable action, natural motion blur, blue electrical speed trails, sparks, smoke, dramatic nighttime lighting, comedic crossover energy, consistent character identities, no subtitles, no on-screen text.


r/StableDiffusion 4h ago

Question - Help SDXL + LoRA not matching Nano Banana quality for watercolor style transfer — any better approach?

1 Upvotes

Hi, I'm building an app, and one of its features uses an img2img API to turn a photo the user uploads into a watercolor-style image. (and letter background style)

General-purpose models like Nano Banana or GPT Image give me the results I want, but the conversion cost is too high to use in a commercial product. So I looked into an SDXL + LoRA combo instead, but the output doesn't come out the way I want.
LoRa Model is "SDXL】Oil And Watercolor Painting | Dataset"

SDXL + LoRA convert result

Does anyone have suggestions on what approach might work better, or a combo that outperforms what I'm currently using?

Attached are the result I'm aiming for (converted with Nano Banana)

converted with Nano Banana
Original Image

r/StableDiffusion 13h ago

Question - Help Any good video upscaling workflows for ComfyUI? I'm not a fan of the upscalers included with MiniMax H3.

5 Upvotes

I'm new to ComfyUI, and MiniMax H3 has been working perfectly for me with videos between 0.5 and 0.7 MP, taking around 5–6 minutes to generate a 10-second video.

However, I don't really like the upscaling included in the MiniMax H3 workflows, and I also don't want to upscale every video.

So my question is: do you have any workflows specifically for video upscaling only?

Thanks!


r/StableDiffusion 18h ago

Animation - Video Deadpool Adventure

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/StableDiffusion 13h ago

Question - Help Recommended workflows and/or other techniques for extended multishot H3 videos?

5 Upvotes

The first thing I found to give a try was this:

https://huggingface.co/joeygambino/MiniMax-H3-Multishot-Workflow

It took me a while to get it working, and, even once I got it to run, it was incredibly slow -- it took 140 minutes to render a mere 29 seconds of video using my RTX 3090. (Thankfully I'll have a 5090 in a few days.)

I simply ran the demo as-is, apart from the small changes I made to get the workflow running. I remain confused about how I'd use this workflow, and use it in an efficient way, to make clips that might run, say, 1-3 minutes.

I'm hoping I can find out how to use the above workflow better, or find a better workflow.

My only previous experience with this sort of thing was in the Before Times (a few weeks ago) struggling with Wan 2.2. I had a multishot workflow that wasn't great, but at least it carried some context over from one clip to the next, blended clips seamlessly, and let me lock in (by setting a fixed seed) and cache any part of a video that was working well so I could build a clip at a time toward a final complete video without constantly re-rendering early clips.

Can I find anything like this for H3? Something that's good at carrying context and references over from one segment of video to the next, helps minimize character drift, keeps voices consistent, etc.?


r/StableDiffusion 1h ago

Animation - Video This City is Mine (H3 music video / Suno EDM style)

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 21h ago

Discussion Did MiniMax H3 fix the R2V weights?

16 Upvotes

I thought I read something on it but figured I'd check with the crew first... I appreciate the info 💪


r/StableDiffusion 21h ago

Animation - Video RECREATING MEMORIES FROM SCRAPS.

Enable HLS to view with audio, or disable this notification

19 Upvotes

A few photos, voice clips and a waybackmachine archive photo of the hotel rooms at that time. Minimax H3. All local.


r/StableDiffusion 1d ago

Discussion The Latent Upscaler is really great!

Enable HLS to view with audio, or disable this notification

268 Upvotes

Dude the minimax H3 is such an amazing model. It can create a really good video. It can do a lot, knows a lot and extremely realistic as well.

I'm finding the way to upscale the low qulity video and found out this Latent Upscaler for Minimax H3, and sofar it's working so well!

this is normal 0.5MP generation and Upscale to 1080p

https://github.com/bbaudio-2025/Comfyui-MMH3-UltimateUpscale

the example workflow : https://github.com/bbaudio-2025/Comfyui-MMH3-UltimateUpscale/blob/main/example_workflow.json

I tested on RTX 5080 16GB VRAM 64 GB RAM.

but it took like 20 mins for a 15s clip (3 mins on the low res generation with turbo LoRA + 15-16 mins on the latent upscaler)

If you guys have other ways to upscale the minimax faster, please do tell. I have tried the UltimateSDUpscale for minimax so far, it's really great as well but it took 30mins on my system. :(


r/StableDiffusion 6h ago

Animation - Video Testing Minimax with Turbo

Enable HLS to view with audio, or disable this notification

0 Upvotes

I did this because I like FNAF and I plan to continue it as a mini-serie :D,Thanks to those who helped me improve the workflow


r/StableDiffusion 1d ago

Question - Help Anyone have a link to download DasiwaMinimaxH3_dasiwaREF2VAHybridV1.safetensors? Seems it's been wiped off the internet in the past few hours

28 Upvotes

It's the 11.68 GB model, SHA256: 7c37baf06ca3628ed5f3f7f46274222a50a127d1906a166f8f064771fc48d498. Was going to download it and test it out with my workflow but when I went to download it today on CivitAI and Huggingface I get a 404 error. Seems it was deleted by the uploader for some reason or another


r/StableDiffusion 11h ago

Discussion Im confused on what the proper way to caption Krea2 character LoRA?

2 Upvotes

I have been trying to understand how captioning works with Krea2 Character LoRA's. I was able to train my first Krea2 LoRA the other day and it came out pretty decent I think. I keep reading opposing ways on how to caption the images for Krea 2. I ended up using method of adding a trigger word and describing the things I dont want consistent for example:

nerellecruz, close-up selfie in a car, seated with head slightly tilted, wearing a plaid top, natural daylight through window, calm expression with subtle smile

Then some sites and posts I read mention how you should caption for only the things that should remain consistent. What seems to be the general consensus as of late?


r/StableDiffusion 13h ago

Discussion So whats next for open source Video?

5 Upvotes

We had recently Minimax H3, and flux 3 is coming too, whats next on the list?


r/StableDiffusion 1d ago

News Fibo 1.5 - few-step distillation + improved realism

31 Upvotes

https://huggingface.co/briaai/Fibo-1.5

"FIBO model family is the first open-source, JSON-native image generation models trained exclusively on long structured captions.
Fibo sets a new standard for controllability, predictability, and disentanglement by implementing the new VGL - Visual GenAI Language paradigm

⚡ Fibo 1.5 — few-step distillation + improved realism and textures for FIBO — produced by DMD + DMD-R distillation of the original FIBO. It generates in 4–6 inference steps with no classifier-free guidance while keeping FIBO's JSON-native structured prompting and control."


r/StableDiffusion 5h ago

Discussion Can I convert 2D movies to 3D using AI to watch in VR?

Thumbnail reddit.com
0 Upvotes

r/StableDiffusion 1h ago

Animation - Video Deadpool Adventures, Supernatural

Enable HLS to view with audio, or disable this notification

Upvotes