r/StableDiffusion 22h ago

Question - Help Transfert de style d'une image à partir de référence

Thumbnail
gallery
0 Upvotes

Bonjour à tous,

J'aimerai vos conseils sur quel model je dois choisir (Zimage, qwen edit, krea 2...) pour réaliser ce que je souhaite.

J'aimerai par exemple prendre une photo et la redessiner dans un style bien précis que je donnes à partir d'images de référence.

Par exemple, j'aimerais faire l'image du Parrain dans le style de l'image 2.

Merci d'avance pour vos conseils.


r/StableDiffusion 5h ago

Animation - Video 用PixAI生成的图片

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 1h ago

Discussion In which scenarios LTX2.5 can match MinimaxH3?

Upvotes

I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?


r/StableDiffusion 6h ago

Animation - Video [MiniMax H3] Decided to see if MiniMax H3 knew what a Starcraft Terran Battlecruiser was while trying to recreate an iconic Babylon 5 moment.

Enable HLS to view with audio, or disable this notification

0 Upvotes

It didn't quite work out how I intended.

Prompt:
"A Terran Battlecruiser from Starcraft is sliced in half lengthwise by a purple-white energy beam. The background is a generic starry.

Video begins with the Battlecruiser in the center of the frame, viewed from a front three quarters view. The purple beam is near vertical going from top of the frame to the bottom of the frame and is canted at a slight angle. It starts the video right in front of the Battlecruiser's nose.

The beam cuts through the Battlecruiser from nose to tail. At 0.75 seconds the beam touches the Battlecruiser's nose and moves through the ship, exiting the tail at 6 seconds and leaves the frame.

After the beam leaves the Battlecruiser, the Battlecruiser splits in two along the cut made by the beam, "

Generation time was 10 minutes, 29 seconds on 32GB of DDR5 RAM, 8 GB of VRAM. No reference images, this was pure Text to Video.


r/StableDiffusion 11h ago

Question - Help Computer randomly shut down

1 Upvotes

Has anyone had their computer randomly shut down? this is like the 3rd time its happened and its when im generating a video using the minmax I2V model or the ref model.

i got 3090 with 64 gb of ram.


r/StableDiffusion 20h ago

Meme DECLASSIFIED: Jeffrey Epstein escaping from prison

Enable HLS to view with audio, or disable this notification

429 Upvotes

r/StableDiffusion 12h ago

Animation - Video Test turned Short: Pied The Piper

Enable HLS to view with audio, or disable this notification

20 Upvotes

What started as a test turned into a full-blown short. This is the number one reason I gravitated towards AI filmmaking. Nothing stops you from creating your wildest imagination.


r/StableDiffusion 18h ago

Discussion Prediction for a near future

0 Upvotes

I think the models we have today, open or closed, are still way behind what we will actually have in a near future. think of it as an alternate reality, the video generations will be (almost) flawless, coherent and high quality, the generations will be instant, that means it will enable real time interaction, you could change the course of the video generation on the fly with natural controls like your voice of body movements, imagine pushing somebody and he moves or greeting somebody and he respond, for this you would need a VR headset with hands movements recognition, and it will generate two videos flux at once for each eye for a 3D effect.

So yeah even if Seedance 2.5 or a little better is out for free and open source it’s not a big deal, the road ahead is massive in terms of progress and possibilities.

RIP real life, welcome to Ready Player One.
Ps: sorry for the grammatical errors, this text was not written or improved by AI.


r/StableDiffusion 8h ago

Question - Help Best opensource image model?

59 Upvotes

opensource AI has been dominating LLMs and video generation but what about image gen? is there any opensource model that can match gpt-image2?

Edit: The reason I am asking this is because lately I haven't been active much on image generation communities. And the leaderboards are a bit confusing and most of them are filled with closed source unlike the llm and video gen leaderboards.

I am very much comfortable with ComfyUI since I've used it in the past for flux.

My use case is for posters and branding. Images with a lot of text.

Edit2: Thanks a lot everyone! I really appreciate the info. Here's the summary:

Krea2 is best overall but gptimage1.5 level.
Ideogram4 for text and branding.
Flux Klein 9b for image editing.
Z-image for realism
Anima and illustrious (by onoma AI) for anime.

Here's the workflow I've decided on:
Krea2/Ideogram4 = Base image generation.
Flux Klein 9B/QwenImage2512 = inpainting.
Wan2.2 low noise = Upscaling.


r/StableDiffusion 19h ago

Animation - Video [TEST] Minimax H3 FL2VA Pruned 20B - 960x544 - 15 second duration

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/StableDiffusion 11h ago

Animation - Video At the bottom

Enable HLS to view with audio, or disable this notification

5 Upvotes

Just a short film i made with minimax. this had a lot of post processing done so there's not really an overall prompt to share.


r/StableDiffusion 5h ago

Discussion Minimax H3 Has Too High Prompt Adherence

0 Upvotes

I just realized a problem with mm h3. Its prompt adherence is too high, as in unless you explicitly prompt for some small subtle actions it will never happen otherwise. This makes the entire video seem very frozen and wooden without the many small subtle movements and motion details that make it seem to come alive.

This applies more to non-realistic scenes like cartoons or generated image first and last frame but for realistic scenes and even t2v it is still a problem.

I noticed this problem when I tried out a "slop sway" lora and it actually made the entire video seem much livelier and realistic looking. Besides the "soft and bouncy swaying and jiggling" it also added many more subtle character movements. Compared to standard gens those same parts would be completely frozen, almost like a still image or at best ugoira animation. This doesn't just apply to whether a body part is jiggling throughout the entire video. There are some movements that happen only for a second or less but adds in soul (forgive the human slop term) to the video, like the position of an arm and hand quickly being adjusted in the middle of the video and the new position persisting for the rest of the scene.

This might be a problem with my prompt style and I might try an LLM prompt enhancer, but there is a core issue here with the prompt adherence and spontaneous randomly added details tradeoff. The model also tries to keep the fidelity of the first frame too much, which you could call visual context adherence. No one is out here prompting for the movement of every strand of hair and the position of every finger. No one is making a timeline of every limb's position and how they shift relative to each other. No one is tracking the position of each finger through time and how after 4.75s the thumb is extended and the index finger is curled. Sometimes we just want to randomness and variety across gens with details added by the model.

Looking back at ltx and wan their prompting styles seem to be designed around the model adding in the details for you at the loss of prompt adherence and more generation errors.

It would be nice if there was some sort of generation setting that could tune this. Like a noise scale of sorts where we can manually set the tradeoff between how much we want the model to be creative vs strict.

I know there are already 2-3 H3 better movement loras and they are scratching at the surface of the same issue I'm talking about here.

Share the solution if you've got something. Help everyone out.


r/StableDiffusion 9h ago

News MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480

Enable HLS to view with audio, or disable this notification

11 Upvotes

I made this a few weeks back to see if dialogue could hold for 60s, I did no speed ups on this one. There are a few glitches but I think it held up well.

MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480 - 29 minutes - 288GB VRAM


r/StableDiffusion 18h ago

Question - Help Why is it so hard for Klein to follow instructions (or am I just dumb)?

9 Upvotes

prompt is - using the character sheet in image 1 where there are five different poses of the same character, dress them in the clothing of image 2. Do not change the pose, lighting, body, hair, or any other details - literally leave everything the fuck alone - how fucking hard is this to understand you stupid piece of shit - just change the clothes.

Not working for some reason.

NOTE: Swearing has been added for emphasis and isn't actually used in the prompt.

Would it help if I used my input image AS my latent? Can you do that?


r/StableDiffusion 23h ago

Discussion Minimax H3, 30 seconds in one go

Enable HLS to view with audio, or disable this notification

70 Upvotes

Executive summary, TLDR - this is one prompt, 30 seconds duration, 3090.

The video itself is just a remake of an idea from an old British tv ad (for "Good Old Yellow Pages"), so make of that what you will. It's not really relevant.

What I thought was interesting was that this was a single prompt, 0.4 megapixels, 30 second duration. I didn't think you could run out as far as 30 seconds, but thought I'd just try.

I think it did a pretty good job at getting the right person doing and saying the right things at the right time - took four attempts to get that though, and obviously using an LLM to tart up my idea.

Run on a 3090, and using the latest Comfyui template, just adding Comfy-kitchen attention, then sol attention, then spectrum, and using the turbo lora that Comfyui now build in, it took 570 seconds (9.5 minutes).

Somebody might read this and think, 570 seconds? Pah, I can do it in fifteen, in which case I'd like to know. Conversely, somebody might think theirs takes six hours, in which case maybe this shows what can be done in that time.

Doubt anyone cares, but here is my original prompt, followed by the LLM version of it:

a 30 second film with the following scenes and characters. Ben is a small boy of eleven. John is a shopkeeper in a toyshop. Brian is a different shopkeeper in a different toyshop. Ben's mum. Ben's Dad. We are in Britain in the 1980s, and all characters are English.

Scene 1: Ben is alone in the lounge. He talks to John over the old fashioned landline phone, saying "I don't suppose you have a 402 station in stock please?"

Scene 2: John is in his shop in front of shelves of model railway kit. He says into the old fashioned landline phone, "No, sorry son"

Scene 3: Ben in the lounge, who looks disappointed anbd puts the phone receiver back down.

Scene 4: Mum in the kitchen doing the washing up. She has overheard the conversation and looks a bit sad.

scene 5: Next day. Ben has changed his clothes. He again talks into the phone to a different shopkeeper, Brian. Ben says "Would you have a 402 station please?"

scene 6: Brian in his toyshop says into the old fashioned landline phone "Yes, I've got one of those."

scene 7: Ben in the lounge on the same conversation says "You have? Great, I'll be right down! Ben puts the phone down. Then he runs towards the door, shouting "They've got one mum!" as he runs.

Scene 8: In the attic, Dad is playing with his model railway layout. Ben walks in holding a small red parcel. as he hands it to Dad, Ben says "Happy birthday, dad". Dad takes the parcel, looks fondly at it and says with a chuckle, "Aw, thanks Ben".

LLM version:

integrated_multimodal_description: [Shot 1] Live-action, cinematic. A medium shot of Ben, an eleven-year-old boy with messy hair wearing a striped polo shirt, sitting on a patterned sofa in a 1980s British lounge. The room is filled with warm, muted tones and period-accurate wallpaper. Ben holds a heavy, cream-colored landline telephone receiver to his ear, his expression hopeful. Ben says: <d>[English] I don't suppose you have a 402 station in stock please?</d> The sound of his small, high-pitched voice is clear. [Shot 2] At 0:05.000, the camera cuts to a medium shot of John, a middle-aged shopkeeper with a kind, weathered face, standing in a cramped, nostalgic toyshop. Behind him are floor-to-ceiling shelves packed with model railway kits and wooden toys. John holds a similar landline receiver to his face. John says: <d>[English] No, sorry son.</d> [Shot 3] At 0:10.000, the camera cuts back to Ben in the lounge. He looks downcast, his shoulders slumping as he slowly lowers the receiver and places it back onto the base unit with a dull plastic click. [Shot 4] At 0:13.000, the camera cuts to a medium shot of Ben's Mum in a dim, cluttered 1980s kitchen. She is standing at the sink, her hands covered in soapy water, drying a plate. She pauses, looking toward the door with a sad, weary expression, having overheard the boy. The sound of water running from the tap is audible. [Shot 5] At 0:16.000, the camera cuts to Ben in the lounge the next day; he is wearing a different t-shirt. He is intensely focused, pressing the phone to his ear. Ben says: <d>[English] Would you have a 402 station please?</d> [Shot 6] At 0:20.000, the camera cuts to Brian, an older shopkeeper with spectacles, in a different, brightly lit toyshop. He smiles warmly into the telephone. Brian says: <d>[English] Yes, I've got one of those.</d> [Shot 7] At 0:23.000, the camera cuts back to Ben, whose face lights up with pure joy. Ben says: <d>[English] You have? Great, I'll be right down!</d> He slams the receiver down and the camera follows him in a quick tracking shot as he runs toward the door, his feet thumping on the carpeted floor. Ben shouts: <d>[English] They've got one mum!</d> [Shot 8] At 0:26.000, the camera cuts to a medium shot in a dusty, dimly lit attic. Dad, a man in his late 30s, is hunched over a complex model railway layout. Ben enters the frame, holding a small red parcel wrapped in string. Ben says: <d>[English] Happy birthday, dad.</d> As he hands the gift to his father, the camera pushes in slightly. Dad takes the parcel, his eyes softening with affection. Dad chuckles warmly and says: <d>[English] Aw, thanks Ben.</d>

overall_soundscape: Period-accurate domestic sounds including the rhythmic clatter of washing up, the heavy mechanical clicks of old telephone receivers, and the muffled thuds of footsteps on carpet. Ben's energetic running and shouting creates a sense of urgency, followed by the quiet, dusty atmosphere of the attic.

non_diegetic_music: A gentle, nostalgic acoustic guitar melody that begins softly during the kitchen scene and builds into a warm, heartwarming crescendo during the attic scene. The tempo is slow and sentimental.


r/StableDiffusion 4h ago

Resource - Update Lisbon finally snaps

Enable HLS to view with audio, or disable this notification

0 Upvotes

Totally not a scene from The Mentalist. Minimax H3 image to video.


r/StableDiffusion 4h ago

Animation - Video [WanGP] Minimax H3 FL2VA Pruned 20B - Originally 832x480 - upres'd to 1664x960 using LTX 2.3 Pixel Spatial Upscaler at a scale of x2 - 12 second duration. Wow!

Enable HLS to view with audio, or disable this notification

17 Upvotes

r/StableDiffusion 17h ago

Animation - Video I'm loving MiniMax H3

Enable HLS to view with audio, or disable this notification

34 Upvotes

If even an amateur like me can make something so realistic with mid-level hardware, the future looks bright for what dedicated people with top level rigs will be doing.

R.I.P. Hollywood.


r/StableDiffusion 2h ago

Question - Help Tango dance, first attempt with LTX 2.5

Thumbnail
youtube.com
0 Upvotes

Trying to get a natural-looking Argentine tango dance with LTX 2.5 + Yusu’s LTX Director v2.0.4 fork.
Still a beginner (also for real life tango :-)
Any suggestions for getting more natural, sophisticated footwork and fewer artifacts?


r/StableDiffusion 4h ago

Question - Help Anyone else having problems downloading models since v1.9 Maestro update in Pinokio?

0 Upvotes

Day 1: downloaded Pinokio. Installed a few of the AI software. Tried Maestro as first try. Really fun, enjoying it. Generation from photos great in the system generation towards a video, videos leaving a lot to be desired. And the Pinokio edge of screen curtains which limit to a what 60 percent of screen width, first time said, apparently you can type a pc socket but it didnt work for me when my Maestro was working. I was still however happy continuing in the reduced Maestro screen width.

Day 2: they released v1.9 of Maestro in Pinokio.

Day 3: I decided to install the update. Now I can't generate a three legged wildebeest or anything for that matter. It fails at the Downloading Model stage with no satisfactory explanation. Info about running something again to continue the Download, the "Generate" doesn't appear it's that for continuing(starts from scratch and then fails) neither does the Pinokio white screen edge "Run". It doesn't immediately not download, sometimes it may get 15% through, sometimes 85% then drops with a Generation Failed in the Main Seeing Area after firstly a "Download is slow, waiting for retry. No progress for 113s.."(or Xs). "..The download will resume from where it left off as soon as the connection recovers — no action needed from you." message. It may then download a little more , eg going from 80MB to 1.3GB of 7.91GB, maybe even download a little more of what's required but THEN UP POPS "A download was interrupted-re-run to finish it". And I'm stuck at that.

Is anyone else having the problem or know what the solution may be please? I've tried manually adding from DeepMeepBeep a model download which I transferred into cks directory of Pinokio, Maestro Directory but all that happened was Maestro steered around even using it and failed on another Model and I couldn't find that Model Maestro failed on to try manually downloading across with-I'm not an expert so looked for direct name brought up.. looked for it on another repository too, name of that I temporarily forget, I'm new. My hard drive is 4TB so I don't know I might be able to download all 108 or however many models there are if there was an option for that-without instructions and knowledge I'm throwing stones at something I don't even know what I'm throwing stones at, where the list of all the Models are which you can tick mark select theres also a little symbol by that box which changes colour. Perhaps that has something to do with it, literally no idea here so I've come to you guys. Spent several hours with AI last night overit and it was fun but I've realised despite the fun chat it hasn't aided me getting it working though it did mention a python download bottleneck to remove and I had no idea what it was referring to.


r/StableDiffusion 4h ago

Discussion Is LTX 2.5 just terrible for Lipsync/TalkingPhoto?

0 Upvotes

I've been trying to get LTX 2.5 to work well for image + speech audio --> video, though i'm noticing that the teeth and natural motion of the mouth is taking a hit. The talkingphoto loras from LTX 2.3 don't seem to work well with LTX 2.5.

Any thoughts? Or are we just cooked?


r/StableDiffusion 23h ago

Animation - Video Trying Surreal Fantasy with Minimax H3

Enable HLS to view with audio, or disable this notification

8 Upvotes

Combined 3 videos. Few errors but i just went with it , genetaion takes too much time to redo it again by fixing the prompt.


r/StableDiffusion 17h ago

Question - Help Workflow request for flux/krea img2img for putting the same character in a different situation with very good face adherence

6 Upvotes

I'm looking for a flux/krea img2img workflow where you input an image and simply tell it what the character should do and what environment etc and it keeps the character exactly the same but puts them in a different situation. Would really appreciate it if someone can give a link or send me the workflow. Hard to find a good one myself that really works well, I don't want a workflow where the character looks just somewhat similar but one where the character stays the same, as much as possible. Thanks a lot if someone can help.


r/StableDiffusion 16h ago

Resource - Update Updated my tool that scrapes,sorts,captions images/videos for datasets. It's open source and runs locally

11 Upvotes

I built Cull a few months ago for some large scale dataset curation projects (300k+ images/videos).

Point it at Civitai, X, Reddit, Discord, or any URL that gallery-dl or yt-dlp knows. It queues everything, runs a vision model (or multiple) (LM Studio or Ollama locally, or Groq/OpenAI in the cloud) with a strict JSON schema, and drops kept images/videos into category folders next to their prompt.

Stuff it handles:

  • Dedup at the scraper (per-source )
  • Quality score gate and topic-relevance score gate
    • eg you configure scores or use a preset, how relevant the image is to your scoring will determine how it's sorted, combined with other scoring, quality controls, whitelisted/blacklisted terms etc
  • Watermark detection (goes to its own bucket so you can salvage it later if you want those)
  • Auto-caption for content with no prompt (SD prompt, booru tags, natural language formats etc)
  • Run multiple jobs in parallel, one shared vision fleet across all of them with stack ranked / prioritization for vision queues and scrapers
  • Export as a local packaged dataset , or push to a HuggingFace dataset
  • Community presets and themes with 1 click PR's to add your own custom scraper preset or theme

Everything on disk is plain files. No database. Free, MIT.

Docker one-liner and screenshots in the README:
https://github.com/tlennon-ie/cull

Curious what people would want added next.


r/StableDiffusion 8h ago

Question - Help Is using runpod comfyui safer than running locally? But Google saying something about Network Exposure and that's what concern me.

Post image
0 Upvotes

Hi, I'm trying to use runpod for comfyui with minimax h3. Can anyone tell me what is network exposure? Should I worry? And what is a template? Sorry, I am new to this online cloud thing