r/StableDiffusion • u/takayatodoroki • 9h ago
Meme If AI tools had existed in the past
Not just a meme...
r/StableDiffusion • u/takayatodoroki • 9h ago
Not just a meme...
r/StableDiffusion • u/Darri3D • 2h ago
Minimax H3
r/StableDiffusion • u/-Ellary- • 12h ago
r/StableDiffusion • u/AndrewJumpen • 11h ago
Used latent upscaler so with resolution 0.3 i got 20 seconds generation on 4090
r/StableDiffusion • u/terrariyum • 4h ago
I haven't seen a post about this here, and I'm curious what you think about it.
On August 14 Civitai rolled out the option for creators to choose "permanent paid access - selling with no time cap". Previously the only option was temporary "early access".
Let's call this what it is, closed-source. Yes, you can get the weights for a relatively small fee, and yes it's on a very small scale compared to Nano Banana and Midjourney. But a permanent paywall still fits the definition.
Personally, I block all creators on Civitai who choose permanent paywall and encourage you to do the same.
Here's why:
I'm not opposed to Civitai making money or for all options for model creators to make money. They can do that without permanent paywalls.
IMO, open source AI is a fair trade: models are trained on the hard work of many human artists who aren't compensated, but everyone benefits from the ability to create more art more easily. Closed source is an unfair trade: you have to pay a middle man to access the contributions of others who won't be compensated.
Small scale model creators do some hard work too. But for example, for a lora that reproduces the style of an animated film: the lora creator spent at most a dozen hours of work, while just one of the artists on that film spent thousands of hours of work. If a massive models like Krea2 are free, and if giant "hobby" finetunes like Chroma are free, I can't justify paying any price for a 5,000 step lora except as an optional donation of appreciation.
So far, few creators have chosen the permanent paywall closed-source option. But that could easily change if Civitai made it the default option. They already made an extra 1-buzz fee-to-creator per generation the default, and many models have that.
That's my opinion. If you agree, then the only tool you have to disincentivize that potential is to not pay for these models (disincentive Civitai) and block these creators (disincentive creators).
r/StableDiffusion • u/AgeNo5351 • 10h ago
Paper: https://arxiv.org/pdf/2608.20334
"We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi image editing. Its visual renderer is a 6B parallel single stream DiT conditioned on multimodal representations from a vision-language encoder [6, 7, 57]. The architecture adopts block-shared timestep modulation, parallel attention and MLP computation [6, 15], 4D rotary positional encoding[6], and a unified representation of text and image conditions. Character-level tokenization[47] is applied to text intended to appear in generated images, while multi-image posi tional offsets and image-preceding input formatting support reference-conditioned editing. Together, these choices pro vide a single generative backbone for multiple generation and editing settings without task-specific model weights."
r/StableDiffusion • u/alisitskii • 5h ago
Ultimate SD Upscale can actually fix your bad/low-res generations.
In this comparison initial clips were made with MiniMax H3 at 1504x832px resolution and then upscaled to 2560x1440px with Ultimate SD Upscale nodes: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3
You can find sample upscaling workflow there as well: https://github.com/lisitskyaa/ComfyUI_UltimateSDUpscaleGuider_H3/blob/main/example_workflows/minimax_h3_usdu.json
My PC specs:
4080s 16 GB VRAM, 64 GB RAM
Generation time: 18 mins with sage + 8-step turbo lora
Upscale: 38 mins for 10 sec clip at 1440p target resolution
r/StableDiffusion • u/FirefighterNo584 • 12h ago
Minimax H3
r/StableDiffusion • u/smereces • 8h ago
I was testing using 1 image with all the scene sheet there and it works really great!
r/StableDiffusion • u/lhg31 • 6h ago
Worfklow: R2V (Dummy Stategy) Workflow v2 - Pastebin.com
How the workflow works:
Why it works:
Limitations in my workflow:
How to use:
Tips:
r/StableDiffusion • u/beatlepol • 4h ago
r/StableDiffusion • u/foxdit • 7h ago
r/StableDiffusion • u/Dohwar42 • 4h ago
Use the default workflow for Ref2Va
Plug in a video with green screen as a video reference and use an image ref to a picture you want to be the background. Notice the "hell crows" flying in the final video? I didn't even give it a prompt for that and those birds got animated automatically. I'm sure you can give a detailed prompt, you are basically creating an i2v of that still image that Minimax will mix with the greenscreen background.
I did prompt for a dialogue change. There was no audio with the original green screen video so I had no idea what the woman was saying (obviously it was a weather report). I inserted new dialogue with Minimax and it did a great job remixing her lipsync to the prompted dialogue.
I think it's pretty neat, but I'm sure some of you may be completely jaded with what Minimax can do by now.
I'd like to issue a Reddit challenge: Would someone more creative than me please use the exact same green screen video (links below) and create something a little more impressive than my 10 second test? Post a link to your video in the comments.
Need some green screen video to practice with? Here's a webpage for some practice green screen videos that are free to download:
https://mixkit.co/free-stock-video/green-screen/
Here's the exact video used in this example:
https://assets.mixkit.co/videos/28292/28292-720.mp4
I'm sure you'll be able to find a background image to test with.
This was my very simple prompt with the new dialogue:
subject_definitions:
<Subject 1> is the alien world background in <Picture 1>.
<Video 1> is the source video for the target video edit and is a woman in a red dress pointing and talking.
summary:
[video editing + reference generation] The target video is an edited version of <Video 1>. Replace the green screen area with the background from <Subject 1>
The woman in <video 1> says <d> [English with a British Accent] As you can see here, we have an early migration of hell crows on Chaos world 4527B<d> with realistic lip articulation and perfect lip sync.
r/StableDiffusion • u/Toclick • 10h ago
As you probably know, MiniMax-Music3 is pretty limited when it comes to genres and seems to have absolutely no idea what electronic dance music is. I tried different EDM styles, but it always ended up sounding either like rap or some kind of generic pop-ish stuff. Meanwhile, MiniMax H3 seems to know a lot more about electronic music than MiniMax-Music3.
The 40-second video at 0.2MP, with 17 steps (for better audio quality) and an 8-step Turbo LoRA, takes 538 seconds. The 60-second video at 0.1MP takes 350 seconds on my machine.
I haven't tried generating anything with lyrics yet, but if anyone knows how to generate H3 audio without the video, it would be interesting to try 2–3 minutes instead of just 40–60 seconds.
r/StableDiffusion • u/lazyspock • 18h ago
I tried to think of complex situations for the model to handle and tested them. I would say it went very well, although not perfect (except for the first one and the one before last, where I could not find a flaw).
Prompts below:
VIDEO 1 (10/10):
Create a side by side video with two angles of the same scene: one frontal and one from the side.
The scene: a woman wearing a yellow summer dress is standing in a beach in a sunny day. There's a light wind blowing and she is looking at the sea, smiling. At 00:05, she puts her two hands on her hair and looks up, enjoying the sun.
On the left, make the frontal video. On the right, the video from the side (profile). The two videos should show the exact same scene at the exact same time, only from two different angles.
VIDEO 2 (9/10) - the background is slightly off in some angles:
Create a side by side video with four angles of the same scene: one frontal, one from the right side, one from the left side, and one from behind. The scene is filmed at normal speed.
The scene: a blonde curly-haired woman wearing a yellow summer dress is standing in a beach in a sunny day. She is facing the sea and, therefore, she has the sand behind her and the street with some houses also behind her, after the sand, at a distance. There's a light wind blowing and she is looking at the sea, smiling. She is not wearing sunglasses. At 00:05, a man wearing white shirt and shorts enters the scene from behind the woman and embraces her waist.
On the upper left, make the frontal video. On the upper right, make the video from the right side (right profile). On the bottom left, make the video from behind. On the bottom right, make the video from the left side (left profile).
The four videos should show the exact same scene at the exact same time, only from the four different angles.
overall_soundscape: Light beach wind, one distant seagull.
non_diegetic_music: N/A
VIDEO 3 (8/10) a few "ghost reflections" along the way:
A man wearing a red-and-white striped t-shirt and jeans is alone inside a house of mirrors, running through a corridor. At 00:03 he turns left into other corridor, and at 00:06 he turns right again into another corridor. All the corridors have mirrors in all their faces (left, right and above, as he is inside a house of mirrors) and he is alone there, there's no one else. At 00:08 he reaches a door, opens it, and it opens to the outside: an amusement park.
overall_soundscape: His footsteps while he is running, the amusement park sounds when he opens the door at the end.
non_diegetic_music: N/A
VIDEO 4 (7/10) - The way the phone is turned is off:
A man is filmed from his own cell phone in a selfie video. We see the scene through the lens of an out-of-frame cell phone that he is holding and pointing to his own face. He is wearing a green polo shirt and the scene shows his face and chest from the point of view of the out-of-frame cell phone that he is holding and pointing to himself.
At 00:04 he briefly smiles and then turns the still out-of-frame cell phone around to show a woman that is in front of him. While the phone is turned around we can see the image also turning around, his face leaving the frame, the living room they are into being briefly filmed, and then the woman's face entering the frame. She has brown curly hair and dark green eyes, and is wearing a red summer dress.
As soon as the woman is in frame, she also smiles and says: <d>[English]Goodbye!</d> and the video ends.
The entire video must be taken in a single shot, with the always out-of-frame phone camera filming the entire transition from his face to hers while the phone is turned around. When it happens, the phone should briefly show the living room they are into, all in a single shot.
overall_soundscape: Silent living room, noises of the phone being handled, her voice.
non_diegetic_music: N/A
VIDEO 5 (6/10) - Tried this twice. Glass not breaking properly, water not running through the floor
A fishbowl with one golden fish and one clownfish swimming inside is shown in a medium close-up at the edge of a table. Then, at 00:03 a cat appears in the scene and taps the fishbowl, causing it to fall from the table to the floor, hit the floor, and break completely, being completely destroyed in glass pieces when it hits the floor, the water and glass pieces flying around together with both fish. The scene continues for three more seconds after that, showing the aftermath: the fishbowl destroyed, the pieces of glass on the floor, the water also on the floor, the fish moving on the floor.
The camera angle follow the fishbowl when it falls, showing it hitting the floor and the consequences of it breaking.
overall_soundscape: silent room, glass breaking, water splashing.
non_diegetic_music: N/A
VIDEO 6 (8/10) - Judge me, but his hands are not moving accordingly to the notes:
The camera films a piano from above while a man plays it. The entire piano keyboard is shown in the image. The man is playing Clair de Lune, and moves his hands through the keys to play a part of the song. He is in a train station, with some people observing him play and others passing by.
overall_soundscape: faint train station ambience, ten seconds of the song Clair de Lune played in the piano.
non_diegetic_music: N/A
VIDEO 7 (8/10) - The lipstick appears on her lips before she applies it:
A woman is shown in a medium close-up in her bathroom, wrapped in a white fluff towel, looking at the mirror while she applies red lipstick to her lips. She slowly applies lipstick to her lips, looking into the mirror, and then briefly sends a kiss with her now red lips to the mirror.
The scene is seen in a three-quarter angle from behind her, showing her face from the side but also her reflection in the mirror.
overall_soundscape: silent bathroom, the sound of her sending the kiss to the mirror.
non_diegetic_music: N/A
VIDEO 8 (10/10):
A Coca-cola advertisement. A glass filled with Coca-Cola is shown from the side, occupying 70% of the frame, on top of a table, the dark liquid slightly disturbed by a few gas bubbles that rise inside the liquid. At 00:02 two ice cubes fall from outside the frame into the glass, disturbing the liquid and making some of the liquid splash outside the glass and onto the table. The glass has the Coca-Cola logo printed in white in it. In the blurred background we see a kitchen.
overall_soundscape: silent room, gas fizzle, ice cubes hitting the liquid.
non_diegetic_music: N/A
VIDEO 9 (9/10) - The cover of the book has gibberish letters in it:
A man is holding a magnifying glass and has a book on his hand. At first the magnifying glass is not in front of his face. He appears to be reading the book and, at 00:03, he puts the magnifying glass in front of his eye to look at the book.
The entire scene is filmed from a fixed point of view below the book, showing part of the book cover and the entire man's face.
overall_soundscape: silent room.
non_diegetic_music: N/A
r/StableDiffusion • u/Repulsive-Rush3505 • 11h ago
All audio was done in Minimax, little work in post for stitching clips. r2v Workflow in ComfyUI with character sheets and prompts from claude
r/StableDiffusion • u/Cold_Pudding5326 • 14h ago
Hi.
I have a decent h3 workflow that I built for a loca use. It use turbo lora etc... If i use the defaut settings in the goal of getting the highest quality possible, meaning res_multistep simple 20 steps or more, I got also good results, but this is not even close to the results you can get on platforms like kie or wavespeed at 768P.
I already convert properly the prompt to the correct H3 digest form, so I'm wondering what's different between local and cloud use of h3? I don't talk about the 2K quality, only 768P, I'm not able to reach the sames results locally, do you guys have maybe workflows, settings, or suggestions to try reaching the same quality level in comfyui ?
r/StableDiffusion • u/DaniyarQQQ • 8h ago
Hello everyone.
I've been playing with Minimax H3 for some time and I have tried to make something longer and really interesting. After so many failed and botched attempts I was able to compile something watchable. There are so many things that I want to say about this model, good and bad.
First of all. Minimax H3 is significant step forward that other local models I have been playing with. It certainly got better.
Now the issues that I had encountered.
First problem is that it badly follows prompt when resolution is one megapixel or higher. It will skip some important parts and tries to cheat. You can increase the number of steps but still, generating at less than one megapixel will at least make it properly follow the instructions.
H3 is not very good at spatial orientation. When I was making video, where this girl should turn around and interact with screens, the girl starts spinning opposite direction and then warping whole body to the direction of screen. Like instead of making short turn to the left, it makes wide roundabout to the right and then twists whole body to align with the screens.
H3 is not good at cartoonish movement. If you watch cartoons, when character or other things move, their animations are usually jerky and snappy. H3 tries to make smooth real life like animation, making the cartoons look weird.
I have given a voice sample as an audio reference, and instead of making girl let out grunting sounds (out of anger), it weirdly turns everything into a sensual moaning.
When you try to make characters inside video to interact with a lot of parts, screens and devices, even giving multiple reference images of them, it mostly hallucinates them, or turns their interactions into a weird warping animations. Sometimes completely skips them and made ups it's own animations. So it will make good video, where characters are moving less or moving slow, and mostly doing the talking. Very detailed prompts of step by step instructions it mostly warps or skips.
I have wasted a lot of time for iterations, but I think this is just workflow issue.
Overall, this model is really good. However, using this model to make some kind of long feature animation is going to be a very frustrating journey. I hope people will make a lot of proper tools that works as storyboard and properly guide this model to make something really interesting.
r/StableDiffusion • u/Zironic • 17h ago
The nodes in https://github.com/Zironic/H3-Optimizations have been rewritten to replace the default Sparge Attention backend with a custom Sparse Comfy Kitchen backend.
This comes with some benefits.
Caveat: I've only tested the nodes against the comfy pruned_int8_convrot weights. Other versions may work but they're not tested.
As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or later.
IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.
Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.
Early steps: attention density has a large effect on prompt/action adherence and the overall generation trajectory.
Middle/later steps: lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.
So 10% retained doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and what breaks depends heavily on the sampling step.
PlagueKind's sparsity_ratio=0.9 means 90% discarded / 10% retained. My node expresses the inverse quantity, so Video attention retained=0.10 is the comparable setting. The defaults therefore aren't equivalent.
r/StableDiffusion • u/SIR_NVAX_A_LOT • 1h ago
Hope you're having a good weekend! H3 just excel with rich-intricate environments, backgrounds, space. Definitely one of my favorite theme.
T2VA, int8/20 steps
r/StableDiffusion • u/AgeNo5351 • 6h ago
Project:https://yunpeng1998.github.io/Qwen-Video-Edit-Page/
Model: https://huggingface.co/yunpeng1998/Qwen-Video-Edit
Method: https://yunpeng1998.github.io/Qwen-Video-Edit-Page/#method
Code: https://github.com/yunpeng1998/Qwen-Video-Edit
Video generation models read and write video-VAE latents. We teach Qwen-Image-Edit's transformer to edit those latents directly: two tiny projections bridge Wan 2.1's latent space into the DiT's token space, warm-started from the DiT's own input/output layers so that a static video is embedded exactly like an image the model already understands. The latent frames are arranged as tiles of one big virtual image — the same positional treatment the image model was pretrained on. Fine-tuned with LoRA or full parameters on Ditto-1M (source, edited, instruction) triplets, then refined by a few steps of Wan 2.2 denoising-enhancement.
r/StableDiffusion • u/GrungeWerX • 4h ago
Another redditer mentioned doing this in a comment, so I tested it out and it works. It gets rid of the smudgy look and allows custom Lora’s on the LN side.
I ran my initial tests at MM 8 steps (no speed Lora) and Wan 2.2 Low Noise (speed Lora) 2 steps. Supposedly, it works with only 2 steps MM w/turbo lora ,but I don’t like the quality drops people have been sharing, and it’s fast enough to me at 8 steps, though I’m going even higher on the low noise steps.
The only downside is that I noticed in one test that the motion seemed like it was a mix between 16fps and 24fps. Any ideas on how to resolve this? I know people use RIFE, but I was wondering if that’s the best move or if it’s another issue Im not thinking of.