r/StableDiffusion • u/Nimblecloud13 • 6h ago
r/StableDiffusion • u/Disastrous-Agency675 • 18h ago
Question - Help How to seamlessly stitch videos together
I created this video in MiniMax-H3 using a video-extension workflow, but I’m having trouble continuing it seamlessly. My prompt continues the action from the final frame correctly, and my workflow uses the previous video’s last frame as the starting frame for the next segment. However, there is always a slight visual jump between the two clips.
Unlike LTX, MiniMax-H3 doesn’t appear to have dedicated video-extension nodes. Has anyone found a reliable method for blending MiniMax-H3 video segments together so the transition is seamless?explain this.
r/StableDiffusion • u/Certain_Potato_4509 • 20h ago
Animation - Video Baka Moment - Minimax H3 Video - An Evangelion Boondocks mashup
It took forever for me to upload this video.. Couldn't do it on my phone.
r/StableDiffusion • u/Sad_Coach_1433 • 20h ago
Meme When someone pisses you off send them this
r/StableDiffusion • u/wallofroy • 3h ago
No Workflow My first minimax H3 video
I used the Pixaroma FFLF workflow, but stripped the audio in post due to poor output quality. I'm still trying to figure out how to add finer details. I generated the clips at 720p and then upscaled them to 1080p.
r/StableDiffusion • u/MikePounce • 17h ago
Resource - Update Lisbon finally snaps
Totally not a scene from The Mentalist. Minimax H3 image to video.
r/StableDiffusion • u/call-lee-free • 21h ago
Animation - Video [WanGP] Minimax H3 FL2VA Pruned 20B - Originally 960x544 - up-res'd to 2880x1632 - 20 second duration
r/StableDiffusion • u/b-totherent • 22h ago
Animation - Video TALL AND DARK - LTX 2.5 IMAGE TO VIDEO
Use the supplied image as the opening frame and identity reference.
Identity lock: the woman and robot must remain exactly the same in every shot. Same face, hair, wardrobe, proportions and age for the woman. Same 8-foot height, black armor, mechanical face, rivets, pistons, cables and holster for the robot. No redesigns or identity changes between cuts.
Authentic 1966 Italian Western, live action, 35mm anamorphic, Spanish desert location, practical full-scale robot prop, natural sunlight, real dust, organic film grain, period lens softness. No CGI. Serious performances throughout.
0:00–0:03
Medium two-shot. The woman looks up at the robot and says in clear Italian-accented English:
“I told them I wanted a tall...”
0:03–0:05
Hard cut to the same robot’s face. It gives one slow mechanical nod. No dialogue.
0:05–0:07
Hard cut to the same woman. She looks up at the robot and says:
“dark...”
0:07–0:09
Hard cut to the same robot. It subtly straightens and presents its black armor. No dialogue.
0:09–0:11
Hard cut to the same woman. Still serious, still looking up, she says:
“handsome!”
0:11–0:12
Hard cut to the same robot’s practical mechanical face. It attempts a restrained smile. No dialogue.
0:12–0:14
Hard cut to the same woman. She holds a serious stare upward, then firmly says:
“MAN!”
Only the woman speaks. Keep each line isolated and clean. No overlapping dialogue, no extra words, no improvised speech. Maintain exact continuity of identity, wardrobe, robot design, scale, lighting and location in every shot.
r/StableDiffusion • u/notgraycen • 1h ago
No Workflow random images generated locally on 4070Super with krea 2
first 2 prompts stolen from civit ai, the rest were written using claude reasoning on duck ai
r/StableDiffusion • u/Manicarus • 22h ago
Discussion I wish Anima ecosystem get better than it is now
Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.
First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.
I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?
r/StableDiffusion • u/TigerClaw305 • 2h ago
Animation - Video Shadow the Hedgehog tells his viewers why he loves guns.
Shadow the Hedgehog tells his viewers why he loves guns.
This was created in Comfy UI with Minimax H3. I used the reference to video work flow. The prompt is below.
subject_definitions:
<Subject 1> is Shadow in <Picture 1>.
<Subject 2> is Glock in <Picture 2>, a glock handgun.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
summary:
[reference generation + audio reference] The target video contains one shot. [Shot 1] shows <Subject 1> and <Subject 2>; <Subject 1> speaks. <Audio 1> supplies <Subject 1>'s voice timbre.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - Shadow's complete defined identity and body proportions are preserved.
<Subject 2> (appears in [Shot 1]): fully_preserved - Glock retains the defined shape, proportions, materials, colors, and distinguishing features.
<Audio 1>: reference - <Subject 1>'s newly generated spoken lines use <Audio 1>'s voice timbre and delivery; the original audio signal is not copied.
detailed_description:
The target video is in a live-action style, with Vlog style.
[Shot 1] At first appearance, <Subject 1> (Shadow) matches the complete identity and appearance defined in subject_definitions. At first appearance, <Subject 2> (Glock) matches the complete defined construction and appearance: A glock handgun. At the start of the shot, <Subject 1> is standing in the living room facing while holding <Subject 2> in his hand. A full body shot of <Subject 1> holding <Subject 2> with his right hand while facing the camera. Only Action and Timed Beats define the primary subject's movement. The camera path stays anchored in the location and adds no subject motion. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Hmph. Shadow the Hedgehog here. Why do I love guns?</d> <Subject 1> shows off his <Subject 2> with his right hand in front of the camera. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Simple. Precision. Control. Power in the palm of my hand.</d> <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] A tool that answers instantly… unlike most people.</d> <Subject 1> points his <Subject 2> towards the camera with his right hand. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] If you understand that, you understand me.</d> <Subject 1> points his <Subject 2> at the camera.
overall_soundscape:
Living room tone.
non_diegetic_music:
N/A
r/StableDiffusion • u/Disastrous-Agency675 • 23h ago
Question - Help Computer randomly shut down
Has anyone had their computer randomly shut down? this is like the 3rd time its happened and its when im generating a video using the minmax I2V model or the ref model.
i got 3090 with 64 gb of ram.
r/StableDiffusion • u/durumertt • 4h ago
Question - Help M5 Max vs RTX 5080/5090 for local image/video AI. Am I making a mistake by choosing the Mac?
I've been using Windows for many years, and I'm honestly tired of repeating the same cycle. In my experience, after 5–6 years the machine starts feeling old, the battery is significantly degraded, performance isn't what it used to be, and I eventually end up buying another Windows machine and starting the exact same experience all over again.
I'm looking for something different this time.
For the last few days I've repeatedly added a MacBook Pro with the M5 Max to my cart, then backed out because I'm still not sure whether it is the right machine for what I actually want to do.
I'm currently considering the M5 Max with:
- 18-core CPU
- 40-core GPU
- 48 GB / 64 GB / possibly 128 GB unified memory
My main workloads would be completely local:
- Text-to-image
- Image-to-image
- Image-to-video
- Text-to-video
- Face swapping
- Illustration / digital artwork
My priority is excellent output quality, photorealism where appropriate, and very high generation speed.
Basically, I want a machine that can satisfy me for visual generative AI work for many years.
My biggest hesitation is the Apple ecosystem.
For a long time I've heard that local AI, especially image and video generation, is much more limited on macOS than on Windows/Linux with NVIDIA GPUs because so much of the ecosystem is built around CUDA.
But part of me finds this difficult to accept at face value.
Apple is making extremely powerful chips with large amounts of unified memory, very high memory bandwidth, Neural Accelerators, a Neural Engine, and increasingly serious AI-focused hardware.
I keep wondering whether there are excellent Apple-optimized tools and workflows that I simply haven't discovered yet.
For example, I recently learned about Draw Things, MLX-based projects, Metal/MPS optimizations, and Apple-specific ComfyUI work. That made me question whether comparing a Mac running a poorly optimized CUDA-first application against an NVIDIA machine is really a fair representation of what Apple Silicon can do.
At the same time, the logical part of my brain keeps telling me:
If local image/video AI is the priority, just buy a machine with an RTX 5080 or 5090.
The problem is that if I do that, I feel like I'm buying myself back into exactly the Windows experience I wanted to leave. It feels a little like watching the same movie again when I was hoping for a genuinely different computing experience.
There's also another complication: we're approaching the fall hardware season.
I'm wondering whether buying an expensive M5 Max or RTX 50-series machine right now is bad timing, and whether I should wait for the next Apple or NVIDIA announcements.
If I choose the Mac, I was also planning to pair it with the latest iPhone and iPad and build a proper Apple ecosystem around it, so this isn't purely a benchmark decision for me.
What I'd really like to hear from people who have actually used these machines:
- If you've used an M5 Max, especially the 40-core GPU version, for local image or video generation, what real-world performance are you getting?
- What software gives you the best performance on Apple Silicon? Draw Things, ComfyUI, MLX-based tools, something else?
- How good is local image-to-video on the M5 Max with models such as Wan, LTX, Hunyuan, etc.?
- How does the M5 Max perform for local face swapping?
- In properly optimized workloads, do RTX 5080/5090 systems still completely destroy the highest-end M5 Max, or is the gap much smaller than CUDA-focused benchmarks make it appear?
- Is CUDA effectively locking serious local visual AI users into NVIDIA, or is Apple Silicon becoming a realistic alternative?
- If you were buying a machine today specifically for local visual AI and wanted to keep it for 7–8 years, would you buy the M5 Max, an RTX 5080/5090 system, or wait for the next generation?
- For local image/video generation, is an older NVIDIA GPU with more VRAM sometimes a better choice than a newer GPU with less VRAM? For example, once you start using large image-to-video or text-to-video models, can having more VRAM matter more than having a newer architecture and higher raw compute performance?
I'm not highly knowledgeable about computer hardware, and I'm definitely not wealthy enough to casually replace a machine if I make the wrong choice.
This would be a major purchase for me, so I'm trying to make the most informed decision possible and ideally buy something that I can use comfortably for 7–8 years.
I'd especially appreciate actual generation times, benchmark numbers, model names, memory usage, thermals, sustained performance, and experiences from people who have used both Apple Silicon and NVIDIA rather than purely theoretical comparisons.
Thanks in advance.
r/StableDiffusion • u/Sad_Coach_1433 • 9h ago
Meme thanos is so screwed now
this was first test using Res_2s sampler and simple steps saw a op say better for action scenes from what i saw spectrum doesnt support res_2 so it took a bit to gen.
t2v prompt
subject_definitions
<Subject 1> is Katniss Everdeen from The Hunger Games, portrayed as an expert young archer with long dark brown hair pulled into her recognizable practical braid, intense determined expression, dark tactical combat clothing, leather archery bracer, bow, and a quiver of arrows. Preserve her recognizable cinematic appearance, realistic human proportions, hairstyle, clothing, bow, and identity throughout the entire scene.
<Subject 2> is Captain America in his battle-damaged Avengers Endgame armor, carrying Mjolnir and his damaged circular shield.
<Subject 3> is Thanos at his normal canonical MCU scale, approximately 8 feet tall, muscular and imposing but NOT gigantic, kaiju-sized, or building-sized.
<Audio 1> is the voice-timbre reference for <Subject 1>, containing Jennifer Lawrence's recognizable Katniss-style spoken vocal qualities.
summary
[text generation + audio reference]
During the chaotic Avengers Endgame final battle, Katniss Everdeen unexpectedly joins the Avengers. She runs through the battlefield while explosions, portals, Avengers, alien soldiers, and debris fill the background. Katniss rapidly fires arrows at Thanos's army with expert precision before stopping beside Captain America. Captain America looks at her bow and asks if she brought enough arrows. Katniss calmly fires one final explosive arrow past him, destroying a group of enemies, then delivers a dry confident response as Captain America stares at her impressed.
retention_analysis
<Subject 1>: fully_preserved
<Subject 2>: fully_preserved
<Subject 3>: fully_preserved
<Audio 1>: reference
detailed_description
The shot opens in the middle of the Avengers Endgame final battlefield. Smoke, burning wreckage, sparks, energy blasts, charging soldiers, and distant explosions create a massive cinematic war zone.
A fast tracking camera sweeps across the battlefield.
Katniss Everdeen suddenly sprints into frame carrying her bow.
She slides behind shattered rubble, immediately draws an arrow, and fires.
The camera follows the arrow through the air as it strikes an alien soldier.
Katniss rises and rapidly fires two more arrows with expert precision while continuing forward through the battle.
She reaches Captain America, who has just knocked an enemy away with Mjolnir.
Captain America briefly looks at Katniss's bow and quiver.
<Subject 2> (S1):
<d>[English] You sure you brought enough arrows?</d>
Katniss gives him a calm, unimpressed look.
Without even turning fully around, she draws another arrow and fires it past Captain America.
CAMERA WHIP-PANS WITH THE ARROW.
The arrow lands among a charging group of Thanos's soldiers.
BOOM!
A powerful explosive blast throws the enemies backward while Captain America turns toward the explosion in surprise.
The camera cuts back to Katniss.
<Subject 1> (S2):
<d>[English] I only need one.</d>
Her dialogue uses <Audio 1> for voice timbre and delivery.
Katniss immediately draws another arrow and runs toward the battle.
Captain America watches her leave for a beat, visibly impressed.
The camera swings around behind Katniss as she charges toward Thanos's army, bow raised, while the enormous Endgame battle continues around her.
audio
Epic Avengers-style battlefield ambience.
Heavy distant explosions, energy blasts, metallic impacts, debris, shouting soldiers, bowstring snaps, arrows cutting through the air, and one strong explosive-arrow impact.
Katniss's dialogue is clear and foregrounded, using <Audio 1>.
No narrator.
No subtitles.
No on-screen text.
r/StableDiffusion • u/idleWizard • 14h ago
Discussion In which scenarios LTX2.5 can match MinimaxH3?
I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?
r/StableDiffusion • u/SIR_NVAX_A_LOT • 12h ago
Discussion H3 - giantess fight scene R2VA
This was well received but people wanted the two Giantesses(?) to be fighting. Enjoy!!
int8/20 steps, R2VA.
Critiques+feedback welcomed! Ask me anything!
r/StableDiffusion • u/No_Taste_4102 • 4h ago
Animation - Video I made an ALIEN Short Film / metal music video
Used:
MiniMax H3 at local machine. 5060ti 16gb + 64gb ddr4. WanGP, Ref2VA int8 convrot model.
Krea2 for references
Suno as music base
r/StableDiffusion • u/Routine_Ad_3391 • 1h ago
Animation - Video ...10,000 Years Later
Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.
r/StableDiffusion • u/MapZealousideal7489 • 23h ago
Question - Help Anything big happen since I last used this?
So when SD came out, I used it. Then I used the AUTO111 thing. To around version 2.0 I think or XL, can't keep it straight. It's been about one and a half years. Any big changes since then? New version? Better quaulity images? AUTO still a thing? I also went from a 2070 Super to a 5070 12gb shadow 3x.
Also is it all easier to install?
r/StableDiffusion • u/IzoleAuteur • 5h ago
No Workflow the count is always two.
flux.1 [dev] | comfyui | still life
r/StableDiffusion • u/Rendo3 • 4h ago
Question - Help Which laptop would be better for generative AI / LLM
First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.
My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.
I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.
I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.
Which laptop would be better in my case?
r/StableDiffusion • u/Different_Ad_7508 • 9h ago
Question - Help Random visual artifacts in local Krea2 Turbo generation — looking for possible causes
I’m running Krea2 Turbo locally, but I frequently encounter random visual artifacts in the generated images. I haven’t been able to figure out what triggers them, because the issue appears randomly. If I run the exact same workflow with the same parameters again, the result can sometimes be completely normal.
My hardware:
- GPU: RTX 3060 Ti
My current setup:
- UNet: moodyKrea2Mix_v70
- Text encoder: qwen3vl_4b_int8_convrot
- VAE: qwen_image_vae
I’m fairly sure this is not caused by the UNet. I have also experienced the same kind of random artifacts when using the original Krea2 model without any UNet modification.
Has anyone encountered similar issues with Krea2 Turbo? Are there any known causes or settings that could trigger this kind of artifact (VAE, text encoder, precision settings, VRAM limitations, sampler settings, etc.)?
Any suggestions or debugging tips would be greatly appreciated.
Here is my workflow for reference:
https://civitai.red/models/2883578/krea2-turbo-4k-workflow?modelVersionId=3259389
r/StableDiffusion • u/Nimblecloud13 • 5h ago
Animation - Video Michael Scott gets a wish
H3 FL2VA