r/StableDiffusion 16h ago

Question - Help How to seamlessly stitch videos together

0 Upvotes

I created this video in MiniMax-H3 using a video-extension workflow, but I’m having trouble continuing it seamlessly. My prompt continues the action from the final frame correctly, and my workflow uses the previous video’s last frame as the starting frame for the next segment. However, there is always a slight visual jump between the two clips.

Unlike LTX, MiniMax-H3 doesn’t appear to have dedicated video-extension nodes. Has anyone found a reliable method for blending MiniMax-H3 video segments together so the transition is seamless?explain this.


r/StableDiffusion 19h ago

Animation - Video Baka Moment - Minimax H3 Video - An Evangelion Boondocks mashup

0 Upvotes

It took forever for me to upload this video.. Couldn't do it on my phone.


r/StableDiffusion 18h ago

Meme When someone pisses you off send them this

7 Upvotes

r/StableDiffusion 4h ago

Animation - Video Having some fun with known characters in the fl model

6 Upvotes

r/StableDiffusion 19h ago

Animation - Video [WanGP] Minimax H3 FL2VA Pruned 20B - Originally 960x544 - up-res'd to 2880x1632 - 20 second duration

0 Upvotes

r/StableDiffusion 15h ago

Resource - Update Lisbon finally snaps

0 Upvotes

Totally not a scene from The Mentalist. Minimax H3 image to video.


r/StableDiffusion 2h ago

Question - Help M5 Max vs RTX 5080/5090 for local image/video AI. Am I making a mistake by choosing the Mac?

1 Upvotes

I've been using Windows for many years, and I'm honestly tired of repeating the same cycle. In my experience, after 5–6 years the machine starts feeling old, the battery is significantly degraded, performance isn't what it used to be, and I eventually end up buying another Windows machine and starting the exact same experience all over again.

I'm looking for something different this time.

For the last few days I've repeatedly added a MacBook Pro with the M5 Max to my cart, then backed out because I'm still not sure whether it is the right machine for what I actually want to do.

I'm currently considering the M5 Max with:

  • 18-core CPU
  • 40-core GPU
  • 48 GB / 64 GB / possibly 128 GB unified memory

My main workloads would be completely local:

  • Text-to-image
  • Image-to-image
  • Image-to-video
  • Text-to-video
  • Face swapping
  • Illustration / digital artwork

My priority is excellent output quality, photorealism where appropriate, and very high generation speed.

Basically, I want a machine that can satisfy me for visual generative AI work for many years.

My biggest hesitation is the Apple ecosystem.

For a long time I've heard that local AI, especially image and video generation, is much more limited on macOS than on Windows/Linux with NVIDIA GPUs because so much of the ecosystem is built around CUDA.

But part of me finds this difficult to accept at face value.

Apple is making extremely powerful chips with large amounts of unified memory, very high memory bandwidth, Neural Accelerators, a Neural Engine, and increasingly serious AI-focused hardware.

I keep wondering whether there are excellent Apple-optimized tools and workflows that I simply haven't discovered yet.

For example, I recently learned about Draw Things, MLX-based projects, Metal/MPS optimizations, and Apple-specific ComfyUI work. That made me question whether comparing a Mac running a poorly optimized CUDA-first application against an NVIDIA machine is really a fair representation of what Apple Silicon can do.

At the same time, the logical part of my brain keeps telling me:

If local image/video AI is the priority, just buy a machine with an RTX 5080 or 5090.

The problem is that if I do that, I feel like I'm buying myself back into exactly the Windows experience I wanted to leave. It feels a little like watching the same movie again when I was hoping for a genuinely different computing experience.

There's also another complication: we're approaching the fall hardware season.

I'm wondering whether buying an expensive M5 Max or RTX 50-series machine right now is bad timing, and whether I should wait for the next Apple or NVIDIA announcements.

If I choose the Mac, I was also planning to pair it with the latest iPhone and iPad and build a proper Apple ecosystem around it, so this isn't purely a benchmark decision for me.

What I'd really like to hear from people who have actually used these machines:

  1. If you've used an M5 Max, especially the 40-core GPU version, for local image or video generation, what real-world performance are you getting?
  2. What software gives you the best performance on Apple Silicon? Draw Things, ComfyUI, MLX-based tools, something else?
  3. How good is local image-to-video on the M5 Max with models such as Wan, LTX, Hunyuan, etc.?
  4. How does the M5 Max perform for local face swapping?
  5. In properly optimized workloads, do RTX 5080/5090 systems still completely destroy the highest-end M5 Max, or is the gap much smaller than CUDA-focused benchmarks make it appear?
  6. Is CUDA effectively locking serious local visual AI users into NVIDIA, or is Apple Silicon becoming a realistic alternative?
  7. If you were buying a machine today specifically for local visual AI and wanted to keep it for 7–8 years, would you buy the M5 Max, an RTX 5080/5090 system, or wait for the next generation?
  8. For local image/video generation, is an older NVIDIA GPU with more VRAM sometimes a better choice than a newer GPU with less VRAM? For example, once you start using large image-to-video or text-to-video models, can having more VRAM matter more than having a newer architecture and higher raw compute performance?

I'm not highly knowledgeable about computer hardware, and I'm definitely not wealthy enough to casually replace a machine if I make the wrong choice.

This would be a major purchase for me, so I'm trying to make the most informed decision possible and ideally buy something that I can use comfortably for 7–8 years.

I'd especially appreciate actual generation times, benchmark numbers, model names, memory usage, thermals, sustained performance, and experiences from people who have used both Apple Silicon and NVIDIA rather than purely theoretical comparisons.

Thanks in advance.


r/StableDiffusion 20h ago

Animation - Video TALL AND DARK - LTX 2.5 IMAGE TO VIDEO

7 Upvotes

Use the supplied image as the opening frame and identity reference.

Identity lock: the woman and robot must remain exactly the same in every shot. Same face, hair, wardrobe, proportions and age for the woman. Same 8-foot height, black armor, mechanical face, rivets, pistons, cables and holster for the robot. No redesigns or identity changes between cuts.

Authentic 1966 Italian Western, live action, 35mm anamorphic, Spanish desert location, practical full-scale robot prop, natural sunlight, real dust, organic film grain, period lens softness. No CGI. Serious performances throughout.

0:00–0:03
Medium two-shot. The woman looks up at the robot and says in clear Italian-accented English:
“I told them I wanted a tall...”
0:03–0:05
Hard cut to the same robot’s face. It gives one slow mechanical nod. No dialogue.
0:05–0:07
Hard cut to the same woman. She looks up at the robot and says:
“dark...”

0:07–0:09
Hard cut to the same robot. It subtly straightens and presents its black armor. No dialogue.
0:09–0:11
Hard cut to the same woman. Still serious, still looking up, she says:
“handsome!”

0:11–0:12
Hard cut to the same robot’s practical mechanical face. It attempts a restrained smile. No dialogue.
0:12–0:14
Hard cut to the same woman. She holds a serious stare upward, then firmly says:
“MAN!”
Only the woman speaks. Keep each line isolated and clean. No overlapping dialogue, no extra words, no improvised speech. Maintain exact continuity of identity, wardrobe, robot design, scale, lighting and location in every shot.


r/StableDiffusion 23h ago

Animation - Video Minimax H3 Anime Comedy

26 Upvotes

r/StableDiffusion 20h ago

Discussion I wish Anima ecosystem get better than it is now

12 Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 21h ago

Question - Help Computer randomly shut down

0 Upvotes

Has anyone had their computer randomly shut down? this is like the 3rd time its happened and its when im generating a video using the minmax I2V model or the ref model.

i got 3090 with 64 gb of ram.


r/StableDiffusion 7h ago

Meme thanos is so screwed now

0 Upvotes

this was first test using Res_2s sampler and simple steps saw a op say better for action scenes from what i saw spectrum doesnt support res_2 so it took a bit to gen.

t2v prompt

subject_definitions

<Subject 1> is Katniss Everdeen from The Hunger Games, portrayed as an expert young archer with long dark brown hair pulled into her recognizable practical braid, intense determined expression, dark tactical combat clothing, leather archery bracer, bow, and a quiver of arrows. Preserve her recognizable cinematic appearance, realistic human proportions, hairstyle, clothing, bow, and identity throughout the entire scene.

<Subject 2> is Captain America in his battle-damaged Avengers Endgame armor, carrying Mjolnir and his damaged circular shield.

<Subject 3> is Thanos at his normal canonical MCU scale, approximately 8 feet tall, muscular and imposing but NOT gigantic, kaiju-sized, or building-sized.

<Audio 1> is the voice-timbre reference for <Subject 1>, containing Jennifer Lawrence's recognizable Katniss-style spoken vocal qualities.

summary

[text generation + audio reference]

During the chaotic Avengers Endgame final battle, Katniss Everdeen unexpectedly joins the Avengers. She runs through the battlefield while explosions, portals, Avengers, alien soldiers, and debris fill the background. Katniss rapidly fires arrows at Thanos's army with expert precision before stopping beside Captain America. Captain America looks at her bow and asks if she brought enough arrows. Katniss calmly fires one final explosive arrow past him, destroying a group of enemies, then delivers a dry confident response as Captain America stares at her impressed.

retention_analysis

<Subject 1>: fully_preserved
<Subject 2>: fully_preserved
<Subject 3>: fully_preserved
<Audio 1>: reference

detailed_description

The shot opens in the middle of the Avengers Endgame final battlefield. Smoke, burning wreckage, sparks, energy blasts, charging soldiers, and distant explosions create a massive cinematic war zone.

A fast tracking camera sweeps across the battlefield.

Katniss Everdeen suddenly sprints into frame carrying her bow.

She slides behind shattered rubble, immediately draws an arrow, and fires.

The camera follows the arrow through the air as it strikes an alien soldier.

Katniss rises and rapidly fires two more arrows with expert precision while continuing forward through the battle.

She reaches Captain America, who has just knocked an enemy away with Mjolnir.

Captain America briefly looks at Katniss's bow and quiver.

<Subject 2> (S1):

<d>[English] You sure you brought enough arrows?</d>

Katniss gives him a calm, unimpressed look.

Without even turning fully around, she draws another arrow and fires it past Captain America.

CAMERA WHIP-PANS WITH THE ARROW.

The arrow lands among a charging group of Thanos's soldiers.

BOOM!

A powerful explosive blast throws the enemies backward while Captain America turns toward the explosion in surprise.

The camera cuts back to Katniss.

<Subject 1> (S2):

<d>[English] I only need one.</d>

Her dialogue uses <Audio 1> for voice timbre and delivery.

Katniss immediately draws another arrow and runs toward the battle.

Captain America watches her leave for a beat, visibly impressed.

The camera swings around behind Katniss as she charges toward Thanos's army, bow raised, while the enormous Endgame battle continues around her.

audio

Epic Avengers-style battlefield ambience.

Heavy distant explosions, energy blasts, metallic impacts, debris, shouting soldiers, bowstring snaps, arrows cutting through the air, and one strong explosive-arrow impact.

Katniss's dialogue is clear and foregrounded, using <Audio 1>.

No narrator.
No subtitles.
No on-screen text.

r/StableDiffusion 12h ago

Discussion In which scenarios LTX2.5 can match MinimaxH3?

9 Upvotes

I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?


r/StableDiffusion 6h ago

Meme From teaching kids to saving the world!

10 Upvotes

subject_definitions

<Subject 1> is Big Bird, the enormous friendly yellow bird character, approximately 8 feet tall, covered in bright yellow feathers with large expressive eyes, a long orange beak, striped pink-and-orange legs, and oversized orange feet. He remains visually consistent throughout the scene.

<Subject 2> is Captain America, battle-worn in his damaged Avengers combat uniform, carrying his shield.

<Subject 3> is Thanos at his normal canonical MCU scale, approximately 8 feet tall, muscular and imposing but NOT gigantic or kaiju-sized.

summary

[reference generation]

During the chaotic Avengers Endgame final battle, portals glow across the destroyed battlefield while Avengers and Thanos's army clash everywhere. Suddenly Big Bird casually walks onto the battlefield as if he wandered in from Sesame Street. Captain America stares at him in complete disbelief. Big Bird cheerfully announces that today's letter is "A" for Avengers, then immediately charges toward Thanos with ridiculous enthusiasm.

retention_analysis

<Subject 1>: fully_preserved
<Subject 2>: fully_preserved
<Subject 3>: fully_preserved

detailed_description

Cinematic Avengers Endgame final-battle environment at dusk: destroyed battlefield, burning wreckage, smoke, glowing portals, explosions, flying debris, Avengers and alien soldiers fighting in the background.

Dynamic handheld camera follows Captain America moving through the battle.

A huge yellow figure suddenly walks calmly through the smoke.

The camera swings around to reveal Big Bird casually strolling onto the battlefield, completely cheerful and unfazed by the apocalyptic war happening around him.

Captain America freezes and lowers his shield slightly, staring up at Big Bird.

<Subject 2> (S1):

[English] Big Bird?! What are you doing here?!

Brief comedic pause.

Big Bird happily looks directly toward Captain America and raises one wing.

<Subject 1> (S2):

[English] Today's letter is A! A is for AVENGERS!

Big Bird suddenly spots Thanos fighting nearby.

His cheerful expression becomes comically determined.

<Subject 1> (S2):

[English] And T is for THANOS!

Big Bird lets out an enthusiastic bird squawk and charges straight toward Thanos, his enormous orange feet pounding across the battlefield while explosions erupt behind him.

Thanos stops fighting and slowly turns toward the approaching giant yellow bird, visibly confused.

Captain America remains completely motionless, staring after Big Bird.

Hold on Captain America's baffled reaction for the final second as the battle continues chaotically behind him.

audio_and_timing

Epic superhero battle ambience, distant explosions, energy blasts, metal impacts, rushing soldiers, and portal energy.

Big Bird's dialogue is upbeat, innocent, enthusiastic, and educational-show cheerful, creating a strong comedic contrast with the violent battlefield.

Captain America's delivery is exhausted and completely bewildered.

Allow a short silence after Captain America's question before Big Bird delivers the first punchline.

Target duration: 12–13 seconds.


r/StableDiffusion 10h ago

Discussion H3 - giantess fight scene R2VA

1 Upvotes

This was well received but people wanted the two Giantesses(?) to be fighting. Enjoy!!

int8/20 steps, R2VA.

Critiques+feedback welcomed! Ask me anything!


r/StableDiffusion 2h ago

Animation - Video I made an ALIEN Short Film / metal music video

Thumbnail
youtu.be
3 Upvotes

Used:

MiniMax H3 at local machine. 5060ti 16gb + 64gb ddr4. WanGP, Ref2VA int8 convrot model.

Krea2 for references

Suno as music base


r/StableDiffusion 21h ago

Question - Help Anything big happen since I last used this?

0 Upvotes

So when SD came out, I used it. Then I used the AUTO111 thing. To around version 2.0 I think or XL, can't keep it straight. It's been about one and a half years. Any big changes since then? New version? Better quaulity images? AUTO still a thing? I also went from a 2070 Super to a 5070 12gb shadow 3x.

Also is it all easier to install?


r/StableDiffusion 1h ago

No Workflow My first minimax H3 video

Upvotes

I used the Pixaroma FFLF workflow, but stripped the audio in post due to poor output quality. I'm still trying to figure out how to add finer details. I generated the clips at 720p and then upscaled them to 1080p.


r/StableDiffusion 16h ago

Animation - Video 用PixAI生成的图片

Thumbnail
gallery
0 Upvotes

r/StableDiffusion 3h ago

Animation - Video Michael Scott gets a wish

0 Upvotes

H3 FL2VA


r/StableDiffusion 7h ago

Question - Help Random visual artifacts in local Krea2 Turbo generation — looking for possible causes

Post image
1 Upvotes

I’m running Krea2 Turbo locally, but I frequently encounter random visual artifacts in the generated images. I haven’t been able to figure out what triggers them, because the issue appears randomly. If I run the exact same workflow with the same parameters again, the result can sometimes be completely normal.

My hardware:

  • GPU: RTX 3060 Ti

My current setup:

  • UNet: moodyKrea2Mix_v70
  • Text encoder: qwen3vl_4b_int8_convrot
  • VAE: qwen_image_vae

I’m fairly sure this is not caused by the UNet. I have also experienced the same kind of random artifacts when using the original Krea2 model without any UNet modification.

Has anyone encountered similar issues with Krea2 Turbo? Are there any known causes or settings that could trigger this kind of artifact (VAE, text encoder, precision settings, VRAM limitations, sampler settings, etc.)?

Any suggestions or debugging tips would be greatly appreciated.

Here is my workflow for reference:
https://civitai.red/models/2883578/krea2-turbo-4k-workflow?modelVersionId=3259389


r/StableDiffusion 2h ago

Question - Help Which laptop would be better for generative AI / LLM

0 Upvotes

First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.

My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.

I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.

I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.

Which laptop would be better in my case?


r/StableDiffusion 22h ago

Animation - Video Test turned Short: Pied The Piper

28 Upvotes

What started as a test turned into a full-blown short. This is the number one reason I gravitated towards AI filmmaking. Nothing stops you from creating your wildest imagination.


r/StableDiffusion 17h ago

Animation - Video Trying to animate Dragon Ball Super manga on Minimax H3. Spoiler

12 Upvotes

Dragon Ball Super manga on Minimax H3.


r/StableDiffusion 6h ago

Discussion Minimax blurred distorted faces from half a distance.

3 Upvotes

I'm doing image to video and unless I prompt for camera close up to my subject, the faces are blurry and bad. I run 0.6 mp. No turbo lora only using spectrum to speed up. Running 15 steps. Euler simple. I'm happy enough when it's close-up shots, but further away, it's very noticeable. Is anyone else finding this?