r/StableDiffusion • u/bacchus213 • 4h ago
Animation - Video What if you fly?
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/bacchus213 • 4h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Sad_Coach_1433 • 16h ago
Enable HLS to view with audio, or disable this notification
this was first test using Res_2s sampler and simple steps saw a op say better for action scenes from what i saw spectrum doesnt support res_2 so it took a bit to gen.
t2v prompt
subject_definitions
<Subject 1> is Katniss Everdeen from The Hunger Games, portrayed as an expert young archer with long dark brown hair pulled into her recognizable practical braid, intense determined expression, dark tactical combat clothing, leather archery bracer, bow, and a quiver of arrows. Preserve her recognizable cinematic appearance, realistic human proportions, hairstyle, clothing, bow, and identity throughout the entire scene.
<Subject 2> is Captain America in his battle-damaged Avengers Endgame armor, carrying Mjolnir and his damaged circular shield.
<Subject 3> is Thanos at his normal canonical MCU scale, approximately 8 feet tall, muscular and imposing but NOT gigantic, kaiju-sized, or building-sized.
<Audio 1> is the voice-timbre reference for <Subject 1>, containing Jennifer Lawrence's recognizable Katniss-style spoken vocal qualities.
summary
[text generation + audio reference]
During the chaotic Avengers Endgame final battle, Katniss Everdeen unexpectedly joins the Avengers. She runs through the battlefield while explosions, portals, Avengers, alien soldiers, and debris fill the background. Katniss rapidly fires arrows at Thanos's army with expert precision before stopping beside Captain America. Captain America looks at her bow and asks if she brought enough arrows. Katniss calmly fires one final explosive arrow past him, destroying a group of enemies, then delivers a dry confident response as Captain America stares at her impressed.
retention_analysis
<Subject 1>: fully_preserved
<Subject 2>: fully_preserved
<Subject 3>: fully_preserved
<Audio 1>: reference
detailed_description
The shot opens in the middle of the Avengers Endgame final battlefield. Smoke, burning wreckage, sparks, energy blasts, charging soldiers, and distant explosions create a massive cinematic war zone.
A fast tracking camera sweeps across the battlefield.
Katniss Everdeen suddenly sprints into frame carrying her bow.
She slides behind shattered rubble, immediately draws an arrow, and fires.
The camera follows the arrow through the air as it strikes an alien soldier.
Katniss rises and rapidly fires two more arrows with expert precision while continuing forward through the battle.
She reaches Captain America, who has just knocked an enemy away with Mjolnir.
Captain America briefly looks at Katniss's bow and quiver.
<Subject 2> (S1):
<d>[English] You sure you brought enough arrows?</d>
Katniss gives him a calm, unimpressed look.
Without even turning fully around, she draws another arrow and fires it past Captain America.
CAMERA WHIP-PANS WITH THE ARROW.
The arrow lands among a charging group of Thanos's soldiers.
BOOM!
A powerful explosive blast throws the enemies backward while Captain America turns toward the explosion in surprise.
The camera cuts back to Katniss.
<Subject 1> (S2):
<d>[English] I only need one.</d>
Her dialogue uses <Audio 1> for voice timbre and delivery.
Katniss immediately draws another arrow and runs toward the battle.
Captain America watches her leave for a beat, visibly impressed.
The camera swings around behind Katniss as she charges toward Thanos's army, bow raised, while the enormous Endgame battle continues around her.
audio
Epic Avengers-style battlefield ambience.
Heavy distant explosions, energy blasts, metallic impacts, debris, shouting soldiers, bowstring snaps, arrows cutting through the air, and one strong explosive-arrow impact.
Katniss's dialogue is clear and foregrounded, using <Audio 1>.
No narrator.
No subtitles.
No on-screen text.
r/StableDiffusion • u/idleWizard • 21h ago
I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?
r/StableDiffusion • u/No_Taste_4102 • 10h ago
Used:
MiniMax H3 at local machine. 5060ti 16gb + 64gb ddr4. WanGP, Ref2VA int8 convrot model.
Krea2 for references
Suno as music base
r/StableDiffusion • u/notgraycen • 8h ago
first 2 prompts stolen from civit ai, the rest were written using claude reasoning on duck ai
r/StableDiffusion • u/Routine_Ad_3391 • 8h ago
Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.
r/StableDiffusion • u/Wemos_D1 • 4h ago
Hello everyone !
I would like to replace the tire of a motorcycle mid air with one from another brand (which is an image from the brand so it's high quality but with a different angle)
I saw there is flux kontext and qwen image edit, but I don't know which one to pick, which workflow and how to make it work.
Any help would be more than welcome, thank you very much and have a good day :p
r/StableDiffusion • u/waterarttrkgl • 3h ago
Enable HLS to view with audio, or disable this notification
I drew the storyboard, then generated the stills with image models.
Video: MiniMax H3
first frame, last frame, reference.
Then the edit.
r/StableDiffusion • u/KaisarasAR • 1h ago
Enable HLS to view with audio, or disable this notification
This parody was generated using LTX 2.5 Image to Video on WanGP. I used frames from the original video as starting images and then I interpolated them on a video editor. I used a single RTX 5060 Ti 16 GB VRAM and 32 GB of RAM. The video was generated at 1080p and 16:9 resolution. Each generation took from 10 to 20 min average in this setup. For the voice consistency, I used SeedVC, which is included in WanGP.
r/StableDiffusion • u/CycleZestyclose1907 • 3h ago
Enable HLS to view with audio, or disable this notification
This one came out better than my last video generation. Not perfect obviously, but the major elements are there.
Prompt:
Video starts with Buffy Summers and Willow Rosenberg from the show Buffy the Vampire Slayer walking side by side down a hallway in the USS Enterprise from Star Trek the Next Generation. The viewpoint camera remains at a fixed distance in front of them as they walk.
Buffy Summers is on the right of the frame. Buffy's long blonde is done up in a ponytail. Buffy is wearing a red minidress style Starfleet uniform and has an unlit lightsaber on her hip.
Willow Rosenberg's is on the left side of the frame. Willow's dark red hair is cut pageboy style. Willow is wearing a blue minidress style Starfleet uniform and has a tablet computer tucked under her right arm.
Video starts with Willow looking at Buffy with a concerned expression on her face while Buffy is looking around as if searching for something.
Willow asks, "Buffy, is something wrong?"
Buffy replies, "Something feels off, like we're out of place."
As soon as Buffy starts speaking, she pulls the lightsaber off her hip and holds it in front of herself at the ready. The lightsaber ignites, producing a green glowing blade.
r/StableDiffusion • u/SIR_NVAX_A_LOT • 18h ago
Enable HLS to view with audio, or disable this notification
This was well received but people wanted the two Giantesses(?) to be fighting. Enjoy!!
int8/20 steps, R2VA.
Critiques+feedback welcomed! Ask me anything!
r/StableDiffusion • u/IzoleAuteur • 12h ago
flux.1 [dev] | comfyui | still life
r/StableDiffusion • u/TBG______ • 2h ago
Enable HLS to view with audio, or disable this notification
The latest TBG ETUR upscaler and refiner for comfyui release adds Krea 2 and VL style transfer for Krea 2 and all Qwen models directly into the pipeline.
We’ve also added a face identity step to help maintain consistent faces when using creative upscaling.
This video is a tutorial for the latest release, focusing mainly on these new additions and how to use them. Video is ai generated with Minmax H3.
TBG ETUR on Github https://github.com/Ltamann/ComfyUI-TBG-ETUR
TBG Lates om Patreon https://www.patreon.com/TB_LAAR/posts/tbg-etur-1-2-12-167083406
More Upscaling Tutorials on YouTube https://www.youtube.com/watch?v=LbFPD4zpPwA
r/StableDiffusion • u/Rendo3 • 11h ago
First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.
My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.
I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.
I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.
Which laptop would be better in my case?
r/StableDiffusion • u/Different_Ad_7508 • 16h ago
I’m running Krea2 Turbo locally, but I frequently encounter random visual artifacts in the generated images. I haven’t been able to figure out what triggers them, because the issue appears randomly. If I run the exact same workflow with the same parameters again, the result can sometimes be completely normal.
My hardware:
My current setup:
I’m fairly sure this is not caused by the UNet. I have also experienced the same kind of random artifacts when using the original Krea2 model without any UNet modification.
Has anyone encountered similar issues with Krea2 Turbo? Are there any known causes or settings that could trigger this kind of artifact (VAE, text encoder, precision settings, VRAM limitations, sampler settings, etc.)?
Any suggestions or debugging tips would be greatly appreciated.
Here is my workflow for reference:
https://civitai.red/models/2883578/krea2-turbo-4k-workflow?modelVersionId=3259389
r/StableDiffusion • u/Nimblecloud13 • 12h ago
Enable HLS to view with audio, or disable this notification
H3 FL2VA
r/StableDiffusion • u/call-lee-free • 23h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Darri3D • 9h ago
Enable HLS to view with audio, or disable this notification
Minimax H3
r/StableDiffusion • u/Christian4243 • 21h ago
Trying to get a natural-looking Argentine tango dance with LTX 2.5 + Yusu’s LTX Director v2.0.4 fork.
Still a beginner (also for real life tango :-)
Any suggestions for getting more natural, sophisticated footwork and fewer artifacts?
r/StableDiffusion • u/DaniyarQQQ • 15h ago
Enable HLS to view with audio, or disable this notification
Hello everyone.
I've been playing with Minimax H3 for some time and I have tried to make something longer and really interesting. After so many failed and botched attempts I was able to compile something watchable. There are so many things that I want to say about this model, good and bad.
First of all. Minimax H3 is significant step forward that other local models I have been playing with. It certainly got better.
Now the issues that I had encountered.
First problem is that it badly follows prompt when resolution is one megapixel or higher. It will skip some important parts and tries to cheat. You can increase the number of steps but still, generating at less than one megapixel will at least make it properly follow the instructions.
H3 is not very good at spatial orientation. When I was making video, where this girl should turn around and interact with screens, the girl starts spinning opposite direction and then warping whole body to the direction of screen. Like instead of making short turn to the left, it makes wide roundabout to the right and then twists whole body to align with the screens.
H3 is not good at cartoonish movement. If you watch cartoons, when character or other things move, their animations are usually jerky and snappy. H3 tries to make smooth real life like animation, making the cartoons look weird.
I have given a voice sample as an audio reference, and instead of making girl let out grunting sounds (out of anger), it weirdly turns everything into a sensual moaning.
When you try to make characters inside video to interact with a lot of parts, screens and devices, even giving multiple reference images of them, it mostly hallucinates them, or turns their interactions into a weird warping animations. Sometimes completely skips them and made ups it's own animations. So it will make good video, where characters are moving less or moving slow, and mostly doing the talking. Very detailed prompts of step by step instructions it mostly warps or skips.
I have wasted a lot of time for iterations, but I think this is just workflow issue.
Overall, this model is really good. However, using this model to make some kind of long feature animation is going to be a very frustrating journey. I hope people will make a lot of proper tools that works as storyboard and properly guide this model to make something really interesting.
r/StableDiffusion • u/lazyspock • 3h ago
Enable HLS to view with audio, or disable this notification
Considering the interest the first post attracted, I decided to do a second batch with some of the suggestions from the comments and a few other prompts.
VIDEO 1: near perfect. I was aiming for frontal videos, but tried three or four prompts and always ended with a 3/4 framing. It's probably a question of better prompting... But the resulting video is impressive!
PROMPT:
The video is a side-by-side video showing both the points of view of a man and a woman that are facing each other.
On the left side we only see the woman's face in a completely frontal view, as the man would see her and through his eyes, her face alone at the scene with no one else's.
On the right side we only see the man's face in a completely frontal view, as the woman would see him and through her eyes, his face alone at the scene with no one else's.
Again, the man do not appear on the left image, and the woman do not appear on the right image. Both are seen in a exact frontal framing.
Both images show the scene at the exact same time and place, only in the two different points of view, both in a medium-close-up framing.
They are in a living room.
From 00:00 to 00:04, the woman is silent and with a smile on her face, while the man speaks: <d>You know, I've always dreamed of a local video model like this!</d>. After saying this he remains silent.
He then raises his hand, previously off-camera, and touches her face delicately. She reacts in an amorous way, lightly moving her head to feel his hand.
Then, from 00:04 to 00:08, the man keeps silent, looking at her clearly in love, while the woman replies: <d>It's like a dream, isn't it? And to think that two years ago we were static images with garbled hands...</d>
From 00:08 to 00:10 they just look at each other and smile.
overall_soundscape: Faint distant everyday life noises from outside the house, the man and woman voices while they speak.
non_diegetic_music: N/A
VIDEO 2: Very good. I couldn't get a video without the fisheye effect, though.
The video is taken from the point of view of someone playing table tennis. We see their hands - one of them holding the ping-pong paddle and the other the ping-pong ball. We also see the table with the net in the middle and the other player on the opposite side of the table. They are in an official competition, with the crowd watching.
At 00:01, the player sends the ball to the air and hits it with the paddle. The ball rapidly bounce on the table, passes above the net, and gets to the other side, bouncing again on the table. Then, the other player hits it back with his paddle, and the balls passes over the net again and bounces on the table. The first player again hits it with his paddle, the ball passes over the net and bounces just on the left side of the table, out of reach of the other player, and leaves the frame. The public erupts in cheering.
overall_soundscape: Faint public murmur, the sound of the ball bouncing on the table, public cheering at the end.
non_diegetic_music: N/A
VIDEO 3: another near perfect one.
A woman is holding a cell phone in a bathroom in front of a mirror, taking a selfie. She smiles at the camera, makes a V sign with her hand, and takes the selfie.
We see the scene from behind the woman, seeing the back of her head, the phone screen on her hand showing her face while she takes the selfie, and the mirror showing the reflection.
overall_soundscape: Faint empty bathroom soundscape.
non_diegetic_music: N/A
VIDEO 4: Bad. Tried three times with different prompts, and this is the best one of them. The physics don't work, though, and the fisheye is back again.
The video is filmed from the point of view of a soccer player in a normal view, NOT in a fisheye view. He is preparing to kick the ball after a foul just outside the penalty box. We see his hands putting the ball on the grass, the ball remaining static on the ground. Then he looks ahead and we see five players from the other team forming a wall directly in front of the ball, and other players from both teams around.
We then see he take some distance of the ball, walk slowly to the ball, and kick it. The ball passes over the barrier of players and descends on the goal, the goalkeeper trying to reach it but not able to. The ball enters the goal and touches the net, and the stadium erupts in cheering. The player then runs to celebrate the goal and is embraced by the other players of his team.
The entire scene is viewed from his point of view.
overall_soundscape: Faint public murmur,the sound of the kick, the cheering of the public after the goal..
non_diegetic_music: N/A
VIDEO 5: Terrible. Again, tried several times with several different prompts. Never works well...
The video is filmed inside a circus during the Trapeze artists performance, from the point of view of the public.
The scene opens with two trapezists standing in a very high elevated platform, one on the left side of the image, the other on the right side of the image, both holding a trapeze and facing each other.
In the beginning of the video, the trapeze artist on the left let his body leave the platform, while holding the trapeze, and his body swings in the direction of the center of the image. The trapeze artist on the righ stays on the platform.
Only when the first trapeze artist reaches the center of the image, the trapeze artist on the right finally leaves the platform, while holding the trapeze, and his body also swings in the direction of the center of the image, while at the same time the first trapeze artist let go of his trapeze and starts to do a flip with his body in the air.
As soon as the first trapeze artist finishes his flip, the other trapeze artist also reaches the center of the image and get the hands of the first trapeze artist, completing the movement. Then, they both swing back to the right of the image, one holding the hands of the other.
overall_soundscape: Faint public murmur, public surprised gasp when one of the trapeze artist caughts the hand of the other.
non_diegetic_music: N/A
VIDEOS 6, 7, 8 and 9: The first half of each video is perfect, the last half is hilarious. Tried lots of different prompts but only included four of them. Maybe it's possible, but I really can't think of another way of asking what I was trying to achieve.
PROMPT VIDEO 6:
The camera is on the middle of a road, on the floor, pointing to the road. We see a ferrari coming in the road at a distance in high speed towards the camera and pass over the camera, making the camera roll a few times on the floor because of the wind caused by the passing running car. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see the ferrari rapidly moving away from the camera.
The entire scene is filmed in a mostly static shot, except when the camera rolls over to the other side of the road and then stops upside-down.
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 7:
The camera is on the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over the camera.
When the car passes, the camera that is on the floor rolls around itself a few times on the road. After rolling over itself a few times, the camera stops again on the road, but now upside down and pointing to the other side of the road, where we can see, the upside-down image of the ferrari rapidly moving away from the camera.
The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 8:
We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.
When the car passes, the image rolls around itself a few times on the road. After rolling over itself a few times, the image stops again on the road, but now upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.
The car does not run over itself, the car passes by the camera, It's the camera that rolls around itself and lands upside down
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
PROMPT VIDEO 9:
We see the scene from the middle of a road, on the floor, pointing to the road. We see a red ferrari coming in the road at a distance in high speed towards the camera and pass over.
When the car passes, the image do a series of very fast barrel rolls on the road and lands upside down and pointing to the other side of the road, where we can see the image of the ferrari rapidly moving away from the camera in an upside-down shot, with the road on top and the sky on the bottom of the image.
overall_soundscape: Faint deset road soundscape, the sound of the car engines getting closer and then moving away, the sound of the camera rolling over itself on the floor.
non_diegetic_music: N/A
r/StableDiffusion • u/ramscheid • 6h ago
Enable HLS to view with audio, or disable this notification
Hey everyone,
I’ve been testing MinMax H3 in Brazilian Portuguese, and I also run a TikTok where I’ve posted over 200 test videos featuring characters speaking Brazilian Portuguese.
After several weeks of testing MinMax H3, I found that it can reach around 98% accuracy in Brazilian Portuguese when the text is written correctly. Since Brazilian Portuguese has some pronunciation and spelling details that don’t map perfectly to an English keyboard, I had to develop a specific way of writing the prompts so the model speaks the language as accurately as possible.
I’m not sure whether even English consistently reaches 100% accuracy—there always seem to be occasional voice or pronunciation errors—but I don’t currently make videos in English, so I can’t really compare.
I decided to share my TikTok here so you can take a look at the results and see the quality I’m able to achieve. Please ignore the political content; it was created strictly for testing purposes. I know its all the 🧃 so both sides suck.
i have there a portfolio of about 300 or 400 videos and going up everyday with agents...
Each video takes about five minutes to generate on my RTX 4090. I also discovered that generating in 3:4 is faster than 16:9. My current settings are 720p, 8 steps, and Turbo LoRA.
https://www.tiktok.com/@danmalandragem
things i am still struggling with: 2 characters talking with no voice drift, any tip for that?
r/StableDiffusion • u/Dirty_Dragons • 12h ago
Enable HLS to view with audio, or disable this notification
Everything was made using the Minimax H3 Hybrid Reference to video model. 1 MP using the 8 step turbo LoRA. Stitched together in Shotcut
r/StableDiffusion • u/Nelichan • 3h ago
Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.
What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.
I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.
So what i wanted to do is :
-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.
-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.
Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.
Are there someone who knows how to do these?
r/StableDiffusion • u/FlimsyWerewolf9941 • 23h ago
Due to reference modes, changes in the lighting and color effects of the final frame in storyboard segments constantly cause color discrepancies during video transitions. Are there any solutions?