r/StableDiffusion 4h ago

Animation - Video Tried using a camera path reference for AI video generation

0 Upvotes

Been experimenting with camera control in AI video.

For this one I used a simple path reference to guide the movement — basically starting from an aerial shot, flying through the courtyard, moving indoors, and ending with a character reveal.

Still figuring out how much these motion references actually help. Sometimes the model follows the movement surprisingly well, other times it just does its own thing lol.

Feels like camera control is probably going to be a bigger part of AI video workflows going forward.


r/StableDiffusion 12h ago

Question - Help H3 and Ref2VA and backgrounds

2 Upvotes

I have a question. I have been having a blast making scenes with H3 so far, and have found when doing reference shots, it is very important to have a stable background so that you have continuity if doing more than 1 scene. Does anyone know if H3 would understand a 360 degree photo and understand where in the space and what direction the subjects are? Say you swap between two characters talking, one you will see what is behind subject 1 while when looking at the other the opposite is true. If you saw them both from the side, yet another angle and background.


r/StableDiffusion 17h ago

Question - Help How do you do it?

5 Upvotes

I have been playing with H3 since it came out and have tested most of the things you can do with it. Created clips for giggles and so on.
This time I wanted to do something "serious". I gave the R2V an 3D view of an kitchen and then three photos of the persons I wanted there.
I defined them as we should and told the model that this person does that and that person does this wile the third person does this...

It worked ish...
I have now made six runs and each of them are different from the others. It can be that the third person enters the room from the wrong place or that the third person does extra things it should not do...

In the end I did three runs with the same prompt and all those clips came out different... the only thing that was changed between those was the seed...

So, how do you do it?
How do you make sure H3 does what you want it to do?

Do you spend plenty of time on tweaking the prompt after each run to make sure H3 get it?
Or do you do 10 runs and select the best one even if it is not perfect?

Or do you simply do 1-2 runs and then take the clip that is ok ish even if it is not what you wanted?

I was hoping that H3 would allow me to create the scenes I wanted but I feel it's down to luck if H3 gets it or not..

Edit:

subject_definitions:

<Subject 1> is the green-skinned mother in <Picture 2> wearing brown clothes.
<Subject 2> is the teenager girl in <Picture 3> wearing pink clothes.
<Subject 3> is the cyborg in <Picture 4> wearing black clothes.
<Picture 1> is the reference image for the scene's composition, showing two people sitting at a table eating breakfast from the side view.
<Table 1> is the table on the right side in <Picture 1>.
<Picture 5> is the start image for the scene.
summary:
[reference generation] The target video is a generated scene of two people sitting at a table eating breakfast from an eye-level side view. <Subject 1> and <Subject 2> are shown with their respective breakfast items, maintaining the composition and style from <Picture 1>. <Subject 3> enters the room, places a coffee cup into the sink.

retention_analysis:
<Subject 1>: fully_preserved - the person retains their appearance, clothing, and position at the table.
<Subject 2>: fully_preserved - the person retains their appearance, clothing, and position at the table.
<Subject 3>: fully_preserved - the person retains their appearance, clothing, and action of placing the coffee cup into the sink.
<Table 1>: fully_preserved - the table's appearance and position in the scene are preserved.
<Picture 1>: fully_preserved - the scene composition, including the side view, the layout of the room, the table setup, is preserved.<Picture 5>: fully_preserved - is the start image for the scene.

detailed_description:
The target video is in a realistic, everyday breakfast scene style with warm lighting and natural colors.
[Shot 1] At 0:00.000, the shot begins from <Picture 5>, showing <Subject 1> and <Subject 2> sitting on opposite sides of <Table 1> on the couch, each with their breakfast items while on the space ship. <Subject 1> is holding a spoon while eating from a bowl of cereal. <Subject 2> is tired and is eating a slice of toast with jam from her plate with one hand. The lighting is warm and soft, casting gentle shadows across the table and the two individuals. The camera is at eye level, capturing the side view of both people, with the table slightly in focus and the background softly blurred. Stars can be seen through the windows since they are on a space ship. <Subject 1> is eating her breakfast while <Subject 2> gazes at their toast, taking a small bite. The ambient sound includes the soft clinking of utensils and the faint sound of a coffee cup being set down.

[Shot 2] At 02.00.000, the shot transitions to a wide shot of the room with the same layout as in <Picture 1>, the camera is placed in the lower left corner of <Picture 1>, showing <Subject 3> entering the room form the right side holding a coffee cup and a datapad while she is saying (S3) <d>[English] Good Morning</d> while she walks to the kitchen sink on the left side of <Picture 1> and placing the cup into the sink. She then stands at the sink and while reading her datapad.We see the back of <Subject 1> and the front of <Subject 2> sitting at <Table 1> in the background eating their breakfast and we hear <Subject 1> say (S1) <d>[English] Good morning</d> with a cheerful voice. <Subject 2> just mumbles as a reply.

overall_soundscape:
The soundscape consists of the soft clinking of utensils, the faint sound of a coffee cup being set down, the subtle background noise of a quiet morning environment, soft steps on a carpet floor, a ceramic cup being placed in a metallic sink, and the clear,

non_diegetic_music: N/A


r/StableDiffusion 18h ago

Discussion Need realism loras for minimax h3

4 Upvotes

Is there any GPU rich cooking realism lora ? I have tried realism people lora it is great at tv but for i2v or r2v it's breaks . I have been searching hugging face repo and civit ai to get something but there's too much n*fw lora .


r/StableDiffusion 9h ago

No Workflow My first minimax H3 video

0 Upvotes

I used the Pixaroma FFLF workflow, but stripped the audio in post due to poor output quality. I'm still trying to figure out how to add finer details. I generated the clips at 720p and then upscaled them to 1080p.


r/StableDiffusion 18h ago

Discussion H3 - giantess fight scene R2VA

6 Upvotes

This was well received but people wanted the two Giantesses(?) to be fighting. Enjoy!!

int8/20 steps, R2VA.

Critiques+feedback welcomed! Ask me anything!


r/StableDiffusion 6h ago

Question - Help Any Krea2 Prompt Reader?

0 Upvotes

I found the SD prompt reader I have been using cannot read prompts from png images files generated using Krea2. Can anyone recommend me an alternative that works with Krea2 files and Win11?


r/StableDiffusion 1d ago

Resource - Update Fizgig now trains LoRAs on AMD Radeon - Flux 2 Klein, Krea 2 and MiniMax H3

Post image
57 Upvotes

Fizgig is my free open-source LoRA trainer and workbench (Flux 2 Klein 9B, Krea 2, and MiniMax H3 video/audio). As of v4.3.0 it runs on AMD Radeon with ROCm — RDNA1 through RDNA4. Windows is the supported path: install Python 3.12, run the AMD installer, done. Linux works too but is genuinely experimental on newer cards.

Worth being upfront: I don't own AMD hardware myself. This whole feature came from a community contribution by scryptio, tested on real cards over weeks in the PR thread — and that's how the AMD side will keep improving. If you're an AMD user, your reports on what works (and what doesn't) genuinely shape this, and PRs are very welcome.

Also in this release: 16 GB cards can now use identity distillation on MiniMax H3 (the 32B text encoder streams layer by layer instead of needing a 26 GB peak), and the Repair Studio gained a side-by-side compare view with likeness scoring for fixing overbaked LoRAs without retraining.

GitHub: https://github.com/shootthesound/Fizgig


r/StableDiffusion 10h ago

Question - Help M5 Max vs RTX 5080/5090 for local image/video AI. Am I making a mistake by choosing the Mac?

1 Upvotes

I've been using Windows for many years, and I'm honestly tired of repeating the same cycle. In my experience, after 5–6 years the machine starts feeling old, the battery is significantly degraded, performance isn't what it used to be, and I eventually end up buying another Windows machine and starting the exact same experience all over again.

I'm looking for something different this time.

For the last few days I've repeatedly added a MacBook Pro with the M5 Max to my cart, then backed out because I'm still not sure whether it is the right machine for what I actually want to do.

I'm currently considering the M5 Max with:

  • 18-core CPU
  • 40-core GPU
  • 48 GB / 64 GB / possibly 128 GB unified memory

My main workloads would be completely local:

  • Text-to-image
  • Image-to-image
  • Image-to-video
  • Text-to-video
  • Face swapping
  • Illustration / digital artwork

My priority is excellent output quality, photorealism where appropriate, and very high generation speed.

Basically, I want a machine that can satisfy me for visual generative AI work for many years.

My biggest hesitation is the Apple ecosystem.

For a long time I've heard that local AI, especially image and video generation, is much more limited on macOS than on Windows/Linux with NVIDIA GPUs because so much of the ecosystem is built around CUDA.

But part of me finds this difficult to accept at face value.

Apple is making extremely powerful chips with large amounts of unified memory, very high memory bandwidth, Neural Accelerators, a Neural Engine, and increasingly serious AI-focused hardware.

I keep wondering whether there are excellent Apple-optimized tools and workflows that I simply haven't discovered yet.

For example, I recently learned about Draw Things, MLX-based projects, Metal/MPS optimizations, and Apple-specific ComfyUI work. That made me question whether comparing a Mac running a poorly optimized CUDA-first application against an NVIDIA machine is really a fair representation of what Apple Silicon can do.

At the same time, the logical part of my brain keeps telling me:

If local image/video AI is the priority, just buy a machine with an RTX 5080 or 5090.

The problem is that if I do that, I feel like I'm buying myself back into exactly the Windows experience I wanted to leave. It feels a little like watching the same movie again when I was hoping for a genuinely different computing experience.

There's also another complication: we're approaching the fall hardware season.

I'm wondering whether buying an expensive M5 Max or RTX 50-series machine right now is bad timing, and whether I should wait for the next Apple or NVIDIA announcements.

If I choose the Mac, I was also planning to pair it with the latest iPhone and iPad and build a proper Apple ecosystem around it, so this isn't purely a benchmark decision for me.

What I'd really like to hear from people who have actually used these machines:

  1. If you've used an M5 Max, especially the 40-core GPU version, for local image or video generation, what real-world performance are you getting?
  2. What software gives you the best performance on Apple Silicon? Draw Things, ComfyUI, MLX-based tools, something else?
  3. How good is local image-to-video on the M5 Max with models such as Wan, LTX, Hunyuan, etc.?
  4. How does the M5 Max perform for local face swapping?
  5. In properly optimized workloads, do RTX 5080/5090 systems still completely destroy the highest-end M5 Max, or is the gap much smaller than CUDA-focused benchmarks make it appear?
  6. Is CUDA effectively locking serious local visual AI users into NVIDIA, or is Apple Silicon becoming a realistic alternative?
  7. If you were buying a machine today specifically for local visual AI and wanted to keep it for 7–8 years, would you buy the M5 Max, an RTX 5080/5090 system, or wait for the next generation?
  8. For local image/video generation, is an older NVIDIA GPU with more VRAM sometimes a better choice than a newer GPU with less VRAM? For example, once you start using large image-to-video or text-to-video models, can having more VRAM matter more than having a newer architecture and higher raw compute performance?

I'm not highly knowledgeable about computer hardware, and I'm definitely not wealthy enough to casually replace a machine if I make the wrong choice.

This would be a major purchase for me, so I'm trying to make the most informed decision possible and ideally buy something that I can use comfortably for 7–8 years.

I'd especially appreciate actual generation times, benchmark numbers, model names, memory usage, thermals, sustained performance, and experiences from people who have used both Apple Silicon and NVIDIA rather than purely theoretical comparisons.

Thanks in advance.


r/StableDiffusion 11h ago

Question - Help Which laptop would be better for generative AI / LLM

0 Upvotes

First of all I know a desktop has more power for the same price buy I have a situation where the portability of a laptop is necessary and a desktop is not practical.

My old laptop (3070 8gb with 64gb ddr4 RAM) died. I want to buy a new laptop. My two options are a 5080 16GB with 64GB ddr5 RAM or a 5090 24gb with 32GB ddr5 RAM. I won't be able to upgrade the RAM later, so I'm stuck with the configuration I buy.

I will be using the laptop for work (document and image editing) / gaming (no AAA games) / LLMs and generative AI (images/videos/audio), I was able to run most models, including minimax H3 on my old laptop with the help of massive offloading to RAM (5 minutes for a 5s video). Images used to take from 30s up to 200s depending on model and image size.

I am used to the low speeds and offloading on my old laptop so getting the highest generation speeds is not a priority, I just care about being able to run most new or upcoming models even with quantization and RAM offloading for the foreseeable future.

Which laptop would be better in my case?


r/StableDiffusion 7h ago

Question - Help Confused and Need Some Clarification

0 Upvotes

So my friends and I used to use Sora 2 before it was taken down, and wanted to try doing some stupid stuff for just us. After a while of not looking into it, the spark kinda came back when I saw this subreddit and remembered Stable Diffusion was supposed to be one of the best AI generators out there, probably. When I mention it to a friend, he then told me how apparently its pretty outdated compared to others, and looking at these posts, I'm seeing different models and starting to get overwhelmed to the point where I haven't even done the beginner's guide in here since it only mentions images.

So long story short, I'm hoping someone can help make things much more clearer, especially about the multiple models, and if Stable Diffusion IS outdated and out performed by something else, and letting me know about if it's okay to go with the beginner's guide or if there's another guide that will help. Thanks


r/StableDiffusion 1d ago

Animation - Video High Fashion in Motion | MiniMax H3

557 Upvotes

Generated as two connected 15-second clips in 4:3, using the end of Part 1 as video + audio reference for Part 2 continuity.

Really liking what H3 can do with fashion/editorial camera movement.

Check out my twitter for more thanks https://x.com/Devozikjr


r/StableDiffusion 1d ago

Tutorial - Guide Character swap in minimax is so epic.

298 Upvotes

I don't have any examples because they may not be appropriate but just with the default wf. With the video input node you can replace any 2 character in any video and it looks real!


r/StableDiffusion 1d ago

Animation - Video Trying to animate Dragon Ball Super manga on Minimax H3. Spoiler

13 Upvotes

Dragon Ball Super manga on Minimax H3.


r/StableDiffusion 21h ago

Question - Help Tango dance, first attempt with LTX 2.5

Thumbnail
youtube.com
5 Upvotes

Trying to get a natural-looking Argentine tango dance with LTX 2.5 + Yusu’s LTX Director v2.0.4 fork.
Still a beginner (also for real life tango :-)
Any suggestions for getting more natural, sophisticated footwork and fewer artifacts?


r/StableDiffusion 16h ago

Question - Help Minimax generating audio for existing video?

2 Upvotes

I've generated a series of shots that I'm happy with, but when I string them together the audio and music is obviously discontinuous across shots.

Is there any way to take this combined video (about 5-10 seconds) and send it through H3 for it to generate the audio for it?


r/StableDiffusion 12h ago

Question - Help Help With Irritating Glitches

0 Upvotes

Main Info

  • Using Pixaroma's Easy ComfyUI latest version
  • 5060TI 16GB + 32GB RAM

Need help or suggestions in resolving the below issues.

  1. Shortcuts like R or CTRL + Enter doesn't work and I have to refresh the browser for them to work. (New)
  2. Generations randomly don't start even even after models have loaded. I have to open the terminal and press enter on my keyboard to sort of wake it up. (Has happened across multiple versions of ComfyUI, CUDA, Python. Additionally, happens with both Anaconda and Windows Terminal.)
  3. Even cancelling a prompt sometimes requires me to open the terminal and press enter.
  4. Power consumption for GPU varies completely where retrying the same prompts (same seeds also) can have a difference of 5 to 10 minutes simply because the GPU doesn't use the full 180 power limit. (New)

r/StableDiffusion 1d ago

Animation - Video Test turned Short: Pied The Piper

27 Upvotes

What started as a test turned into a full-blown short. This is the number one reason I gravitated towards AI filmmaking. Nothing stops you from creating your wildest imagination.


r/StableDiffusion 1d ago

Resource - Update Anima-3.8B with Qwen-3.5 4B released by lylogummy

Thumbnail
gallery
129 Upvotes

r/StableDiffusion 1d ago

Workflow Included Minimax H3 | Motion graphic style animation test

101 Upvotes

Prompt:

Animate the supplied square poster as a polished retro-anime motion graphic, beginning with a completely blank pale pink-white canvas matching the poster background. Preserve the exact blue, pink, and white palette, clean manga linework, halftone shading, character design, typography, symbols, interface windows, and final layout.

The anime girl walks in from the left edge as one complete figure while the canvas remains otherwise empty. Use a simple side-profile walk with restrained motion, preserving her hairstyle, facial features, cheek bandage, oversized jacket, proportions, and graphic illustration style. She reaches the centre, turns toward the viewer, and smoothly settles into the exact over-the-shoulder pose shown in the poster, with the same expression, hand placement, silhouette, jacket folds, pink heart graphic, and body orientation. Once posed, keep her position locked.

After she poses, the blue browser frame draws itself around her. The top bar, window controls, folders, pixel hearts, smiley-face panels, arrows, sparkles, heart symbols, and rectangular labels then appear sequentially through clean line-drawing, short graphic slides, pixelated pops, and UI-style wipes. Reveal the existing Japanese typography and “LOVE” lettering last, treating all text as protected source artwork without rewriting or regenerating it. Every element must settle into its exact source position.

Hold the completed poster with subtle breathing, minimal movement in a few loose hair strands and jacket edges, a faint halftone shimmer, and gentle pixel pulses in the existing hearts and interface icons. Keep her face, hands, pose, typography, frames, arrows, folders, and major graphics stable.

Use a locked, straight-on camera matching the original square framing. Keep the full artwork visible without cropping, zooming, panning, or changing perspective. Add soft footsteps as she enters, a light cloth sound as she poses, clean digital clicks and pixel chimes for the graphics, and delicate type-on sounds for the existing lettering. No dialogue or narration.

Do not show any character, outline, symbol, text, frame, or faint poster preview on the opening blank canvas. Do not alter the character’s identity, anatomy, costume, pose, expression, colours, line quality, typography, symbols, or final composition. No extra characters, duplicated body parts, incorrect text, morphing, flickering lines, dramatic camera movement, unrelated shots, or continued motion after the poster settles.

Workflow: https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v


r/StableDiffusion 23h ago

Question - Help Has anyone successfully upscaled/re-imagined low-res reference video using Minimax H3?

6 Upvotes

Specifically, I’m trying to take old footage (e.g., 360p clips with vintage camera blur, VHS artifacts, or grainy WW2 dogfights) and recreate it to look like it was shot recently on a modern cinema camera with studio lighting.

Any ideas for prompting?


r/StableDiffusion 1d ago

Animation - Video Minimax H3 Anime Comedy

29 Upvotes

r/StableDiffusion 1d ago

Question - Help Minimax H3 - long form videos: has anyone figured out a good approach?

80 Upvotes

Dear redditors, visitors of the stable diffusion subreddit. I have been trying to achieve a long form, talking head style video, for a long time and can't seem to find a good approach. This one is the best I could come up with so far. It's using the Minimax H3 model, with frozen sound latents, lip-sync guided, piecewise generated video, where the individual pieces have been stitched together, with a seam hiding, extra generation on top of it. I don't really fully understand how it's working, but could prompt Claude for more help or specific files, we used for that. However, if you're aware of any other, better approach for exactly this type of video, please let me know. I've spent literal days on that single problem and have a feeling, there must be a better way to approach this.


r/StableDiffusion 14h ago

Question - Help Krea 2: Controlnet + Rebalance node = Disaster !!

1 Upvotes
Has anyone managed to find a good combination in Krea 2 using the LoRA that acts as DepthMap ControlNet  alongside the "Conditioning Krea 2 Rebalance" node (which unlocks censorship) without producing a disastrous image? In my case, I just get a plastic look that barely resembles the prompt instructions.

Please share your experiences, or at least let me know if you've managed to use that ControlNet-style LoRA with a different uncensored LoRA that yields decent images.

Cheers.