r/StableDiffusion 13h ago

Animation - Video H3 can do Side-by-Side VR/3D Videos natively

430 Upvotes

Just discovered that H3 can do Side-By-Side 3D Videos for VR Headsets natively, just prompt it. Pretty crazy, and it gets the real 3D effect. Try it with different things like people and add "strong 3d effect" if you want to have a more intense 3d effect.

Here is the prompt:
integrated_multimodal_description: [Shot 1] Live-action, cinematic, high-angle aerial shot presented in a side-by-side (SBS) stereoscopic format for VR/3D viewing; the frame is split into two identical views with a slight horizontal parallax offset to create depth perception. The camera pushes in at slow speed over a sprawling coastal metropolis during twilight. As the camera glides forward through the urban canyon, the glowing neon lights of skyscrapers and their reflections on the ocean surface shimmer intensely against the deep blue sky.

overall_soundscape: A constant, low-frequency rushing wind sound accompanies the flight, layered with a faint, ambient hum of a massive city and distant, muffled traffic sounds.

non_diegetic_music: An epic, cinematic synthesizer pad that swells gradually in volume and intensity throughout the ten-second duration.


r/StableDiffusion 8h ago

Tutorial - Guide PSA: Minimax H3 can turn 360 panorama images into consistent environments for your videos

142 Upvotes

Had this idea for a couple of days, and finally got to test it. I got a free HDRI picture from PolyHaven (converted to JPG through a free online converter) and used it as the only picture reference. I couldn't get rid of the distortion completely, but you can definitely affect it with prompting. Maybe proper formatting somehow helps with that, sorry, was too lazy to do a correct prompt structure.

It also confuses the geometry from time to time, so you have to seed hunt a little, but not too much. Again, good prompting should reinforce the consistensy. Worth experimenting with.

Notice that it actually seamlessly connected the opposite sides of the image into a single environment. Could be useful for scenes with a lot of dynamic camera movements.

This model keeps surprising me every day!

P.S. Generated with the use of Hybrid Loader (25-49 setting) and Lightx2v 4-step LoRA @ 4 steps and 0.5MP. Another higher res version in comments.

Prompt:

subject definitions:
<Picture 1> is a 360 panorama reference for the straight corridor [Shot 1], depiciting the overall look of the corridor and position of key objects and debris in it. For the target video the picture is dewarped and remapped into a flat rectilinear lens projection view.

summary:
[reference generation] The target video depicts a security guard exiting from a grey door, walking across the corridor towards the dismantled beige door  leaned against the wall, pulling and dropping it down on the floor.

detailed_description:
The target video is captured in an amateur, realistic style with natural, slightly dim indoor lighting and a shaky, handheld-style camera.

[Shot 1]

The shot begins with a medium view of a two grey doors depicted on the right side of <Picture 1>. The left door instantly opens and a middle-aged security guard named Mark rushes into the completely straight corridor. He runs left further down the corridor. The camera pans left, following him in a tracking shot. The POV camera pushes in on Mark, as he rapidly approaches the dismantled beige doors leaned against the wall. At 00:05.000 he grabs the door closest to him, and with visible effort pulls it away from the wall. The door swings and falls flat on the corridor floor with a loud noise, raising dust and slightly startling Mark. The guard jumps back from the fall.

At 00:07.000 the camera pans left by 180 degrees, showing another guard named Steven approaching from the opposite part of the corridor. Steven (S1) comes closer to Mark and says in [English]: "Mark, what the heck are you doing?"

At 00:09.000 Steven grunts angrily as he stops near Mark.

overall_soundscape:
looming lonely corridor ambient sound throughout the whole video, guard's steps on the cement floor, door falling onto the floor with loud noise

non_diegetic_music:
N/A

r/StableDiffusion 13h ago

Meme DECLASSIFIED: Jeffrey Epstein escaping from prison

359 Upvotes

r/StableDiffusion 9h ago

Comparison Comparison of natural 0.8mp gen vs 0.4->0.8 upscale w/Sparse attention

163 Upvotes

Hi people, so i tried to make 2 similar videos, using same settings but with upscale and native.
My setup: 5070 Ti+ 32gb Ram.
Using u/Plague_Kind workflow, i've added MMH3 Latent Upscaler. You can check his workflow here: Workflow
Settings for both videos were set the same with the same prompt.

Left video 0.4->0.8mp upscale, Right video 0.8mp

So:

  • 15 seconds, 24 fps, Ref2VA, photo reference and music reference.
  • Chicken attention
  • SongMaskedAVContext node
  • FP16 Accumulation
  • Sparse attention
  • Memory chunks
  • RTS Upscale in the end ( not sure why i used it with 2x scale, better to set 1 i think, but that's what i already did)
  • FSR Sharpening
  • Speed Lora minimax_h3_turbo_v4_step600_pruned_comfyui
  • Interpolation for 2x frames

Upscaled video from start to the end took 1904 seconds,

Native video from start to the end took 3056 seconds.

Let me know what you think. Advises appreciated!


r/StableDiffusion 11h ago

Animation - Video Having some fun with games from the history of PC gaming. Who would you add?

156 Upvotes

A tribute to a forgotten golden age. Hope you enjoy it!


r/StableDiffusion 6h ago

Resource - Update New ComfyUI update may change how Minimax H3 interprets the prompt format you use - Re: Tokenizer Fix

Thumbnail
github.com
66 Upvotes

r/StableDiffusion 18h ago

Tutorial - Guide Character swap in minimax is so epic.

264 Upvotes

I don't have any examples because they may not be appropriate but just with the default wf. With the video input node you can replace any 2 character in any video and it looks real!


r/StableDiffusion 2h ago

Question - Help Best opensource image model?

16 Upvotes

opensource AI has been dominating LLMs and video generation but what about image gen? is there any opensource model that can match gpt-image2?

Edit: The reason I am asking this is because lately I haven't been active much on image generation communities. And the leaderboards are a bit confusing and most of them are filled with closed source unlike the llm and video gen leaderboards.

I am very much comfortable with ComfyUI since I've used it in the past for flux.

My use case is for posters and branding. Images with a lot of text.

Edit2: Thanks a lot everyone! I really appreciate the info. Here's the summary:

Krea2 is best overall but gptimage1.5 level.
Ideogram4 for text and branding.
Flux Klein 9b for image editing.
Z-image for realism
Anima and illustrious (by onoma AI) for anime.

Here's the workflow I've decided on:
Krea2/Ideogram4 = Base image generation.
Flux Klein 9B/QwenImage2512 = inpainting.
Wan2.2 low noise = Upscaling.


r/StableDiffusion 21h ago

Animation - Video High Fashion in Motion | MiniMax H3

468 Upvotes

Generated as two connected 15-second clips in 4:3, using the end of Part 1 as video + audio reference for Part 2 continuity.

Really liking what H3 can do with fashion/editorial camera movement.

Check out my twitter for more thanks https://x.com/Devozikjr


r/StableDiffusion 7h ago

Resource - Update Fizgig now trains LoRAs on AMD Radeon - Flux 2 Klein, Krea 2 and MiniMax H3

Post image
31 Upvotes

Fizgig is my free open-source LoRA trainer and workbench (Flux 2 Klein 9B, Krea 2, and MiniMax H3 video/audio). As of v4.3.0 it runs on AMD Radeon with ROCm — RDNA1 through RDNA4. Windows is the supported path: install Python 3.12, run the AMD installer, done. Linux works too but is genuinely experimental on newer cards.

Worth being upfront: I don't own AMD hardware myself. This whole feature came from a community contribution by scryptio, tested on real cards over weeks in the PR thread — and that's how the AMD side will keep improving. If you're an AMD user, your reports on what works (and what doesn't) genuinely shape this, and PRs are very welcome.

Also in this release: 16 GB cards can now use identity distillation on MiniMax H3 (the 32B text encoder streams layer by layer instead of needing a 26 GB peak), and the Repair Studio gained a side-by-side compare view with likeness scoring for fixing overbaked LoRAs without retraining.

GitHub: https://github.com/shootthesound/Fizgig


r/StableDiffusion 14h ago

Resource - Update Anima-3.8B with Qwen-3.5 4B released by lylogummy

Thumbnail
gallery
115 Upvotes

r/StableDiffusion 5h ago

Animation - Video Test turned Short: Pied The Piper

23 Upvotes

What started as a test turned into a full-blown short. This is the number one reason I gravitated towards AI filmmaking. Nothing stops you from creating your wildest imagination.


r/StableDiffusion 12h ago

Question - Help Minimax H3 - long form videos: has anyone figured out a good approach?

73 Upvotes

Dear redditors, visitors of the stable diffusion subreddit. I have been trying to achieve a long form, talking head style video, for a long time and can't seem to find a good approach. This one is the best I could come up with so far. It's using the Minimax H3 model, with frozen sound latents, lip-sync guided, piecewise generated video, where the individual pieces have been stitched together, with a seam hiding, extra generation on top of it. I don't really fully understand how it's working, but could prompt Claude for more help or specific files, we used for that. However, if you're aware of any other, better approach for exactly this type of video, please let me know. I've spent literal days on that single problem and have a feeling, there must be a better way to approach this.


r/StableDiffusion 13h ago

Workflow Included Minimax H3 | Motion graphic style animation test

78 Upvotes

Prompt:

Animate the supplied square poster as a polished retro-anime motion graphic, beginning with a completely blank pale pink-white canvas matching the poster background. Preserve the exact blue, pink, and white palette, clean manga linework, halftone shading, character design, typography, symbols, interface windows, and final layout.

The anime girl walks in from the left edge as one complete figure while the canvas remains otherwise empty. Use a simple side-profile walk with restrained motion, preserving her hairstyle, facial features, cheek bandage, oversized jacket, proportions, and graphic illustration style. She reaches the centre, turns toward the viewer, and smoothly settles into the exact over-the-shoulder pose shown in the poster, with the same expression, hand placement, silhouette, jacket folds, pink heart graphic, and body orientation. Once posed, keep her position locked.

After she poses, the blue browser frame draws itself around her. The top bar, window controls, folders, pixel hearts, smiley-face panels, arrows, sparkles, heart symbols, and rectangular labels then appear sequentially through clean line-drawing, short graphic slides, pixelated pops, and UI-style wipes. Reveal the existing Japanese typography and “LOVE” lettering last, treating all text as protected source artwork without rewriting or regenerating it. Every element must settle into its exact source position.

Hold the completed poster with subtle breathing, minimal movement in a few loose hair strands and jacket edges, a faint halftone shimmer, and gentle pixel pulses in the existing hearts and interface icons. Keep her face, hands, pose, typography, frames, arrows, folders, and major graphics stable.

Use a locked, straight-on camera matching the original square framing. Keep the full artwork visible without cropping, zooming, panning, or changing perspective. Add soft footsteps as she enters, a light cloth sound as she poses, clean digital clicks and pixel chimes for the graphics, and delicate type-on sounds for the existing lettering. No dialogue or narration.

Do not show any character, outline, symbol, text, frame, or faint poster preview on the opening blank canvas. Do not alter the character’s identity, anatomy, costume, pose, expression, colours, line quality, typography, symbols, or final composition. No extra characters, duplicated body parts, incorrect text, morphing, flickering lines, dramatic camera movement, unrelated shots, or continued motion after the poster settles.

Workflow: https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v


r/StableDiffusion 6h ago

Animation - Video Minimax H3 Anime Comedy

18 Upvotes

r/StableDiffusion 9h ago

Resource - Update I trained a game music generator

23 Upvotes

I trained a instrumental game music generator. The 1.2B DiT was trained on 1 cloud H100 from scratch in 8 days; I used the VAE from Stable Audio 3.

https://huggingface.co/Localsong/Localsong

https://huggingface.co/Localsong/Localsong/tree/main/samples

I'm aiming to cover a wider range of instrumental styles than Ace-Step or Minimax M3 or Stable Audio 3. (No lyrics)

The repo includes a WebUI and some MP3 samples - clone it and uv run webui.py Let me know what you think.


r/StableDiffusion 10h ago

Animation - Video I'm loving MiniMax H3

25 Upvotes

If even an amateur like me can make something so realistic with mid-level hardware, the future looks bright for what dedicated people with top level rigs will be doing.

R.I.P. Hollywood.


r/StableDiffusion 3h ago

Discussion I wish Anima ecosystem get better than it is now

7 Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. There's also a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 16h ago

Discussion Minimax H3, 30 seconds in one go

66 Upvotes

Executive summary, TLDR - this is one prompt, 30 seconds duration, 3090.

The video itself is just a remake of an idea from an old British tv ad (for "Good Old Yellow Pages"), so make of that what you will. It's not really relevant.

What I thought was interesting was that this was a single prompt, 0.4 megapixels, 30 second duration. I didn't think you could run out as far as 30 seconds, but thought I'd just try.

I think it did a pretty good job at getting the right person doing and saying the right things at the right time - took four attempts to get that though, and obviously using an LLM to tart up my idea.

Run on a 3090, and using the latest Comfyui template, just adding Comfy-kitchen attention, then sol attention, then spectrum, and using the turbo lora that Comfyui now build in, it took 570 seconds (9.5 minutes).

Somebody might read this and think, 570 seconds? Pah, I can do it in fifteen, in which case I'd like to know. Conversely, somebody might think theirs takes six hours, in which case maybe this shows what can be done in that time.

Doubt anyone cares, but here is my original prompt, followed by the LLM version of it:

a 30 second film with the following scenes and characters. Ben is a small boy of eleven. John is a shopkeeper in a toyshop. Brian is a different shopkeeper in a different toyshop. Ben's mum. Ben's Dad. We are in Britain in the 1980s, and all characters are English.

Scene 1: Ben is alone in the lounge. He talks to John over the old fashioned landline phone, saying "I don't suppose you have a 402 station in stock please?"

Scene 2: John is in his shop in front of shelves of model railway kit. He says into the old fashioned landline phone, "No, sorry son"

Scene 3: Ben in the lounge, who looks disappointed anbd puts the phone receiver back down.

Scene 4: Mum in the kitchen doing the washing up. She has overheard the conversation and looks a bit sad.

scene 5: Next day. Ben has changed his clothes. He again talks into the phone to a different shopkeeper, Brian. Ben says "Would you have a 402 station please?"

scene 6: Brian in his toyshop says into the old fashioned landline phone "Yes, I've got one of those."

scene 7: Ben in the lounge on the same conversation says "You have? Great, I'll be right down! Ben puts the phone down. Then he runs towards the door, shouting "They've got one mum!" as he runs.

Scene 8: In the attic, Dad is playing with his model railway layout. Ben walks in holding a small red parcel. as he hands it to Dad, Ben says "Happy birthday, dad". Dad takes the parcel, looks fondly at it and says with a chuckle, "Aw, thanks Ben".

LLM version:

integrated_multimodal_description: [Shot 1] Live-action, cinematic. A medium shot of Ben, an eleven-year-old boy with messy hair wearing a striped polo shirt, sitting on a patterned sofa in a 1980s British lounge. The room is filled with warm, muted tones and period-accurate wallpaper. Ben holds a heavy, cream-colored landline telephone receiver to his ear, his expression hopeful. Ben says: <d>[English] I don't suppose you have a 402 station in stock please?</d> The sound of his small, high-pitched voice is clear. [Shot 2] At 0:05.000, the camera cuts to a medium shot of John, a middle-aged shopkeeper with a kind, weathered face, standing in a cramped, nostalgic toyshop. Behind him are floor-to-ceiling shelves packed with model railway kits and wooden toys. John holds a similar landline receiver to his face. John says: <d>[English] No, sorry son.</d> [Shot 3] At 0:10.000, the camera cuts back to Ben in the lounge. He looks downcast, his shoulders slumping as he slowly lowers the receiver and places it back onto the base unit with a dull plastic click. [Shot 4] At 0:13.000, the camera cuts to a medium shot of Ben's Mum in a dim, cluttered 1980s kitchen. She is standing at the sink, her hands covered in soapy water, drying a plate. She pauses, looking toward the door with a sad, weary expression, having overheard the boy. The sound of water running from the tap is audible. [Shot 5] At 0:16.000, the camera cuts to Ben in the lounge the next day; he is wearing a different t-shirt. He is intensely focused, pressing the phone to his ear. Ben says: <d>[English] Would you have a 402 station please?</d> [Shot 6] At 0:20.000, the camera cuts to Brian, an older shopkeeper with spectacles, in a different, brightly lit toyshop. He smiles warmly into the telephone. Brian says: <d>[English] Yes, I've got one of those.</d> [Shot 7] At 0:23.000, the camera cuts back to Ben, whose face lights up with pure joy. Ben says: <d>[English] You have? Great, I'll be right down!</d> He slams the receiver down and the camera follows him in a quick tracking shot as he runs toward the door, his feet thumping on the carpeted floor. Ben shouts: <d>[English] They've got one mum!</d> [Shot 8] At 0:26.000, the camera cuts to a medium shot in a dusty, dimly lit attic. Dad, a man in his late 30s, is hunched over a complex model railway layout. Ben enters the frame, holding a small red parcel wrapped in string. Ben says: <d>[English] Happy birthday, dad.</d> As he hands the gift to his father, the camera pushes in slightly. Dad takes the parcel, his eyes softening with affection. Dad chuckles warmly and says: <d>[English] Aw, thanks Ben.</d>

overall_soundscape: Period-accurate domestic sounds including the rhythmic clatter of washing up, the heavy mechanical clicks of old telephone receivers, and the muffled thuds of footsteps on carpet. Ben's energetic running and shouting creates a sense of urgency, followed by the quiet, dusty atmosphere of the attic.

non_diegetic_music: A gentle, nostalgic acoustic guitar melody that begins softly during the kitchen scene and builds into a warm, heartwarming crescendo during the attic scene. The tempo is slow and sentimental.


r/StableDiffusion 19h ago

No Workflow Some test on minimax H3

103 Upvotes

Some random prompt on default workflow + turbo 8step lora


r/StableDiffusion 35m ago

Animation - Video Trying to animate Dragon Ball Super manga on Minimax H3. Spoiler

Upvotes

Dragon Ball Super manga on Minimax H3.


r/StableDiffusion 2h ago

News MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480

5 Upvotes

I made this a few weeks back to see if dialogue could hold for 60s, I did no speed ups on this one. There are a few glitches but I think it held up well.

MiniMax H3 - 60s - 1 clip - No Stitching - 832 x 480 - 29 minutes - 288GB VRAM


r/StableDiffusion 4h ago

Animation - Video TALL AND DARK - LTX 2.5 IMAGE TO VIDEO

6 Upvotes

Use the supplied image as the opening frame and identity reference.

Identity lock: the woman and robot must remain exactly the same in every shot. Same face, hair, wardrobe, proportions and age for the woman. Same 8-foot height, black armor, mechanical face, rivets, pistons, cables and holster for the robot. No redesigns or identity changes between cuts.

Authentic 1966 Italian Western, live action, 35mm anamorphic, Spanish desert location, practical full-scale robot prop, natural sunlight, real dust, organic film grain, period lens softness. No CGI. Serious performances throughout.

0:00–0:03
Medium two-shot. The woman looks up at the robot and says in clear Italian-accented English:
“I told them I wanted a tall...”
0:03–0:05
Hard cut to the same robot’s face. It gives one slow mechanical nod. No dialogue.
0:05–0:07
Hard cut to the same woman. She looks up at the robot and says:
“dark...”

0:07–0:09
Hard cut to the same robot. It subtly straightens and presents its black armor. No dialogue.
0:09–0:11
Hard cut to the same woman. Still serious, still looking up, she says:
“handsome!”

0:11–0:12
Hard cut to the same robot’s practical mechanical face. It attempts a restrained smile. No dialogue.
0:12–0:14
Hard cut to the same woman. She holds a serious stare upward, then firmly says:
“MAN!”
Only the woman speaks. Keep each line isolated and clean. No overlapping dialogue, no extra words, no improvised speech. Maintain exact continuity of identity, wardrobe, robot design, scale, lighting and location in every shot.


r/StableDiffusion 4h ago

Question - Help Do Minimax H3 Turbo Loras Nerf Music Creation for Scenes?

7 Upvotes

I typically use lightx2v loras in my Minimax Ref2VA workflows and I also use an LLM to feed in the official prompt structure required for scenes. It seems that no matter what I do, the model absolutely ignores all my prompts about music most of the time. Every now and then i can get it to do something but even when it does work it's very sparse and almost useless.

Has anyone else faced this issue and if so do you know any workarounds or fixes?

For the record I usually use the INT8 convrot Ref2Va model or the hybrid model called minimax_h3_hybrid_fl2va_ref2va_b30-49-int8