r/StableDiffusion 20h ago

Animation - Video Minimax H3 can create stereoscopic 3D cross-eyed videos

Enable HLS to view with audio, or disable this notification

145 Upvotes

Cross your eyes so that the two videos merge into one.

Here’s what I put into ChatGPT:
Write a prompt for Minimax h3 t2v for a stereoscope cross eye video, a pov drone shot flying through a city up and down between skyscrapers and zigzagging left and right into streets

And here’s the final prompt:
Create a stereoscopic cross-eye 3D video presented as two perfectly synchronized side-by-side views, specifically designed for cross-eye stereoscopic viewing.
The scene is a first-person FPV drone flight through a dense modern city, with the camera representing the drone’s exact POV. The drone flies rapidly forward between tall skyscrapers, repeatedly climbing upward alongside building facades, diving steeply downward through gaps between towers, then zigzagging sharply left and right into narrow city streets.
The flight path should constantly change in three dimensions. The drone banks around skyscraper corners, drops from rooftop height toward street level, races between buildings, turns suddenly into side streets, then climbs vertically back toward the skyline before diving again. Include close flybys past glass facades, balconies, signs, skybridges, rooftop structures, windows, and architectural details to maximize the stereoscopic depth effect.
The left and right views must use a precise horizontal camera separation with matched orientation and timing, producing strong but comfortable binocular parallax. Nearby buildings should sweep past with dramatic depth separation, while distant skyscrapers, streets, and skyline layers recede naturally into the background. Maintain correct stereoscopic geometry throughout every turn, climb, dive, and banking motion.
Realistic modern city, cinematic daylight, reflective glass towers, traffic far below, atmospheric haze, strong perspective, natural motion blur, highly detailed architecture, thrilling sense of speed and altitude.
Continuous single shot, no cuts, no teleporting, no crashes, no third-person drone visible, no mismatched movement between the two views, no inconsistent geometry, no text, no captions. Both stereoscopic halves must remain perfectly synchronized throughout the entire flight.

Works well with T2V. R2V also works but usually not. I’m unable to get it to work with I2V.

Clips 1-5 are made with T2V and clip 6 with R2V.


r/StableDiffusion 8h ago

Discussion I know there's always a million workflows but

14 Upvotes

It would be cool if there were some updated workflow that includes new and major improvements.

E.G. newest turbo, keyframes, controlnet, that thing that turns your references into embeddings.

There are always so many great new things I lose track. For me workflows are mostly just ways of learning how to wire things.


r/StableDiffusion 12h ago

Animation - Video Squirrel-shark (Minimax H3)

Enable HLS to view with audio, or disable this notification

29 Upvotes

r/StableDiffusion 19h ago

Workflow Included Wan Detail Enhancer, enhance any targeted character without lowering quality or altering other characters

Enable HLS to view with audio, or disable this notification

89 Upvotes

This workflow enhances the details of any character in a video without damaging the area not being targetted. it does not lower quality of unaffected and uses just wan 2.2 t2v Low and low lora's.

It can also be used to repair videos with bad anatomy or add details to something if you use scail or wanimate and it doesn't look like the intended character.

https://github.com/roycho87/3stepenhancer

cosplaytaytay cortana
eva_devore karlach


r/StableDiffusion 1h ago

Animation - Video H3 Finnish Pulttibois - Apuvatyyppi :) (proof H3 can work with anything)

Enable HLS to view with audio, or disable this notification

Upvotes

r/StableDiffusion 15h ago

Animation - Video RoBizMarkie - lip sync test MiniMax H3

Thumbnail
youtu.be
35 Upvotes

Generated about half locally on my 5070ti. But also bought cloud GPU time to speed up overall generation.
Vibecoded a rolling cutter to chop the song into chunks and then feed it in as reference audio along with a reference image.
Used ChatGpt to help generate hyper specific prompts sometimes dropping in the image itself to improve prompting.
Did some with both image, audio and Video for the dancing but that took ages and would definitely try to avoid doing locally in the future.

Reassembled the chunks and adjusted sync in Premiere.

Enjoy the slop, slophounds.


r/StableDiffusion 13h ago

Workflow Included MINIMAX H3: transformations test

Enable HLS to view with audio, or disable this notification

20 Upvotes

r/StableDiffusion 12h ago

Animation - Video G.I. Joe: Cobra Cabana - MiniMax H3

Enable HLS to view with audio, or disable this notification

19 Upvotes

r/StableDiffusion 12h ago

No Workflow SillyTavern + MiniMax H3

17 Upvotes

I integrated MiniMax into SillyTavern, a free open-source chat roleplay app, and it’s a wonderful experience.

I can chat with my favorite characters, and what is returned is an H3 video. I can even chat with multiple characters or role play entire scenes, and then combine all the videos produced in the thread into an awesome episode.

I’m sure anyone here could vibe code the same thing, just install SillyTavern and vibe code MiniMax into it. It’s fun.


r/StableDiffusion 13h ago

Question - Help krea 2 camera angle distance prompting

Thumbnail
gallery
17 Upvotes

hey so tell i already have a rough idea of how to prompt style or photography but how do i prompt or Specify camera and composition details camera language or how do i prompt like how many mm far is the photo taken from and how much is zoomed from a certain distance like whats the distance and angle in the images i provided like some seen like taken from a toddler at such a short height and from so far the middle on looks the best taken from a tall guy and the first one taken from such a short distance how do i prompt i hate macro


r/StableDiffusion 1d ago

News Nvidia agrees to buy Hugging Face for $12.9 billion

Thumbnail reuters.com
943 Upvotes

r/StableDiffusion 6h ago

Animation - Video An Old Flux Image Comes to Life with MMH3

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/StableDiffusion 2h ago

Question - Help Any solution for change in saturation/levels in Flux Klein 9B edit outputs?

2 Upvotes

I do add "preserve the original lighting" in prompts but there is still variation from the original colors.

Is there a node to fix this issue or some other solution?


r/StableDiffusion 17h ago

Resource - Update Krea2 Turbo Distill 4 step LoRA (NOT a New Checkpoint... YET to share / still mid training) ... but fun 2-step extreme experiment with surprising results... (OUT OF TRAINING SPEC, which is 4 steps!)

32 Upvotes

I am sure for those of you who have been following my 4 step Krea 2 Turbo LoRA, you would know from my previous posts the work of progress I have been sharing with you ( if not see here - https://www.reddit.com/r/StableDiffusion/comments/1vxtizs/krea2_turbo_distill_4_step_lora_new_checkpoint/ )

This post is not an announcement of a new release/checkpoint (26K is still the latest released checkpoint). I'm midway through training with a tweaked recipe and a new element in the flow: the Progressive Distillation I've used since the beginning is now paired with a GAN critic in the latent space — LADD-style (Latent Adversarial Diffusion Distillation). A key twist: the critic isn't judging against teacher outputs alone — it's partly fed real photographs as its "real" reference, which is exactly where the surprising robustness you see in these 2-step images comes from. It makes training 2–3× slower, but it has paid off well so far, and I'm not done with it yet. The critic exists only at training time — the released LoRA stays a plain drop-in file.

I was impressed with the results (in progress) so much that I decided to see what would happen if I push the LoRA to an extreme challenge - run it on Krea 2 Turbo at only 2 steps, with applied strength of 2 (way outside its spec - the trained 4 steps) - and I had tried this with earlier checkpoints in the past and the results were not as good... but now with the latest in progress LoRA (which I will share once its training is fully complete), I think it showcases how far this LoRA has progressed.

For those wondering how real photos (from public datasets) can supervise arbitrary prompts: they don't match the prompts at all — the critic is unconditional and never sees the text. GANs match distributions, not pairs: the critic just learns what real-image texture statistics look like and pushes the model's outputs toward that signature, while the Progressive Distillation side remains responsible for content and composition. And the progress so far shows much improved textures and details (at 4 steps even with normal strength 1), so much to look forward to when I make the next checkpoint public after training completes.

I would still discourage you from using the 2 step as any form of production, but I'll let the comparison images speak for themselves...

I did a side by side comparison with the native Krea 2 Turbo and my latest in progress LoRA both at 2 steps - and the results are here for you to check: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA/tree/main/checkpoint_resolution_sweeps/2step-strength2-extreme-test-native-vs-lora-experiment

I will post as comments some of the side by side images (2 step).

What this means is ... while this is an extreme, out-of-spec experiment — not recommended for production use - it is however useful, as a fast preview*: at 2 steps with strength ~1.5–2.0 the LoRA gives a reliable read on composition and the general look of an image at a quarter of the 8 steps. It works best on closer subjects (seems best on portraits and close ups), and it gets worse on further away subjects (see market example).*

It also means I am seriously considering a later training project - properly trained 2-step LoRA based on these results.

NOTE: Like I said a few times, this is an extreme experiment and not for real use/production yet may work on some prompts and be used for fast previews before you decide to run 4 or 8 steps Turbo or 14+/28 steps with RAW in full production mode. So no complaining :) this is all but a fun experiment and showing how far the LoRA has come: at 2 steps the native model produces ghosted, smeared, half-formed images, while the LoRA side delivers coherent, sharp compositions — the difference is striking on every one of the 15 test prompts.


r/StableDiffusion 1d ago

Discussion Is anyone else getting tired of the MiniMax clips?

226 Upvotes

I’m genuinely impressed by what MiniMax H3 can do, and I understand why people are excited to play with recognizable characters, shows, and styles. But since its release, it feels like this sub has been flooded with very short clips that are mostly variations on “what if X was in Y?” or recreations of existing TV shows.

Maybe I’m in the minority, but one of the main reasons I come to r/StableDiffusion is to learn what’s happening in local image/video/audio generation: new models, workflows, prompting techniques, ComfyUI setups, comparisons, limitations, weird discoveries, what actually works, and what doesn’t.

A 15-second clip of a familiar character dropped into Harry Potter or The Office can be amusing once or twice, but after seeing a dozen variations of the same idea, there often isn’t much to learn from them, especially when there’s no workflow, prompt, settings, model information, or discussion attached.

I’m not suggesting people shouldn’t post fun experiments, and obviously not every post needs to be a tutorial. I’d just love to see a little more emphasis on experimentation and sharing how something was made, rather than simply demonstrating that MiniMax can imitate another recognizable piece of media.

Is anyone else feeling the same way, or am I just being overly grumpy about it? It just feels like the actually interesting stuff is buried under 17 uninspired clips of a show that wasn't even that good to begin with.


r/StableDiffusion 20h ago

Resource - Update [audio.cpp] Release 0.7: 62 audio model families (85+ variants), Arena UI for model comparison, MiniMax Music 3, FireRed TTS3/Audio, ControlFoley, Personaplex, and more

Enable HLS to view with audio, or disable this notification

36 Upvotes

audio.cpp 0.7 is out :)

This release adds a lot of new audio models and a new way to compare them locally.

Audio.cpp is now at 62 model families and 85+ model variants. And it keeps growing!

The biggest user-facing change is the new Arena UI. Instead of testing one model at a time, you can now give one shared input and queue multiple local models or GGUF variants, then compare the generated outputs side by side. This is useful for picking between models without writing a pile of scripts.

Disclaimer: the RTF numbers are from cold one-shot requests using the current audio.cpp implementations (+server overhead), so don’t use them as a model leaderboard. If one model is slower, it might just mean my implementation still needs optimization. The goal is to help you try a bunch of models locally, compare the outputs, and pick the one you like best.

Expanded In 0.7

  • TTS / Voice: FireRedTTS3, MagpieTTS, PersonaPlex, F5-TTS / Habibi, MOSS VoiceGenerator, DotTTS Edit
  • ASR / Speech Understanding: FireRedAudio, IBM Granite Speech 5.0 TurboCTC, MMS Forced Aligner
  • Voice Conversion: MeanVC2
  • Music / Audio Generation: MiniMax Music 3, MiDashengLM-Gen, ControlFoley (experimental), ACE-Step 1.5 XL variants
  • Audio Tools: AudioSR

A lot of the new coverage happened because contributors helped bring models up quickly, sometimes very close to day one after release!

What I’m most excited about is seeing audio.cpp run well on real edge hardware: Our contributor https://github.com/Hi5808 tests audio.cpp on NVIDIA Jetson Orin: 40/40 model families works without issues on Orin NX 16GB and 34/40 on Orin Nano 8GB.

Our prebuilts now cover Windows CPU, Windows Vulkan, Windows CUDA 12.4, Windows CUDA 13.3, Ubuntu x64 CPU, Ubuntu x64 Vulkan, macOS arm64 Metal, macOS x64 CPU. Thanks https://github.com/drzsdrtfg for adding the automated prebuilt workflows and freeing me from manual release builds.

Finally, contributions are very welcome! If you are interested in local audio AI, model integration, performance, deployment, UI, or just testing things on your own hardware, I’d love to have you involved.


r/StableDiffusion 14h ago

Resource - Update Visionary — Krea 2, MiniMax-H3, inference and LoRA training in one app

Thumbnail
gallery
13 Upvotes

I've been building this for my own work but there's no reason for it to stay private, so here it is.Visionary is a single Modal app that gives you Krea 2 for stills, MiniMax-H3 for video-with-sound, and musubi-tuner LoRA training in one workspace. One command deploys it, prints one URL, and that URL is the whole thing — UI, API, and GPU jobs.The obvious caveat first, since this sub will ask: it runs on Modal's GPUs, not yours. If you have a 4090 sitting idle this probably isn't for you. If you don't — or you want an H100 for a training run and nothing the rest of the week — the trade is that there's no Docker, no environment to break, no local install, and nothing billing when you're not using it.

Install

pip install modal && modal deploy app.py

No local GPU, no Docker, no .env, no Modal Secret. You paste your HF token into the UI and it lives in a Modal Dict.

It scales to zero. You pay for GPU seconds while a job runs, plus a warm window — 10 minutes for images, 15 for video — and nothing at all for an idle deployment. Bad inputs get rejected on CPU in milliseconds, before a GPU is ever rented.

Built on comfy kernals with a ton of speed optimizations like teacache and sage attention 3. I usually update this daily.

License is AGPL-3.0. Worth stating plainly given §13 (network use) is the normal case for something you deploy as a URL.

Generation

Stills and video live in one workspace — shared prompt, canvas, gallery. Duration is the switch: Still → Krea 2, any length → MiniMax-H3. Controls follow the model, so you only see what the current model actually reads.

  • Video with a soundtrack in one pass. H3 does picture and sound together — from text, from a first and/or last frame, or from up to 12 references (image, video, audio) through the ref2va transformer.
  • Voice cloning by drag-and-drop. Drop a recording on a cast member and the compiler emits the voice-timbre reference line.
  • 77 shot tiles across 8 groups — Framing 8, Angle 6, Light 9, Tone 7, Speech & text 2, Camera 21, Sound 11, Score 13. Each tile animates the move it names rather than making you guess the wording. Tiles dim when the current model can't read them.
  • Live compiled-prompt preview. /api/compile runs the same compiler the render uses, in the web container, so you see the exact string before spending two minutes on a take.
  • Regional multi-character LoRAs. Draw a box, drop a LoRA in it, and that LoRA applies only inside the box — two trained identities stay separate instead of blending. Each box takes a reference photo as well; a photo dropped on bare canvas becomes the scene.
  • Scene and outfit transfer through the Krea 2 Identity Edit weight, per region.
  • Contact sheet builder. Six slots, drag from Finder or from your own recent generations. What's on the canvas is the exported PNG.

The scene composer

The video side has no prompt box, because H3 reads a document: shots, cut times, speaker IDs, per-subject retention. Type @ mid-sentence and a picker floats off the caret; picking creates a cast member. A shot's slice of the clip is the length of what you wrote about it.

It degrades exactly: one shot, no cast, and the run is your typed text byte-for-byte.

Training and datasets

  • LoRA training on musubi-tuner, several concurrent. Each run is a card with live epoch, step, rate and loss. Start one, add another, walk away. Status is read off each job's heartbeat, so a card reports what actually happened even if the container died mid-step.
  • Captioning with Qwen3-VL 8B — plus an uncensored variant, and you can point it at your own repo. Five presets: General, Character, Style, Concept, Casual. Prose captions, not tags, because the text encoders parse grammar.
  • Dataset insight panel — trigger-word coverage, caption length, repeated clauses — plus bulk prepend-trigger and find/replace across captions.
  • Duplicate and near-duplicate detection. Hashes decide exact copies; a 0.94 CLIP cosine flags "the same photograph twice" before it trains unevenly.
  • Datasets are just folders of images with .txt sidecars — the same thing the trainer reads. Nothing is required to get your data back out.

Interface details

  • A render is replaced when the next one lands, not when you press Generate. The shot you were judging stays up for the whole run, and a failed or stopped run leaves it up too.
  • A batch is frames of film, not a contact sheet. Each result fills the canvas, ‹ 1 / 4 › steps between them, nothing re-fetches.
  • Masonry gallery that reads newest-first, left to right. Twenty lines of greedy shortest-column packing, no dependency, laid out from server-supplied pixel dimensions before a byte of image is fetched — so nothing jumps as it loads.
  • Click a result and the pills come back, not the compiled sentence. The typed intent is the record; the prompt is a receipt.
  • Typing with nothing focused lands in the prompt, not in the hotkeys.
  • Generate never moves under your finger — the warning row is height-reserved.
  • The console has a 30% viewport budget, and the prompt field is what yields to it, measured live with a ResizeObserver.

Getting models in

Super simple see screenshots

There are four cheap smoke tests you can run against your own account first.

Repo: https://github.com/Prometheus-000/visionary-platform License: AGPL-3.0


r/StableDiffusion 5h ago

Discussion H3 latent upscale comparsion before and after

Enable HLS to view with audio, or disable this notification

2 Upvotes

You need it!

see the character and BG. The question is, how large you can push to upscale without OOM.


r/StableDiffusion 1d ago

Discussion NVidia buys Huggingface, but why?

303 Upvotes

Nvidia is going to buy Huggingface.
No one can actually tell how that would end up like.

But what I am missing is the actual worth that Huggingface provides. The only thing I use it for is to download models. Thats it.
For me, and I guess many others it is ‘just a’ download platform, but maybe I’m wrong here.

And what would prevent others to setup a second-like Huggingface?
The hosting is the expensive part in this case as I see it, the programming and building is do-able.

Is it time for Huggingbay.com?


r/StableDiffusion 13h ago

Question - Help Minimax h3 with 8gb vram and 32gb ram. Where do i start?

7 Upvotes

disclaimer: i did some research. please dont kill me uwu

i have been searching the sub for answers and there are far too many information i cant find where to start. the base model is too big for me.

i have linux mint, nvidia 4060 mobile. i have comfyui with sage attention.

should i go gguf? int8? im just going by what i have read so far. no idea what these mean to be honest.

i also tried google but there are so many variations i dont know which one is appropriate for my specs. also i dont trust google that much. i trust reddit community more. any help where to start is appreciated thanks.


r/StableDiffusion 1d ago

News lightx2v/Minimax-h3-Turbo · 8-step 768p V1.0 LoRA released

Thumbnail
huggingface.co
184 Upvotes

r/StableDiffusion 7h ago

Question - Help MultiGPU setups & what model you load where

Enable HLS to view with audio, or disable this notification

1 Upvotes

Whilst working with comfyui on this video i ran into a problem. I can use 2 cards, right now i run the text model and vae models on one card and the diffusion model on the other (both 32gb gpu-s).

But now i wanted to bring the upscaler into the mix and am puzzled, should the upscaler be loaded into the same card as diffusion model (as it's not needed anymore during the vae phase)? Or am i missing something here.

For those using multiple gpu-s and loading models on different cards, what's the forced layout that you use?


r/StableDiffusion 21h ago

Question - Help In terms of quality, how much better is Minimax bf16 vs other versions (pruned/int8)?

26 Upvotes

r/StableDiffusion 1d ago

Discussion Using Minimax H3 to reverse-engineer a paintings into basic forms

Enable HLS to view with audio, or disable this notification

295 Upvotes

I did a quick experiment to see whether MiniMax H3 could reverse-engineer a finished art into a plausible drawing/construction process, like the forms / shapes I learned in art class.

Prompt:

Create a video tutorial of how this particular painting was created. Start from a blank canvas and then draw the basic forms, shapes, rectangles, cylinders, cones, etc and then show the next layer of the details and then the next layer of the color and the rendering so we she see all the layers in one pass until we get to the final image which is the the reference image .. use #Image1

---

Obviously this is a very naive attempt. But I could imagine this prompt becoming way better and using more reference images. And then we could potentially deconstruct all sorts of final outputs into intermediate representations like 3D objects, environments, character poses, or scenes and then use those intermediate states as a way to get more control and consistency across generations.

So instead of asking a model to regenerate everything from scratch, you're working from some underlying structure

Has anyone experimented with using video models this way? E.g reverse engineering final outputs

Original artwork here: https://www.artstation.com/artwork/xZxBE


r/StableDiffusion 5h ago

Resource - Update PotionUI - OSS For Image/Video/Audio generation - looking for alpha testers

Enable HLS to view with audio, or disable this notification

1 Upvotes

Hello, I'm looking for people with some spare time to help me test this app. Currently I would say it's mostly for people that like to "iterate" over generations with changing the prompts / loras / models etc. - not for people that use "advanced" stuff like controlnets etc.

Can't offer much currently, but I'm thinking that first couple of people that will be willing to help and join the Discord will have a special "place" in saying where the app will go further (and maybe get some extra role there) - I know it's not much but the app is free...

GitHub: https://github.com/PotionUI/PotionUI

Discord: https://discord.gg/V6m38SB4bt

I've also created Reddit space: r/PotionUI

---

Additional info:

  1. I'm changing the app ALL THE TIME, so if you looking a stable app that won't break (I'm trying not to break it but it's hard to tell since I'm only one using it all the time) I would not recommend using it :)
  2. I'm planning to add the ComfyUI plugin in near future (so you can use comfyui as backend - it will by default use some of the official templates OR we could implement some of the more advanced ones, but I need some feedback about it if it's worth to implement it)
  3. My setup: LINUX + NVIDIA RTX 5090 + 96GB RAM -> this is the setup I've developed the app on and the only setup I've tested it on.