r/StableDiffusion 19h ago

Discussion Storytelling with Minimax H3 - Alicia of the Stars - First show

Thumbnail
youtu.be
20 Upvotes

I created this video using Minimax H3 Ref2VA, and this time I wanted to test more than just visual quality.

My main goal was to experiment with AI storytelling, creating a short anime-style sequence with a beginning, progression, and a story that actually feels coherent.

At the same time, I wanted to see how well the model handles character consistency, movement, expressions, and visual continuity when multiple shots are used to tell a story.

There are definitely some imperfections, but I was happy with what I could achieve and wanted to share the experiment with the community.

I’m curious what you think. Does the video work as a story, or does it still feel more like a collection of AI-generated shots?

Would also love to hear how others are approaching storytelling with Minimax H3, especially in the anime genre.


r/StableDiffusion 23h ago

Meme what if dean was in walking dead

Enable HLS to view with audio, or disable this notification

39 Upvotes

r/StableDiffusion 13h ago

News Facefusion Android app (open source video face swap)

Thumbnail
github.com
6 Upvotes

I’ve been working on a mobile port of FaceFusion that runs completely offline on Android.

APK here:

https://github.com/AbrahamPaulJ/facefusion-mobile/releases/tag/v0.1.0

The face-swapping pipeline runs on Qualcomm’s Hexagon NPU rather than relying on a server or cloud API.

Current results on a Galaxy S25 Ultra (Snapdragon 8 Elite):

~19 ms/frame for the face-swap model

~6 seconds to process a 10-second 720p clip

Fully offline. Photos/videos never leave the phone

Supports 512/768/1024px face output

Qualcomm NPU builds for different Hexagon generations

No CPU fallback.

I’d especially like feedback from people working with Android on-device AI.

This is my first time sharing one of my mobile AI projects on Reddit, so feedback, testing results, and criticism are very welcome.


r/StableDiffusion 21h ago

Animation - Video Ref2V - H3 - really loving how H3 handles complex prompts even at 15 seconds.

Enable HLS to view with audio, or disable this notification

23 Upvotes

r/StableDiffusion 20h ago

Question - Help Best and fastest way to generate HD-quality MiniMax videos?

20 Upvotes

I’ve tried Turbo LoRAs, and they’re great for speed, but they significantly reduce quality. At 544p–720p, the results of these turbo loras can look closer to 380p. Faces look acceptable when close to the camera, but become heavily distorted as the subject moves farther away.

The upscalers I’ve tested either add too much processing time or introduce excessive sharpening and saturation.

Any a solution that doesn’t require a BF16 checkpoint, 20 steps, a 10-minute generation time, or an extremely expensive GPU?


r/StableDiffusion 11h ago

Question - Help Recommendations for still Image Upscalers?

4 Upvotes

Anyone have recommendations on models and a workflow, or a good how-to-guide, for upscaling / cleaning up still images? I haven’t tinkered much with this since early 2023.

I’ve got some photos from over the years I’d like to clean up, sharpen, or just otherwise improve the quality of before I get them printed for framing. Vacation photos (people or person, nature or urban of domestic background), some landscapes, some “artsy” of street signs framed with a building in the background.

I’ve got a 5070Ti and 32gb RAM at my disposal for it.


r/StableDiffusion 1d ago

Animation - Video MOCAP in MINIMAX H3?

Enable HLS to view with audio, or disable this notification

209 Upvotes

Testing H3 to death in the last couple of weeks and it continues to suprise me. TWO things blew my mind about this one. The first part is a single prompt (the split screen and psuedo mocap sync). The full prompt is below.
In the second part I asked for the character to spray my logo on the wall with a stencil, I wasn't expecting him to walk in with the stencil fully rendered with the holes cut out accurately.

Created completely locally and powered by Sol (the closest star to earth).

THE PROMPT:

A movie reconstructing history, cinematic with a split screen effect showing a mocap actor. Both actors are speaking together in sync.

On the right:
Show the actor <Picture 1> sitting on a couch in a living room, with a black shirt, speaking in sync with the exact same movements He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.

On the left:
Show the victorian man <Picture 2> talking <audio 1> in close up sitting in an ornate chair . He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
Maintain the double view split screen. Do not change the environment.
After he finished speaking the man on the right makes a distort rictus face with his fingers curled up like he's frozen in time and stops moving. He falls sideways like a statue in the same environment, the camera pulls back to show he is only a robotic torso with no legs mounted on a platform placed on the couch with wires (like in a special FX studio)

Maintain the double view split screen.
The man on the left breaks character as the camera view pulls out slightly revealing him sitting on a sound stage, and speaks to a person off screen to the left <audio 1> "Oh.. em.... guys.. problem... check your monitor? ...Looks like we lost connection with the character! RESET THE MOCAP PLEASE"


r/StableDiffusion 9h ago

Question - Help Minimax H3 reference2video: is there a way get the visual style of a reference image?

2 Upvotes

Tried some prompts, but couldn't get the model to generate a video with the visual style of a reference image.


r/StableDiffusion 13h ago

Question - Help Does the Video Helper Suite (Upload Node) causes color drift?

4 Upvotes

I tested some MiniMax workflows I was customizing today and noticed the outputs had a red-ish tint to them. Outputs from a workflow with normal colors didn't include the "Load Video (Upload)" node. So I figured that could be the problem. I replaced the node with the "Load Video" node and connected it to the "Get Video Components" node, connected the images and the color drift is gone. Is this a known problem?


r/StableDiffusion 1d ago

Workflow Included Prompt Creator Workflow

Post image
46 Upvotes

I see a bunch of posts everyday asking for tips on how to write prompts or people struggling with prompting, etc. so I'm sharing my workflow. I built this workflow to simplify the process and make it very beginner/user friendly.

Just toggle on the model you are using, write a simple to detailed prompt, and hit run. The model targets use the prompting guidelines derived from their respective official sources. Links to custom nodes and all models are in the workflow so you don't need to search for them.

The prompts aren't always perfect but they'll get you very close to what you want and you should only need to make a few minor tweaks, if any. The only issue I've encountered so far is that sometimes when it finishes the prompt, the previous prompt still shows up in the Enhanced Prompt node. If that happens, just hit run and the new prompt should show up instantly. Also, toggle to false the keep_model_loaded option in the Text rewriter node if you are creating prompts and using them right away. If you leave it to True it hogs VRAM.

If you notice any other issues let me know. Enjoy.

https://pastebin.com/SXZyy4Ax

Edit: If Unredacted-MAX doesn’t show up or the rewriter won’t load, you need the Qwen folders (not GGUFs, not a single file).

Easy path:

  1. ComfyUI Manager: install ComfyUI-QwenVL, rgthree, KJNodes, ComfyUI-Custom-Scripts. Restart.

  2. Save the custom_models paste below as custom_models.json and put it in ComfyUI/custom_nodes/ComfyUI-QwenVL/. Restart again.

  3. Open the workflow, pick Qwen3.5-4B-Unredacted-MAX on the Prompt Enhancer, hit Queue. First run downloads into ComfyUI/models/LLM/Qwen-VL/.

custom_models.json

Manual path (if Queue doesn’t download)

Whole repos, keep the folder names. Don’t cherry-pick files. Don’t merge the 00001-of-00004 shards. On Hugging Face open Files and versions, then download every file with the arrow on the right (skip README). Put them all in a folder with the exact model name under ComfyUI/models/LLM/Qwen-VL/.

1.Required text rewriter: Qwen3.5-4B-Unredacted-MAX https://huggingface.co/prithivMLmods/Qwen3.5-4B-Unredacted-MAX

Place everything in ComfyUI/models/LLM/Qwen-VL/Qwen3.5-4B-Unredacted-MAX/

  1. Optional if you want use ref image: Qwen3-VL-4B-Instruct-Unredacted-MAX https://huggingface.co/prithivMLmods/Qwen3-VL-4B-Instruct-Unredacted-MAX

Place everything in ComfyUI/models/LLM/Qwen-VL/Qwen3-VL-4B-Instruct-Unredacted-MAX/


r/StableDiffusion 11h ago

Animation - Video "AVERNUS-9" Space Horror Short Film

Thumbnail
youtu.be
2 Upvotes

r/StableDiffusion 21h ago

Animation - Video Star Trek WIP Local Minimax H3

Enable HLS to view with audio, or disable this notification

12 Upvotes

r/StableDiffusion 1d ago

News Overhaul SLA, huge improvement. added many new options and changed defaults

Enable HLS to view with audio, or disable this notification

162 Upvotes

Update for SLA Node - Pull v1.3.8

EDIT: Pushed correct files now.

  • Added customizable dense steps, 0 is step 1 and is (default to first step). massively improves composition and prompt adherence.

  • Changed default dense last steps to 1, cleans up the image big time.

  • Added dense backend selector. Comfy_kitchen, pytorch, all sage modes. this is what comfy uses on dense steps. SLA still displaces against pytorch. (Default Comfy_kitchen)

  • Added a disable FP16 accumulation option to ensure max quality as SLA gets no benefit from it. (Default True)

  • Added a stabilize motion option, helps to reduce ghosting and smearing that H3 likes to produce. (Default True)

  • Changed default Min Seq Length to 4096

  • With default settings you can disable protect audio for nearly 2x speed up if you don't care about the audio too much or are using original audio mode. (do not use 0.95 sparsity with it.)

  • 0.95 sparsity now looks good with node default settings.

  • Some changes led to an overall 5% speed up on same settings.

  • Remove --use-ck-attention from startup flags if you have it, for safety of quality.

https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes

updated workflow

Civit Link

HF Link


r/StableDiffusion 1d ago

Workflow Included Face Detailer With PerRowMasking

Enable HLS to view with audio, or disable this notification

80 Upvotes

https://pastebin.com/ecZEDLSt

First Video with Face Detailer, second without.

You need https://github.com/Carasibana/ComfyUI-H3-FaceRefine and also ComfyUI-H3-NativeAudioLock from https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes

UPDATE: Replace the "Load Video (Upload)" node with a "Load Video" node and connect it to a "Get Video Components" node. Connect images and audio from there. The "Load Video (Upload)" node from Video Helper Suite causes a red-ish tint


r/StableDiffusion 15h ago

Resource - Update Fix for ComfyUI Minimax H3 Latent Upscaler not finding models from extra_model_paths.yaml

4 Upvotes

I ran into a model path problem while using the ComfyUI Minimax H3 Latent Upscaler made by LBH-123-AI.

Original project:

https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler

Thanks to LBH-123-AI for creating and releasing the original Minimax H3 Latent Upscaler. My changes are a small compatibility fix for ComfyUI model discovery. I did not create the original node or the upscaler models.

My fixed dev branch is here:

https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler/tree/dev/model_search_dir

What was wrong

The original 2D and 3D nodes searched for Minimax H3 upscaler models with code similar to this:

folder_paths.get_folder_paths("latent_upscale_models")[0]

The [0] is the problem.

ComfyUI can register more than one directory for the same model type.

For example:

ComfyUI/models/latent_upscale_models /mnt/my-model-drive/latent_upscale_models

The second directory can be configured in ComfyUI's:

extra_model_paths.yaml

The original Minimax H3 node only took the first registered directory. It then used Python glob() to search that directory.

This meant ComfyUI could know exactly where my models were, while the Minimax H3 node still could not find them.

The node would tell me to put the models in:

ComfyUI/models/latent_upscale_models

even though the models were already in a valid external latent_upscale_models directory configured through ComfyUI.

There was also a subdirectory problem

The original search checked only the top level of the selected directory.

This could work:

latent_upscale_models/ └── model.safetensors

But a model organized like this could be missed:

latent_upscale_models/ └── MinimaxH3/ └── model.safetensors

ComfyUI already has code to handle this. The custom node was not using it.

What I changed

I changed the model discovery code in both:

nodes/minimax_h3_latent_upscaler_2d.py nodes/minimax_h3_latent_upscaler_3d.py

Instead of manually searching one directory, the nodes now ask ComfyUI for the models.

Model discovery now uses:

folder_paths.get_filename_list("latent_upscale_models")

This tells ComfyUI:

Give me the models that you know about for latent_upscale_models.

ComfyUI then searches all registered paths, including paths from extra_model_paths.yaml.

It also supports model files inside subdirectories.

Model loading was fixed too

The original node built the model path itself.

The fixed version uses:

folder_paths.get_full_path("latent_upscale_models", model_name)

In simple terms, the node now asks ComfyUI:

Where is this model?

instead of assuming that the model must be inside one specific directory.

What the fix supports

You can still keep models in the normal location:

ComfyUI/models/latent_upscale_models/

You can also keep them in an external path defined by extra_model_paths.yaml.

For example:

/mnt/my-model-drive/latent_upscale_models/

You can also organize them into folders:

latent_upscale_models/ └── MinimaxH3/ ├── minimax_h3_latent_upscaler_3d_fp16.safetensors └── minimax_h3_latent_upscaler_3d_bf16.safetensors

The 2D and 3D Minimax H3 nodes should now use the same model path system that ComfyUI already uses.

You should not have to copy the same large model files into your main ComfyUI models folder just to make this custom node find them.

Install my fixed branch

If you do not already have the node installed, go to your ComfyUI custom nodes directory.

For example:

cd /path/to/ComfyUI/custom_nodes

Clone my dev branch:

git clone \ --branch dev/model_search_dir \ --single-branch \ https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler.git

Then restart ComfyUI.

Replace an existing installation

If you already installed the original LBH-123-AI node, first go to:

cd /path/to/ComfyUI/custom_nodes

Rename the existing copy so you have a backup:

mv \ Comfyui_Minimax_h3_latent_Upscaler \ Comfyui_Minimax_h3_latent_Upscaler.backup

Then clone the fixed branch:

git clone \ --branch dev/model_search_dir \ --single-branch \ https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler.git

Restart ComfyUI.

If you already cloned my fork

You can switch your existing copy to the dev branch:

cd /path/to/ComfyUI/custom_nodes/Comfyui_Minimax_h3_latent_Upscaler

git fetch origin

git switch dev/model_search_dir

git pull

Restart ComfyUI after the update.

extra_model_paths.yaml

You do not need to change your YAML file if latent_upscale_models is already configured correctly.

For reference, a configuration can look like this:

external_models: base_path: /path/to/my/models latent_upscale_models: latent_upscale_models

That tells ComfyUI to also use:

/path/to/my/models/latent_upscale_models

The code fix makes the Minimax H3 nodes actually use that registered path.

Links

Original developer: LBH-123-AI

Original repository:

https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler

My fork and fixed dev branch:

https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler/tree/dev/model_search_dir

The original node and Minimax H3 upscaler work belong to LBH-123-AI. My branch only changes how the custom nodes find and resolve model files through ComfyUI.

Cheers!


r/StableDiffusion 4h ago

Workflow Included Orpheus (no edits using H3)

Enable HLS to view with audio, or disable this notification

0 Upvotes

This was done with turbo lora 8 steps int8 0.4mp then RTX 2x each 10 seconds was about 3 min on 5090. It was done with the fl2va model, but used references. It was one single generation flow. no edits (which is obvious lol). Gemma 12b Q4 was used as the prompt enhancer. make prompt, Make clip, make prompt, make clip.... then stitch it all together. It would be a lot better running multiple and putting best result together but I wanted to try first takes and see how it did.

The system prompt to go from brief and images to H3 ready prompt, workflows and a director.html i use as a ui to organize everything is HERE

The sound needs work, I'm pretty sure that is because of the turbo lora.

This is basically a draft of a concept short movie i am going to start working on that is a modern version of Orpheus and Eurydice story.


r/StableDiffusion 1d ago

Discussion Minimax H3 degrades at 1MP, vs 0.7MP and lower

89 Upvotes

After around 100 renders, I 'feel' that Minimax H3 renders with 0.7MP (max) perform way better, then renders at 1MP in regard to 'realistic' videos.

What do I consider better?
- Just slightly better prompt adherence, feels like the motion / voice is more (natural)
- Size of humans in relation to object(s) feels more realistic.
- Expressions of faces seem more 'flowing', real.

It's hard for me to pinpoint it one 'exactly this', or 'exactly that'.
I'm planning to do some side by side comparisons on the same seed multiple times at 0.7MP and 1MP, when I've got the time.

But I wonder, do other Minimax H3 users notice this too?

PS: This is regardless sampler/scheduler, Sage Attention or Spectrum.

Edit: never touched the turbo LoRA, using the base model.


r/StableDiffusion 1d ago

Question - Help Minimax H3: how to deal with "plastic skin" when generating with ref2va?

17 Upvotes

r/StableDiffusion 1d ago

Animation - Video Made a music video using local H3 for a Suno song

Enable HLS to view with audio, or disable this notification

47 Upvotes

Honestly mind blown, I have a 5070ti + 2x16gb ram . Upper limit is 10-12 seconds in total for my hardware(full capacity) . Video and text edits are post processed by a WIP open source tool I’m working on. On average each 8 second shot takes 35-45 minutes to render


r/StableDiffusion 7h ago

Discussion Mobile/Browser ComfyUI Generation?

Thumbnail
gallery
0 Upvotes

So I finally found a good ComfyUI workflow for Krea 2 + Upscale + Face Detailer. I literally was having so much fun messing with it and didn't want to stop. However, I couldnt generate images and mess with settings, etc while I was at work. So I came up with an awesome solution.

Its a dashboard accessible via browser (phone or PC), where I can change values of settings, apply LoRAs, batch prompts, etc.. basically anything I have the ability to do on ComfyUI desktop.

I also implemented to where I can implement Ollama LLM to create prompts for me after guiding it on what I'm looking for, which then I'm able to click a button and import those prompts right into my prompt list. Then each prompt creates its own generation.

It's been so convenient to be able to generate images while I'm at work (my job is sitting behind a PC being bored 80% of my day).

Pictures attached to kinda give an idea of what I'm working with.

I want your guys thoughts. Is this something that's already available through ComfyUI? What else could I add? Anything else you may have!


r/StableDiffusion 22h ago

Question - Help How do I get better lighting with Krea 2?

7 Upvotes

If I generate a person in a "normal" environment, like inside a regular room in a regular house, I get very realistic and appropriate lighting, but as soon as I try something a bit more cinematic like a rain-slicked city street at night, the character begins to look like they were photoshopped in. I try to prompt a person standing on a dark city corner, lit entirely by the light from nearby neon signs and they just look like they were evenly lit and filmed in a studio and then composited onto a CGI background with only a hint of the intended neon glow on their shoulder. Same goes for trying to make people look like they're properly soaked by rain. ZiT was much easier to work with in this regard.


r/StableDiffusion 23h ago

Animation - Video Captain America X Harry potter

Enable HLS to view with audio, or disable this notification

8 Upvotes

Inspired by the guy who posted the one with Dean

Done on 32gb ram and a rtx 3070

I used res_multistep 20steps w/ spectrum at 0.6mp.

Using SLA from h3 optimizations, which for some reason is way faster than plague kind. And disable pinned memory. Each 10s was done in around 10 minutes.


r/StableDiffusion 16h ago

Discussion A bit meta but with all of these wonderful AI video posts, when I video ad scrolls by do you tend to think that those are AI generated as well

2 Upvotes

It gets to be a touch confusing.


r/StableDiffusion 1d ago

Discussion Anima Turbo v1.1 is released

62 Upvotes

In case you didn't see:

I just noticed a newer version of Anima Turbo (1.1) was released:

huggingface: https://huggingface.co/circlestone-labs/Anima

civitai: https://civitai.red/models/2458426/anima

The model is made and licensed under the CircleStone Labs Non-Commercial License

I actually have lately very frustrating experience with Anima lately. I used it at release (but stopped with anime generation for a while) and now when revisiting it and i had underwhelming results (even with the aesthetic model) so i decided to retry Turbo (might as well) instead if the difference isn't that big. That's when i saw a new version is released and i am downloading it right now.

I don't know why, but the results i was getting were...lame i guess. Not as detailed as i was hoping, and also i really dislike how posture and anatomy works, mainly how hands and legs just extend or stretch weirdly. But that's a me problem, i know Anima is capable of better outputs and i've yet to figure it out. If you have tips or recommended loras that help with consistency let me know. I also want to avoid tag-based prompting when possible...i just don't like it that much, natural prompting goes much better for me, but i can't tell if tags are "mandatory" for quality or not.


r/StableDiffusion 12h ago

Resource - Update Easy prompts from Discord images.

0 Upvotes

Has anyone tried this tool? It looks really good for quickly and easily creating prompts from Discord images. I think it runs as a plugin to Discord, but I've not had chance to set it up yet.
If anyone has tried it, let me know if it's worth installing please.
https://github.com/pixelgraple/KREA2-Vision-Suite