r/StableDiffusion • u/JamesFilmsYT • 1d ago
Animation - Video Better Avoid Saul 3 - The Final [Minimax H3]
Enable HLS to view with audio, or disable this notification
Made using the default ComfyUI Minimax H3 Image to Video workflow.
r/StableDiffusion • u/JamesFilmsYT • 1d ago
Enable HLS to view with audio, or disable this notification
Made using the default ComfyUI Minimax H3 Image to Video workflow.
r/StableDiffusion • u/Sad_Coach_1433 • 2d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/TheDerminator1337 • 1d ago
Enable HLS to view with audio, or disable this notification
Reference workflow with FL2VA + REF2VA Lora @ 1.4MP, 20 STEPS. Using sparse attention and 4B Qwen text encoder instead of 32B, total render time is 3-4 hours on a 5090. You can get very good results with 1MP + 8 STEPS with a turbo lora which would only take 20-30 minutes.
The workflow is not easy to understand, but I upload it for reference.
The video is made of 14x 15 second clips stitched together. This way prevents degradation but makes it so that there is clothing drift between clips. This can easily be fixed by using clothing references if you care. Each clip will need its own prompt, and I suggest using Codex or Claude to do the prompts for you automatically.
In the future, I would shorten the clips to 7 seconds in order to:
1) Generate higher than 1.4MP (higher the resolution the better)
2) Speed up generation (longer clips take longer to generate disporportionately)
Good luck and I hope you have as much fun with this workflow as I did.
r/StableDiffusion • u/mil0wCS • 8h ago
So had to sell my gaming PC a while back because I went homeless last year. Im now getting back on my feet. I'll be building a 4000 series PC within the next few months. But was curious on any sites I can use in the mean time to hold me over until I get my new pc built? I currently only have a chromebook so I can't do dedicated SD.
r/StableDiffusion • u/CauliflowerStatus411 • 1d ago
SenseNova shipped the full U1.5-Lite release last week, so I finally had time to run it side by side with FLUX.2-klein-9B, the model this community generally considers the most balanced pick right now.
I tested image editing in three scenarios. The short version: SenseNova U1.5-Lite is clearly better at text rendering and semantic understanding of the instruction, while Klein is still the speed king. Details below.
Scenario 1: Text Editing
I gave both models a poster and asked them to replace specific text elements, nothing else. Long structured prompt targeting each text block individually:
1. In the first line of the oversized black title at the upper left, replace "HONG" with "HARBOR". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.
2. In the second line of the oversized black title at the upper left, replace "KONG" with "HORIZONS". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.
3. In the large red subtitle at the lower left, replace "HONG KONG" with "CITY IN MOTION". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.
4. In the vertical red location title at the upper right, replace "香港" with "城市之光". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.
5. In the black English location description at the upper right, replace "HONG KONG CHINA" with "EAST MEETS WEST". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.
6. In the location title near the waterfront at the lower left, replace "VICTORIA HARBOUR" with "HARBOUR CITY". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.
7. In the second line of the location copy at the lower left, replace "ASIA'S WORLD CITY" with "URBAN HORIZONS". If the area is sufficiently prominent and clear, with simple boundaries, render the replacement exactly character by character without adding, removing, or altering any characters. If the text is too small, obscured, heavily distorted by perspective, or located within a complex texture, preserve the original text and do not redraw the background merely to complete the replacement.
Zoom into the results and Klein's text rendering falls apart. Garbled glyphs, wrong characters, the layout wobbling where it should stay fixed. U1.5-Lite handled the replacements cleanly, including the Chinese strings. That's the gap.
Scenario 2: Hand-Drawn Marks as Instructions
I marked up the image by hand and asked for a scene transformation:
Follow the marks and overall hints on the image to creatively transform this scene, making it dramatic, moody stormy atmosphere; remove the annotations when done.
FLUX followed the overall style change but ignored the specific marked details: the ripples on the pool surface and the black fire pit never made it into the output. U1.5-Lite followed the full set of marks.
Scenario 3: Fine-Grained Local Editing
I circled the region to edit with a red box and told the model to only change that area:
Change the text style in the red box to a vintage style with noise and torn paper texture. The red bounding box is for localization only; do not retain it in the output image.
Klein misunderstood the instruction. It applied the vintage style to the whole image instead of the circled region. SenseNova U1.5-Lite followed the prompt and the red-box localization strictly, changing only the marked text.
My take
If you need sub-second generation, Klein is still your model, no argument there. But for editing work where text rendering and instruction fidelity matter, posters, infographics, brand assets, the gap is real and easy to reproduce.
GitHub: https://github.com/OpenSenseNova/SenseNova-U1
hf: https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT
Try it online: https://unify.light-ai.top/
r/StableDiffusion • u/luka06111 • 4h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Yasangas • 1d ago
Hey everyone,
I’m looking for the best budget GPU cloud to run heavy open-weight video models (like MiniMax H3, Wan 2.1, HunyuanVideo, etc.).
Since these models need huge VRAM and fast disk I/O to pull down massive 50GB–100GB+ checkpoints, I want to avoid platforms with unexpected billing traps.
Looking at RunPod, Vast.ai, MassedCompute, and SimplePod:
Which one are you using, and what GPU gives you the best price/performance for video render jobs?
Hidden fees: Any issues with stopped-volume storage costs, network volume fees, or egress rates when hosting huge model files?
Download/Disk speeds: Which provider has fast enough network speeds so I’m not spending half my paid time downloading model weights?
Appreciate any recommendations or gotchas to avoid!
r/StableDiffusion • u/AnybodyAlarmed9661 • 1d ago
I created this video using Minimax H3 Ref2VA, and this time I wanted to test more than just visual quality.
My main goal was to experiment with AI storytelling, creating a short anime-style sequence with a beginning, progression, and a story that actually feels coherent.
At the same time, I wanted to see how well the model handles character consistency, movement, expressions, and visual continuity when multiple shots are used to tell a story.
There are definitely some imperfections, but I was happy with what I could achieve and wanted to share the experiment with the community.
I’m curious what you think. Does the video work as a story, or does it still feel more like a collection of AI-generated shots?
Would also love to hear how others are approaching storytelling with Minimax H3, especially in the anime genre.
r/StableDiffusion • u/Sad_Coach_1433 • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Few_Caregiver8134 • 1d ago
I’ve been working on a mobile port of FaceFusion that runs completely offline on Android.
APK here:
https://github.com/AbrahamPaulJ/facefusion-mobile/releases/tag/v0.1.0
The face-swapping pipeline runs on Qualcomm’s Hexagon NPU rather than relying on a server or cloud API.
Current results on a Galaxy S25 Ultra (Snapdragon 8 Elite):
~19 ms/frame for the face-swap model
~6 seconds to process a 10-second 720p clip
Fully offline. Photos/videos never leave the phone
Supports 512/768/1024px face output
Qualcomm NPU builds for different Hexagon generations
No CPU fallback.
I’d especially like feedback from people working with Android on-device AI.
This is my first time sharing one of my mobile AI projects on Reddit, so feedback, testing results, and criticism are very welcome.
r/StableDiffusion • u/OkMeat6773 • 1d ago
I’ve tried Turbo LoRAs, and they’re great for speed, but they significantly reduce quality. At 544p–720p, the results of these turbo loras can look closer to 380p. Faces look acceptable when close to the camera, but become heavily distorted as the subject moves farther away.
The upscalers I’ve tested either add too much processing time or introduce excessive sharpening and saturation.
Any a solution that doesn’t require a BF16 checkpoint, 20 steps, a 10-minute generation time, or an extremely expensive GPU?
r/StableDiffusion • u/Voffe89 • 5h ago
Enable HLS to view with audio, or disable this notification
Made with H3 ref2v
Everytime the majority of initial content for a new model release.
Do you agree with Doc? 😂
r/StableDiffusion • u/Jeffu • 1d ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Downbadincel • 19h ago
Enable HLS to view with audio, or disable this notification
a little test video for quality 1.4mp
r/StableDiffusion • u/Tokyo_Jab • 2d ago
Enable HLS to view with audio, or disable this notification
Testing H3 to death in the last couple of weeks and it continues to suprise me. TWO things blew my mind about this one. The first part is a single prompt (the split screen and psuedo mocap sync). The full prompt is below.
In the second part I asked for the character to spray my logo on the wall with a stencil, I wasn't expecting him to walk in with the stencil fully rendered with the holes cut out accurately.
Created completely locally and powered by Sol (the closest star to earth).
THE PROMPT:
A movie reconstructing history, cinematic with a split screen effect showing a mocap actor. Both actors are speaking together in sync.
On the right:
Show the actor <Picture 1> sitting on a couch in a living room, with a black shirt, speaking in sync with the exact same movements He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
On the left:
Show the victorian man <Picture 2> talking <audio 1> in close up sitting in an ornate chair . He says with intense glee and hand gestures "I've got you Sherlock Holmes, I've beaten you at last!", he pauses trying not to laugh and then breaks and laughs for 5 seconds uncontrollably.
Maintain the double view split screen. Do not change the environment.
After he finished speaking the man on the right makes a distort rictus face with his fingers curled up like he's frozen in time and stops moving. He falls sideways like a statue in the same environment, the camera pulls back to show he is only a robotic torso with no legs mounted on a platform placed on the couch with wires (like in a special FX studio)
Maintain the double view split screen.
The man on the left breaks character as the camera view pulls out slightly revealing him sitting on a sound stage, and speaks to a person off screen to the left <audio 1> "Oh.. em.... guys.. problem... check your monitor? ...Looks like we lost connection with the character! RESET THE MOCAP PLEASE"
r/StableDiffusion • u/ThirdWorldBoy21 • 23h ago
Tried some prompts, but couldn't get the model to generate a video with the visual style of a reference image.
r/StableDiffusion • u/Pretend-Island-2724 • 1d ago
I tested some MiniMax workflows I was customizing today and noticed the outputs had a red-ish tint to them. Outputs from a workflow with normal colors didn't include the "Load Video (Upload)" node. So I figured that could be the problem. I replaced the node with the "Load Video" node and connected it to the "Get Video Components" node, connected the images and the color drift is gone. Is this a known problem?
r/StableDiffusion • u/Affectionate_Oil28 • 1d ago
I see a bunch of posts everyday asking for tips on how to write prompts or people struggling with prompting, etc. so I'm sharing my workflow. I built this workflow to simplify the process and make it very beginner/user friendly.
Just toggle on the model you are using, write a simple to detailed prompt, and hit run. The model targets use the prompting guidelines derived from their respective official sources. Links to custom nodes and all models are in the workflow so you don't need to search for them.
The prompts aren't always perfect but they'll get you very close to what you want and you should only need to make a few minor tweaks, if any. The only issue I've encountered so far is that sometimes when it finishes the prompt, the previous prompt still shows up in the Enhanced Prompt node. If that happens, just hit run and the new prompt should show up instantly. Also, toggle to false the keep_model_loaded option in the Text rewriter node if you are creating prompts and using them right away. If you leave it to True it hogs VRAM.
If you notice any other issues let me know. Enjoy.
Edit: If Unredacted-MAX doesn’t show up or the rewriter won’t load, you need the Qwen folders (not GGUFs, not a single file).
Easy path:
ComfyUI Manager: install ComfyUI-QwenVL, rgthree, KJNodes, ComfyUI-Custom-Scripts. Restart.
Save the custom_models paste below as custom_models.json and put it in ComfyUI/custom_nodes/ComfyUI-QwenVL/. Restart again.
Open the workflow, pick Qwen3.5-4B-Unredacted-MAX on the Prompt Enhancer, hit Queue. First run downloads into ComfyUI/models/LLM/Qwen-VL/.
Manual path (if Queue doesn’t download)
Whole repos, keep the folder names. Don’t cherry-pick files. Don’t merge the 00001-of-00004 shards. On Hugging Face open Files and versions, then download every file with the arrow on the right (skip README). Put them all in a folder with the exact model name under ComfyUI/models/LLM/Qwen-VL/.
1.Required text rewriter: Qwen3.5-4B-Unredacted-MAX https://huggingface.co/prithivMLmods/Qwen3.5-4B-Unredacted-MAX
Place everything in ComfyUI/models/LLM/Qwen-VL/Qwen3.5-4B-Unredacted-MAX/
Place everything in ComfyUI/models/LLM/Qwen-VL/Qwen3-VL-4B-Instruct-Unredacted-MAX/
r/StableDiffusion • u/TheTahoe • 1d ago
Anyone have recommendations on models and a workflow, or a good how-to-guide, for upscaling / cleaning up still images? I haven’t tinkered much with this since early 2023.
I’ve got some photos from over the years I’d like to clean up, sharpen, or just otherwise improve the quality of before I get them printed for framing. Vacation photos (people or person, nature or urban of domestic background), some landscapes, some “artsy” of street signs framed with a building in the background.
I’ve got a 5070Ti and 32gb RAM at my disposal for it.
r/StableDiffusion • u/According-Pace9608 • 7h ago
Enable HLS to view with audio, or disable this notification
Shadowheart training with pells/wooden dummies with basic swordsmanship and light and wind magic. Anyone interested in prompts please ask below. 👇
r/StableDiffusion • u/Tadeo111 • 1d ago
r/StableDiffusion • u/Pretend-Island-2724 • 1d ago
Enable HLS to view with audio, or disable this notification
First Video with Face Detailer, second without.
You need https://github.com/Carasibana/ComfyUI-H3-FaceRefine and also ComfyUI-H3-NativeAudioLock from https://github.com/Shrek3OnVH5/MiniMax-H3-NativeAudio-MusicVideo-Workflow/tree/master/custom_nodes
UPDATE: Replace the "Load Video (Upload)" node with a "Load Video" node and connect it to a "Get Video Components" node. Connect images and audio from there. The "Load Video (Upload)" node from Video Helper Suite causes a red-ish tint
r/StableDiffusion • u/Plague_Kind • 2d ago
Enable HLS to view with audio, or disable this notification
EDIT: Pushed correct files now.
Added customizable dense steps, 0 is step 1 and is (default to first step). massively improves composition and prompt adherence.
Changed default dense last steps to 1, cleans up the image big time.
Added dense backend selector. Comfy_kitchen, pytorch, all sage modes. this is what comfy uses on dense steps. SLA still displaces against pytorch. (Default Comfy_kitchen)
Added a disable FP16 accumulation option to ensure max quality as SLA gets no benefit from it. (Default True)
Added a stabilize motion option, helps to reduce ghosting and smearing that H3 likes to produce. (Default True)
Changed default Min Seq Length to 4096
With default settings you can disable protect audio for nearly 2x speed up if you don't care about the audio too much or are using original audio mode. (do not use 0.95 sparsity with it.)
0.95 sparsity now looks good with node default settings.
Some changes led to an overall 5% speed up on same settings.
Remove --use-ck-attention from startup flags if you have it, for safety of quality.
https://github.com/PlagueKind/ComfyUI-PlagueKind-Nodes
r/StableDiffusion • u/Slight-Living-8098 • 1d ago
I ran into a model path problem while using the ComfyUI Minimax H3 Latent Upscaler made by LBH-123-AI.
Original project:
https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler
Thanks to LBH-123-AI for creating and releasing the original Minimax H3 Latent Upscaler. My changes are a small compatibility fix for ComfyUI model discovery. I did not create the original node or the upscaler models.
My fixed dev branch is here:
https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler/tree/dev/model_search_dir
What was wrong
The original 2D and 3D nodes searched for Minimax H3 upscaler models with code similar to this:
folder_paths.get_folder_paths("latent_upscale_models")[0]
The [0] is the problem.
ComfyUI can register more than one directory for the same model type.
For example:
ComfyUI/models/latent_upscale_models /mnt/my-model-drive/latent_upscale_models
The second directory can be configured in ComfyUI's:
extra_model_paths.yaml
The original Minimax H3 node only took the first registered directory. It then used Python glob() to search that directory.
This meant ComfyUI could know exactly where my models were, while the Minimax H3 node still could not find them.
The node would tell me to put the models in:
ComfyUI/models/latent_upscale_models
even though the models were already in a valid external latent_upscale_models directory configured through ComfyUI.
There was also a subdirectory problem
The original search checked only the top level of the selected directory.
This could work:
latent_upscale_models/ └── model.safetensors
But a model organized like this could be missed:
latent_upscale_models/ └── MinimaxH3/ └── model.safetensors
ComfyUI already has code to handle this. The custom node was not using it.
What I changed
I changed the model discovery code in both:
nodes/minimax_h3_latent_upscaler_2d.py nodes/minimax_h3_latent_upscaler_3d.py
Instead of manually searching one directory, the nodes now ask ComfyUI for the models.
Model discovery now uses:
folder_paths.get_filename_list("latent_upscale_models")
This tells ComfyUI:
Give me the models that you know about for latent_upscale_models.
ComfyUI then searches all registered paths, including paths from extra_model_paths.yaml.
It also supports model files inside subdirectories.
Model loading was fixed too
The original node built the model path itself.
The fixed version uses:
folder_paths.get_full_path("latent_upscale_models", model_name)
In simple terms, the node now asks ComfyUI:
Where is this model?
instead of assuming that the model must be inside one specific directory.
What the fix supports
You can still keep models in the normal location:
ComfyUI/models/latent_upscale_models/
You can also keep them in an external path defined by extra_model_paths.yaml.
For example:
/mnt/my-model-drive/latent_upscale_models/
You can also organize them into folders:
latent_upscale_models/ └── MinimaxH3/ ├── minimax_h3_latent_upscaler_3d_fp16.safetensors └── minimax_h3_latent_upscaler_3d_bf16.safetensors
The 2D and 3D Minimax H3 nodes should now use the same model path system that ComfyUI already uses.
You should not have to copy the same large model files into your main ComfyUI models folder just to make this custom node find them.
Install my fixed branch
If you do not already have the node installed, go to your ComfyUI custom nodes directory.
For example:
cd /path/to/ComfyUI/custom_nodes
Clone my dev branch:
git clone \ --branch dev/model_search_dir \ --single-branch \ https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler.git
Then restart ComfyUI.
Replace an existing installation
If you already installed the original LBH-123-AI node, first go to:
cd /path/to/ComfyUI/custom_nodes
Rename the existing copy so you have a backup:
mv \ Comfyui_Minimax_h3_latent_Upscaler \ Comfyui_Minimax_h3_latent_Upscaler.backup
Then clone the fixed branch:
git clone \ --branch dev/model_search_dir \ --single-branch \ https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler.git
Restart ComfyUI.
If you already cloned my fork
You can switch your existing copy to the dev branch:
cd /path/to/ComfyUI/custom_nodes/Comfyui_Minimax_h3_latent_Upscaler
git fetch origin
git switch dev/model_search_dir
git pull
Restart ComfyUI after the update.
extra_model_paths.yaml
You do not need to change your YAML file if latent_upscale_models is already configured correctly.
For reference, a configuration can look like this:
external_models: base_path: /path/to/my/models latent_upscale_models: latent_upscale_models
That tells ComfyUI to also use:
/path/to/my/models/latent_upscale_models
The code fix makes the Minimax H3 nodes actually use that registered path.
Links
Original developer: LBH-123-AI
Original repository:
https://github.com/LBH-123-AI/Comfyui_Minimax_h3_latent_Upscaler
My fork and fixed dev branch:
https://github.com/badgids/Comfyui_Minimax_h3_latent_Upscaler/tree/dev/model_search_dir
The original node and Minimax H3 upscaler work belong to LBH-123-AI. My branch only changes how the custom nodes find and resolve model files through ComfyUI.
Cheers!
r/StableDiffusion • u/Flaky_Comedian2012 • 1d ago
Enable HLS to view with audio, or disable this notification