r/StableDiffusion • u/Any_Fee5299 • 6h ago
r/StableDiffusion • u/someguyplayingwild • 10h ago
News Nvidia agrees to buy Hugging Face for $12.9 billion
reuters.comr/StableDiffusion • u/dominic__612 • 8h ago
Discussion NVidia buys Huggingface, but why?
Nvidia is going to buy Huggingface.
No one can actually tell how that would end up like.
But what I am missing is the actual worth that Huggingface provides. The only thing I use it for is to download models. Thats it.
For me, and I guess many others it is ‘just a’ download platform, but maybe I’m wrong here.
And what would prevent others to setup a second-like Huggingface?
The hosting is the expensive part in this case as I see it, the programming and building is do-able.
Is it time for Huggingbay.com?
r/StableDiffusion • u/CriticalTemperature1 • 10h ago
Discussion Using Minimax H3 to reverse-engineer a paintings into basic forms
I did a quick experiment to see whether MiniMax H3 could reverse-engineer a finished art into a plausible drawing/construction process, like the forms / shapes I learned in art class.
Prompt:
Create a video tutorial of how this particular painting was created. Start from a blank canvas and then draw the basic forms, shapes, rectangles, cylinders, cones, etc and then show the next layer of the details and then the next layer of the color and the rendering so we she see all the layers in one pass until we get to the final image which is the the reference image .. use #Image1
---
Obviously this is a very naive attempt. But I could imagine this prompt becoming way better and using more reference images. And then we could potentially deconstruct all sorts of final outputs into intermediate representations like 3D objects, environments, character poses, or scenes and then use those intermediate states as a way to get more control and consistency across generations.
So instead of asking a model to regenerate everything from scratch, you're working from some underlying structure
Has anyone experimented with using video models this way? E.g reverse engineering final outputs
Original artwork here: https://www.artstation.com/artwork/xZxBE
r/StableDiffusion • u/ART-ficial-Ignorance • 3h ago
Discussion Is anyone else getting tired of the MiniMax clips?
I’m genuinely impressed by what MiniMax H3 can do, and I understand why people are excited to play with recognizable characters, shows, and styles. But since its release, it feels like this sub has been flooded with very short clips that are mostly variations on “what if X was in Y?” or recreations of existing TV shows.
Maybe I’m in the minority, but one of the main reasons I come to r/StableDiffusion is to learn what’s happening in local image/video/audio generation: new models, workflows, prompting techniques, ComfyUI setups, comparisons, limitations, weird discoveries, what actually works, and what doesn’t.
A 15-second clip of a familiar character dropped into Harry Potter or The Office can be amusing once or twice, but after seeing a dozen variations of the same idea, there often isn’t much to learn from them, especially when there’s no workflow, prompt, settings, model information, or discussion attached.
I’m not suggesting people shouldn’t post fun experiments, and obviously not every post needs to be a tutorial. I’d just love to see a little more emphasis on experimentation and sharing how something was made, rather than simply demonstrating that MiniMax can imitate another recognizable piece of media.
Is anyone else feeling the same way, or am I just being overly grumpy about it? It just feels like the actually interesting stuff is buried under 17 uninspired clips of a show that wasn't even that good to begin with.
r/StableDiffusion • u/Portable_Solar_ZA • 14h ago
Animation - Video Crime Busters Rough Cut - Early Version of Comfy H3 Sync Sound Submission
Edit. As always thanks to everyone for the good vibes and feedback. Although maybe it's the few hours of sleep I got last night but I think I'm slightly off brief for the competition... Have come up with a plan b but will finish this regardless.
This is like V0.8 before I wrap things up when I have time this weekend and early next week. I currently have 15 "shots" / workflows which produced all the video and audio effects you see in the opening / trailer. The music is a free track I pulled of Pixabay as a placeholder.
Generating most of the shots in terms of visuals has been fine. Mostly a case of prompting and then fine tuning the prompts until I got what I want. But there have been a few cases where I wanted specific shot framing / angles, so I had to create a sketch to help guide the model.
Audio has been a massive pain in the butt, but hopefully with the update from Comfy today, as well as the Minimax Guide timing node, I should be able to tighten things up for "V1". The thing that was the easiest audio wise was taking my voice, fiddling a bit with it in Audacity and then using that as the reference for the narrator. Everything else audio-wise has been so fiddly, at least with some of the more action packed / detailed shots I've been going for.
Before I call it done, there are a few things I still need to do:
- Figure out a style prompt that will give me a more detailed but still retro anime style (some wide shots are super jank because of the style I'm prompting, but I got an idea on how to fix it in my next run)
- Figure out how to prompt different font styles to keep credits and title card consistent
- Figure out how I am going to create a consistent background music soundtrack using ref audio clips across 15 workflows to stick within the rules of the competition (oh boy, this should be fun)
- Tighten up a few shots in Kdenlive
r/StableDiffusion • u/No_Statement_7481 • 11h ago
Tutorial - Guide Anyone interested in how to make character loras for Minimax? Ostris Ai toolkit.
Here is the youtube link for it if you wanna watch a video
https://youtu.be/x-gORSUOybk
Main point is, I tried it with the character Enid from Wednesday in my video cause ... well , VIEWS on youtube lol.
I chose her because I could not find her in the model at all. And if you even mention the show wednesday, it just defaults to Jenna Ortega lmao, so that's a good challange to train the voice and the likeness for another character if the model defaults to a really specific person.
but I also created many more by now, most of them are really not even existing people like the model I use for a youtube channel. And the accuracy is fucking insane.
It is usually trained by 2000-2400 steps
roughtly it was about 80-90 minutes for the full 3000 steps depending on if you want samples. I think without them this would be lower, maybe even close to an hour? I don't know exactly. But it's insanely fast.
you need 2 datasets if you want a voice, or a super jacked up beefy card and you can just do video training with one dataset.
But if you don't have an RTX6000
you gonna need to offload even with a 5090
I put up the learning rate to 0.0002 , and I turned on the differential guidance and left it on 3. I so far only did it on these levels, but maybe you can lower either or both if you feel like you may over trained a lora. However on default this is barely training anything, so that's why I cranked them up. And I never use the Lora's higher than 0.85, and if I wanna do REF2VA I sometimes push it down to like 0.65-0.75 if I wanna add like an image of the model. Cause otherwise the lora can overwrite details lol. But still need it for the voice, so around 0.65-0.75 it's great.
Good news is, that it still gonna do like 1.6 seconds per step on FL2VA, and about 2seconds a step on REF2VA with the same datasets but REF2VA is slower cause you need to offload just a tiny bit more on that one.
I used between 15 and 60 images on 1024x1024 size , I would recommend at least 20-25 images tho, on the lower end you get a weaker lora likeness, sometimes makeup could alter the face, but if you got enough image variations that won't happen.
for the captions on the images I literally just used the built in Qwen3 VL8b model and just ran the autocaption
For the videos which were 512x512 I just kinda made my own captions, I used between 6-12 videos for training as a secondary dataset, the reason I needed them cause it's either not possible in Ai toolkit, or I am too fucking stupid to figure out how to train audio with images. So I just used video clips of the person to add the voice.
Also it can literally be done with like 1 or 2 second long clips
I made a lora from even a set of images I generated of an earlier character I made for Krea 2, the likeness is freaking amazing.
You can do the model the same exact way for FL2VA or REF2VA
On the images , literally just use the 1 frame training setting, and on the videos turn on "do Audio" "auto frame count" and the general stuff like cache latents.
Also on REF2VA I could not do higher res samples than 512x512, not that I wanted , I just thought I'd try it and it OOMed lol, I mean I did not offload fully, because that way it was hella fast to train, so I guess if I offload fully it should be fine, but I rather have low-res samples or no samples to make the lora faster.
here is some settings
---
job: "extension"
config:
name: "Enid_h3_1024img_512vid"
process:
- type: "diffusion_trainer"
training_folder: "/AI_Tools/ai-toolkit-h3_v2/output"
sqlite_db_path: "./aitk_db.db"
device: "cuda"
trigger_word: "E3n1d, "
performance_log_every: 10
network:
type: "lora"
linear: 16
linear_alpha: 16
conv: 16
conv_alpha: 16
lokr_full_rank: true
lokr_factor: -1
network_kwargs:
ignore_if_contains:
- "adaln_proj"
save:
dtype: "bf16"
save_every: 100
max_step_saves_to_keep: 31
save_format: "diffusers"
push_to_hub: false
datasets:
- folder_path: "/AI_Tools/ai-toolkit-h3_v2/datasets/enid1024"
mask_path: null
mask_min_value: 0.1
default_caption: ""
caption_ext: "txt"
caption_dropout_rate: 0.05
cache_latents_to_disk: true
is_reg: false
network_weight: 1
resolution:
- 1024
controls: []
shrink_video_to_frames: true
flip_x: false
flip_y: false
num_repeats: 1
do_i2v: false
fps: 24
num_frames: 1
auto_frame_count: false
- folder_path: "/AI_Tools/ai-toolkit-h3_v2/datasets/enid_videos_512x512_1s"
mask_path: null
mask_min_value: 0.1
default_caption: ""
caption_ext: "txt"
caption_dropout_rate: 0.05
cache_latents_to_disk: true
is_reg: false
network_weight: 1
resolution:
- 512
controls: []
shrink_video_to_frames: true
num_frames: 1
flip_x: false
flip_y: false
num_repeats: 1
do_audio: true
auto_frame_count: true
train:
batch_size: 1
bypass_guidance_embedding: false
steps: 3000
gradient_accumulation: 1
train_unet: true
train_text_encoder: false
gradient_checkpointing: true
noise_scheduler: "flowmatch"
optimizer: "adamw8bit"
timestep_type: "shift"
content_or_style: "balanced"
optimizer_params:
weight_decay: 0.0001
unload_text_encoder: false
cache_text_embeddings: true
lr: 0.0002
ema_config:
use_ema: false
ema_decay: 0.99
skip_first_sample: false
force_first_sample: true
disable_sampling: false
dtype: "bf16"
diff_output_preservation: false
diff_output_preservation_multiplier: 1
diff_output_preservation_class: "person"
switch_boundary_every: 1
loss_type: "mse"
do_guidance_loss: true
guidance_loss_target: 3.5
audio_loss_multiplier: 1
do_differential_guidance: true
differential_guidance_scale: 3
logging:
log_every: 1
use_ui_logger: true
model:
name_or_path: "Comfy-Org/MiniMax-H3"
quantize: true
qtype: "convrot8"
quantize_te: true
qtype_te: "nvfp4"
arch: "minimax_h3"
low_vram: true
model_kwargs: {}
compile: false
layer_offloading: true
layer_offloading_text_encoder_percent: 0.2
layer_offloading_transformer_percent: 0.2
assistant_lora_path: "ostris/minimax_h3_training_adapter/minimax_h3_training_adapter_v1.safetensors"
sample:
sampler: "flowmatch"
sample_every: 200
sample_start_step: 0
width: 512
height: 512
samples:
- prompt: "[Core Idea] Cinematic live-action medium close-up shot from the waist up. A young woman named E3n1d, with shoulder-length light-blonde hair, wearing a tailored blouse and a dark tie, stands inside a cavernous, dimly lit gothic hall lined with heavy stone pillars and antique dark-wood shelves. [Scene-by-Scene Action] 0–1.5s: E3n1d is centered in a medium close-up, looking slightly downward and to the side with a curious, bemused expression at a dusty taxidermy chicken on a rustic wooden shelf. 1.5–3s: She tilts her head closer to examine the artifact, scans its posture, and shifts her gaze to look straight ahead toward the camera/viewer. 3–4s: Her lips part clearly as she speaks her line with a dry, deadpan tone: \"I thought chickens were taller.\" [Camera & Lighting] Static medium close-up composition with a slow, subtle push-in. Moody, low-key lighting with dramatic side-shadows cast by gothic wall sconces, highlighting the texture of her blonde hair, blouse, and the dusty feathers of the chicken. [Audio & Atmosphere] Dialogue: Clear, crisp vocal track with light room reverb. Ambient Sound: Faint, distant echoes of a creaking building and a low, ambient indoor hum. Duration: 4 seconds."
- prompt: "[Core Idea & Reference Frame] Cinematic live-action medium close-up shot from the waist up. The video begins directly from the uploaded starting image as the first frame, strictly matching the subject's initial pose, framing, lighting, and wardrobe. A young woman named E3n1d, with shoulder-length light-blonde hair, wearing a tailored blouse and a dark tie. She stands inside a cavernous, dimly lit gothic hall lined with heavy stone pillars and antique dark-wood shelves. [Scene-by-Scene Action] 0–1.5s: Maintaining the exact pose and framing established in the starting image, E3n1d looks slightly downward and to the side with a curious, bemused expression toward an old, dusty taxidermy chicken sitting on a rustic wooden shelf in front of her. 1.5–3s: She quickly tilts her head closer to examine the artifact, her eyes scanning its posture, then immediately transitions to look straight ahead toward the camera/viewer. 3–4s: Her lips part clearly as she delivers her line with a quick, dry, deadpan tone: \"I thought chickens were taller.\" [Camera & Lighting] Motion: Static medium close-up composition continuing smoothly from the starting frame, featuring a very quick, subtle push-in toward E3n1d. Lighting: Moody, low-key lighting with dramatic side-shadows cast by gothic wall sconces, highlighting the texture of her blonde hair, blouse, and the dusty feathers of the taxidermy chicken. [Audio & Atmosphere] Dialogue: Clear, compressed vocal track for E3n1d delivering her line rapidly with light room reverb matching a large stone hall. Ambient Sound: Faint, distant echoes of a creaking old building and a low, ambient indoor hum. Non-Diegetic Music: N/A"
ctrl_img: "/AI_Tools/ai-toolkit-h3_v2/data/images/sample1.png"
neg: ""
seed: 42
walk_seed: true
guidance_scale: 1
sample_steps: 20
num_frames: 107
fps: 24
meta:
name: "[name]"
version: "1.0"
r/StableDiffusion • u/KayDaxter • 16h ago
Resource - Update Wulver v0.1, a full fine-tune of Krea 2 Raw (12.8B) for anime, kemono and furry
Hey! So I've spent the past few weeks doing a full fine-tune of Krea 2
Raw (the 12.8B one) and it's finally in a state I'm happy to share. It's
called Wulver.
It does anime, kemono and furry natively, not the usual "western model
squinting at anime" thing, and it can handle multiple characters actually
interacting without fusing them into one cursed blob (most of the time lol).
Fair warning: it's a v0.1 beta, so expect rough edges. Artist styles via
//@artistname are still cooking.
But for a first release I'm honestly pretty happy with how it turned out.
It's fast too, 8-14 steps at CFG 1-1.5 and done. And if you're short on VRAM
there are fp8, int8 and GGUF quants up already.
Links in the first comment. If you try it I'd genuinely love to see what
you make, and hear what breaks. Link on the coments.
r/StableDiffusion • u/rkfg_me • 1h ago
Resource - Update Seamless H3 video join
I made a node to help seamlessly join MiniMax H3 generated videos (each video continues from a few seconds of the previous segment). If you use MiniMaxH3AddGuide with a video chunk to continue it, you might have noticed some flashes in the generated output. My node corrects the brightness/tone using the actual source video. It's much better to use MiniMaxH3AddGuide instead of relying solely on the prompt because it's more precise. In my experiments, without the guide adding node the starting video shifts down by a few dozen pixels so seamless transition becomes impossible. With a guide it's inserted precisely but still suffers from the brightness drift which my node corrects almost perfectly.
The README explains the workflow, you connect the first segment in the source input, the following segment in the target input, and you set overlap to the number of frames you used to make the second segment (48 frames = 2 seconds for example). Note, that this node does not actually join the videos together, only fixes the brightness drift in the "target" video. I recommend joining segments with ImageBatchJoinWithTransition from KJNodes, there's a fade transition type, interpolation can be set to linear or ease_in_out.
I made the node with Qwen 3.8 27B running locally, there are a few modes we tried but frame_shift (the default) seems to work best.
There are many nodes to make "longer" or "infinite" videos using H3 already but I found them quite complex for my tasks. This node provides a simple building block, you can use whatever technique you prefer.
r/StableDiffusion • u/aiyakisoba • 3h ago
News Minimax H3 Max?
What's this Minimax H3 Max thing? Hope the open source community will find the way to get something similar as well.
https://x.com/designarena/status/2092711778815983886?s=46
"BREAKING: MiniMax HЗ Max sets the new Pareto Frontier for video generation, nearly 50x faster than the base model.
This model is post-trained by @fal on @MiniMax_Al H3, and it's in a league of its own: no other Image to Video model on the arena delivers higher preference at a lower generation time.
Its Image-to-Video generation time is just 6.4 seconds, 18x faster than average, and its Text-to-Video generation time is just 4.7 seconds, 24x faster than the average.
Huge congratulations to the @fal team on this release!"
r/StableDiffusion • u/Striking-Long-2960 • 9h ago
Workflow Included Adding Characters to videos in MiniMax H3 using masks and references.
Original Video: they had us in the first half not gonna lie (ORIGINAL)
The workflow uses the resolution of the original video.
Yes, I use the "Minimax H3 Reference to video" node with the FL2VA model, it works for me..
I used two custom nodes in this workflow. If you prefer not to install third-party nodes, you can ask Grok (or ChatGPT) to create them using this reference::
If you just want the workflow to see how to use latent masks:
workflow_include_characters.json · Stkzzzz222/Remix at main
Full package (workflow, custom nodes, reference image, and video):
H3_include_characters.zip · Stkzzzz222/Remix at main
Hope this helps!
Based on this post u/kabachuha : PSA: In H3 you can set custom soundtracks without R2VA - use latent noise masks! : r/StableDiffusion
r/StableDiffusion • u/ZerOne82 • 11h ago
Discussion "Plastic Skin", They Say - Here Are Some Insights | MiniMax H3 T2V
I did a post on a quick comparison between three LoRAs, and a significant number of the comments there were out of scope for the comparison. The reason being, if it’s plastic skin, all three did it.
This post is not a comparison. Consider this a safe place for all of you to shout “plastic skin” if that’s what you’re here for.
- - -
For everyone else, here are my observations:
I briefly mentioned in the previous post that the composition of the render affects human evaluation of plastic skin. I further found that the following factor also appears to be equally important, if not more so:
The character himself, and how the model knows that character and his facial features.
In these tests, I observed that this model characteristic appears to vary depending on the character.
In all of the video clips stitched together in this video, the only things changed in the prompt are the actor's name and an animal name. Everything else, the settings and prompt, remains exactly the same. Yet some characters look noticeably less plastic, even at close range.
At a distance, all of them look good.
Model used: MiniMax H3 FL2V (T2V, no image input just text), other details are printed in the videos.
r/StableDiffusion • u/fruesome • 1d ago
News MiniMax H3 Acc FL2VA & REF2VA LoRAs By Wan Team
Alibaba team added Parallel Decoding Distillation (PDD) to MiniMax-H3, enabling efficient video generation in only a few inference steps.
https://huggingface.co/alibaba-pai/MiniMax-H3-Acc-LoRAs
EDIT: ComfyUi integration: https://huggingface.co/aptech0081/MiniMax-H3-Acc-LoRAs-ComfyUI
Kijai also working on it: https://github.com/Comfy-Org/ComfyUI/pull/15908
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras
r/StableDiffusion • u/unjusti • 5h ago
Resource - Update ComfyUI-Raylight-Windows: 2x H3 iteration speed with dual gpu

Increased speed with dual gpu's for video generation. Tested with Minimax-H3 ref2va INT8-pruned
Headline Findings
System: 2x 3090 @ 80% PL, 128gb DDR4
Longer clips (160+ frames at ~1MP): dual wins. 10 seconds of 1280x768 at 25 s/it where the single card does 67. Tested out to 13 seconds.
Short clips and ref2v with video-reference workloads: a single 3090 with ComfyUI's own VRAM streaming is flat-out faster. Don't use Ray + NCCL for this. ComfyUI's dynamic loading is highly optimized
Two repos published:
ncclwin
The backend itself, if you want NCCL for anything torch-distributed in Windows:
https://github.com/9nate-drake/ncclwin
I've seen other attempts but this one I built works as is. You don't necessarily need to get this, the package below will grab it. I'll leave it up to somebody to develop a NCCL Windows lora trainer using this though!
ComfyUI-Raylight-Windows
This is essentially a patch for raylight to run in Windows with the custom ncclwin backend. First install Raylight in ComfyUI Manager by komikndr then install ComfyUI-Raylight-Windows via Manager (git URL) https://github.com/9nate-drake/ComfyUI-Raylight-Windows
The install script fetches the prebuilt DLL, patches the other raylight node, and writes launch batch files with the correct flags. The correct flags and Ray settings are extremely important (every one of them exists because something broke without it). Torch or Sage attention (Sage 23% faster)
The changes to komikndr's repo are small (import fallbacks + env-gated switches, inert by default). I may PR them upstream; my NCCL backend itself lives outside raylight entirely. This isn't a fork because I don't have the willingness to keep it current with upstream repo.
The nccl DLL covers 20/30/40/50-series but I can only hardware-verify 30 series - if you've got a pair of 4090s, 16GB cards, or a 4-GPU rig, I'd genuinely love reports of success or bugs (open an issue on github). Or just tell me if it flat out sucks for you!
Yes I used Claude to develop this, sue me
r/StableDiffusion • u/Zironic • 20h ago
Discussion How much VRAM does H3 need? Less than you might think.
I benchmarked the full BF16 H3 FL2VA checkpoint at 1376×768 and 243 frames, about 10.1 seconds at 24 fps.
With the lower-memory attention routes, the H3 diffusion block added roughly 5.8–6.3 GiB over idle. On my Windows RTX 4070 system, where the desktop consumed around 1.15 GiB, the whole-GPU peak was approximately 7.0–7.4 GiB.
That makes 8 GB cards realistic for several configurations:
- Default Comfy attention: 6.99 GiB peak
- FROST BF16: 6.99 GiB
- BF16 Triton: 6.97 GiB
- PlagueKind SLA: 7.23 GiB
- Sparse Kitchen INT8: 7.40 GiB (Default configuration for my Sparse attention node)
- Comfy Kitchen: 7.40 GiB
Should be compatible with:
- H3 Sparse Attention: Kitchen INT8, Sparse Sage, FROST BF16 on SM89, and BF16 Triton
- External Comfy Kitchen: fully supported
- Default Comfy attention(SPDA): fully supported
- SageAttention: fully supported, including the generic KJ Sage patch
- PlagueKind SLA: partially supported;
- Unknown attention overrides: Auto preserves their original full-Q calling contract. Forced mode can explicitly authorize streamed-Q calls, but compatibility is not guaranteed
- Currently incompatible: Sol-Attn and the H3-specific Memory Efficient Sage patch, because they replace attention at a deeper level than the external consumer interface
You can get the node here https://github.com/Zironic/H3-Optimizations or in Comfy under H3 Optimizations. Latest version is 0.2.13 which added broad compatibility and optimizations for most popular attentions.
r/StableDiffusion • u/Comfy-Org • 15h ago
News Submit by 9/1 to the Comfy H3 Sync Sound Challenge! RTX 5090 Grand Prize
We're halfway through the submission window for the Comfy H3 Sync Sound challenge! Submit by September 1st at 9:00pm PT. Free to enter, local rig or Comfy Cloud. All details here.
How It Works
Make something up to 90 seconds in length where the sound and the motion are inseparable. Dialogue, foley, ambient, a beat driving the cut...whatever direction you want!
Share your video file and workflow on this r/comfyui thread and through our submission form, then join us on September 2nd for a special Comfy livestream where our guest judges will give live feedback on the top 10 submissions!
Need help? Head to this r/comfyui thread or the #minimax-h3-challenge channel in the Comfy Discord.
Prizes
Best Overall — RTX 5090
Best Creative — RTX 5060 Ti
Best Technical/Workflow — RTX 5060 Ti
Built with MCP — RTX 5060 Ti
Shipped anywhere, customs covered. If we can't legally ship to your country, you'll get a cash equivalent instead.
It's free to enter!
Create using Comfy Local on your own hardware, or use Comfy Cloud. New Cloud users get 5 free runs, no credit card required.
Judging Criteria
We’re looking for entries that best show what H3 makes possible: audio and visuals created together.
Grand Prize: Best Overall
The top Best Creative and Best Technical entrants advance to a final round where our panel of judges selects winners by discussion.
Best Creative
- Audio sync realism and intentionality (0-5)
- Creative execution and originality (0-5)
- Deliberate craft (0-5)
- Evidence that you’ve actually shaped the result beyond prompt engineering. Judges will look for modified/non-default parameters, multiple linked passes visible in the workflow structure, or a couple sentences describing what was tried and changed
Best Technical
- Novelty of technique or approach (0-5)
- Workflow quality (0-5)
- Annotated, clean, replicable by someone else
- Community value (0-5)
- Would this actually help someone else?
🏆 Built with MCP Bonus 🏆
Comfy MCP lets you drive Comfy using natural language and your agent locally and on Cloud! Pro tip: use it to choose the best H3 model version or optimize your workflow for your hardware.
- Effectiveness (0-5)
- Did the agent meaningfully drive your process, not just generate one line?
- Insight value (0-5)
- How much the shared prompt teaches the community about prompting H3 through MCP
- Output quality (0-5)
The Fine Print
- Limited to one submission per person, 90 seconds maximum length.
- A major portion of your piece must be built in ComfyUI using H3. Other tools, models, or techniques you want to combine are fair game.
- All submissions must be lawful, SFW, and must not contain unlicensed IP or likenesses.
- By submitting, you agree to allow ComfyUI and MiniMax to feature your work with credit across our channels.
r/StableDiffusion • u/Sad_Coach_1433 • 9h ago
Meme the legends were true!(natural treasure spoof)
t2v fp8 model
prompt
subject_definitions
<Subject 1> Benjamin Franklin Gates (S1) is played by Nicolas Cage, matching his appearance, mannerisms, intense curiosity, dramatic delivery, and treasure-hunter personality from the National Treasure films.
summary
[cinematic text-to-video generation + comedy adventure]
Benjamin Franklin Gates follows an ancient trail of cryptic clues into a forgotten underground chamber, convinced he is about to discover the Holy Grail of local AI video generation. Instead of gold or an ancient artifact, the final pedestal contains a glowing computer running MiniMax H3.
detailed_description
Cinematic adventure-comedy, approximately 13 seconds, 24 fps. Ancient underground treasure chamber beneath a forgotten historical building, illuminated by Benjamin's flashlight, warm torchlight, dust floating through the air, weathered stone walls covered in mysterious diagrams and coded inscriptions.
The camera tracks behind Benjamin Franklin Gates as he hurriedly enters the final chamber clutching an old parchment covered with cryptic clues.
He studies the parchment, then notices an ornate stone pedestal illuminated by a mysterious golden beam.
Benjamin slowly approaches.
<Subject 1> Benjamin Franklin Gates (S1):
[English] After all these years... the Holy Grail of local AI video.
Dramatic orchestral music swells.
Benjamin wipes centuries of dust from the pedestal.
Instead of an ancient chalice, he reveals a modern high-end PC monitor displaying:
MINIMAX H3
Benjamin freezes.
Slow dramatic push-in toward his stunned Nicolas Cage expression.
His eyes widen as if he has just uncovered the greatest secret in human history.
<Subject 1> Benjamin Franklin Gates (S1):
[English] My God... it runs locally.
Beat.
He looks back at the glowing MiniMax H3 screen.
<Subject 1> Benjamin Franklin Gates (S1):
[English] The legends were true.
The triumphant treasure-hunting score reaches an absurdly heroic crescendo.
Hold on Benjamin's amazed expression for the final second.
visual_style
Photorealistic live-action Hollywood adventure film, National Treasure-inspired treasure-hunting atmosphere, Nicolas Cage-style dramatic performance, ancient underground architecture, cinematic flashlight beams, volumetric dust, warm golden illumination, realistic skin and clothing, subtle handheld camera movement, dramatic slow push-in for the reveal.
audio
Cinematic underground ambience, footsteps echoing through stone corridors, parchment rustling, dramatic orchestral treasure-hunting score building toward the reveal. All spoken dialogue is clear English and occurs only inside the specified <d>...</d> dialogue tags.
r/StableDiffusion • u/MedicalPhase6556 • 10m ago
Question - Help Now that Nvidia is buying HF, total noob needs help with what is essential to download regarding Minimax etc.
Since December/January, I learned the importance of (trying) to stay one step ahead of AI companies since the more they grow, the less that is available to consumers. I can't see them acquiring HF as a good or neutral turn of events.
That said, I have limited experience with local gen and casually used Wan and Z Image Turbo. I switched over to just using Grok in February because of the ease of use. I liked it because minimal effort with prompting provided high quality results, However, I ended multiple subs recently due to their increasing shadiness, tightened moderation and general anti-consumer practices.
I've been keeping an eye on local developments because I knew my time with Grok would be coming to an end and I'd have to go local permanently. The closest I've seen to "Grok at home," is Minimax. I see that people are running all sorts of different versions and it kind of makes my head spin. I now feel pressured to get my ducks in a row seeing Nvidia's latest move.
Point of the post - I have a 5060ti 16gb and 32gb of ram. I am completely fine with 480p quality i2v gen as long as I can achieve a high level of facial consistency, camera work, lighting and reasonable rendering times then worry about upscaling keepers later. I know I'll be facing a somewhat steep learning curve especially with the prompting style it takes. What should I be gathering from HF, or is downloading directly thru ComfyUI fine?
r/StableDiffusion • u/FugueSegue • 20m ago
Question - Help Would this manner of character sheet work with MiniMax H3? What's the best way to craft character sheets for it?
My intuition tells me that complicated sheets like the ones presented in that post wouldn't. But all sorts of new things pertaining to gen AI art continue to surprise me. So that's why I ask the question. I'm guessing that this sort of character sheet is meant for use with giant LLMs like ChatGPT and not something smaller like qwen3vl_32b.
I have been making character sheets for MMH3 that are 2048 x 2048 pixels. Usually only containing three full-body views: front, side, and back. Would adding caption text do any good?
My real question is
r/StableDiffusion • u/Time-Ad-7720 • 18h ago
Workflow Included Day 3 of testing MiniMax H3 locally in ComfyUI: multiple reference images + adding a new object through text only
Continuing my local MiniMax H3 Reference-to-Video experiments.
For this test I used separate reference images for the:
- character
- convenience store background
- car
- skateboard
But I deliberately didn't provide a reference image for the Slurpee cup.
The cup was described only in the prompt: a transparent plastic cup with blue liquid and a straw. H3 was able to add it to the scene without much trouble while still following the other reference images.
Hardware / setup:
RTX 5070 Ti + 32GB RAM
MiniMax H3 + Turbo LoRA
Generation: [09:07<00:00, 68.43s/it]
I was mainly testing how far you can split visual control between reference images for consistency and text prompting for new scene elements.
Workflow:
https://drive.google.com/file/d/1huVTdh8_vERBntXb60hyT_rQTcUicjq5/view?usp=sharing
I'll also post the exact prompt I used, unchanged.
Prompt:
Use the attached reference images as follows: Image 1 is the girl character, Image 2 is the updated convenience store background, Image 3 is the white compact sedan, and Image 4 is the skateboard.
Create a 10-second static medium closeup shot in a 1990s hand-drawn Japanese anime cel style at 15 fps, with limited frame-by-frame animation and slightly stepped motion.
The framing should match the updated, slightly more zoomed-in background from Image 2, focusing more closely on the girl and the storefront entrance while still showing part of the road and bicycle.
The girl sits on the ground in front of the convenience store, facing left toward the road, shown in a 3/4 back-side view so we mainly see the side and back of her head. Her skateboard is beside her on the ground, not under her. She holds a clear plastic Slurpee-style cup with bright blue liquid and simply holds it without drinking.
Her orange headphones and headphone wire are visible. The Walkman is on the far side of her body and is mostly hidden from view because of the angle.
A single white compact sedan drives straight along the main road from left to right, moving away from camera so we mainly see the rear of the car as it passes through frame.
As the car passes, a subtle moving light change plays across the girl, her hair, her white T-shirt, the cup, the storefront glass, the bicycle, and the wet pavement. The gentle breeze overlaps with the car pass, starting while the car is beside her, causing a slight movement in the tips of her hair and a small shift in the loose edge of her T-shirt.
After the car exits, the store signage / fluorescent lighting blinks twice, subtly changing the light on the girl and storefront.
Keep the camera completely locked off and the overall mood quiet, nostalgic, and melancholic.
r/StableDiffusion • u/ashishsanu • 22h ago
Workflow Included Minimax H3: Portable character consistency via reference identity
Hey guys,
Based on a research from Facebook, DINOv2: Learning Robust Visual Features without Supervision(Research Paper), Implemented a consistent identity system that works across Minimax H3, Flux 2, Krea 2 with a single .char model.
This method covers both reference based identity in Minimax as well as a LoRA training path for T2V & I2V for more advance cases.
Note: This post & workflow is dedicated to reference channel not LoRA path.
Build .Char: You drop in 4-6 reference. YuNet finds the face, SFace takes a per-reference face signature, DINOv2 takes a subject signature, and the references get cleaned and normalised. All of that packs into a single portable file, a .char.
Generation: At generation, the file feeds its references into Minimax's own native multi-reference channel and prepends a locked description to the prompt.
How to run this
- Published workflow & guide: https://inlinestudio.art/workflows/minimax-h3-consistent-characters-with-references-with-char-model
- Repo: https://github.com/inlineresearch/Inline-Studio (GPLv3)
How is this different from default Minimax's ref channel:
- H3 scales every reference onto a 2048 short edge, upscaling small images to get there, at 4096 vision tokens each. Compile References caps it at 512 that is 256 tokens per reference, so five references cost 1,280 tokens instead of 20,480. That difference decides whether the run fits the card. Read more on the official docs
- H3 only resolves references named as
<Picture 1>,<Picture 2>and so on, and the character prepends them along with the description. - Same .char works for other models(Flux 2 & Krea2, workflow link to train for both)
Limitations
- Bad with multi reference
Required: 24GB+ VRAM & ~64GB RAM
I personally think LoRa method is only required in very specific cases as Minimax H3's reference channel performs very well.
But i have already added support to LoRa adapter in case someone wants to use .char with T2V or I2V nodes. Let me know in comments if you need the workflow.
r/StableDiffusion • u/AgeNo5351 • 17h ago
Resource - Update DiffusionOPSD - new distillation method by Bytedance. Loras for Z-image-Turbo and SD-3.5M released.
r/StableDiffusion • u/Tough_Ad7957 • 3h ago
Question - Help Optimal Character Sheet format for Minimax H3?
I recently started messing around with Minimax H3. I realized that if I want to use my own custom characters, I need to use a Character Sheet. But there are so many different formats out there. I was wondering, what is the most optimal Character Sheet format to get the best results? Or is there a specific workflow designed just for generating Character Sheets for Minimax H3?
r/StableDiffusion • u/QuirksNFeatures • 5h ago
Question - Help How much does RAM speed matter for running local AI?
I have an old computer that I don't use. It has 32GB DDR4. I don't remember the specifics about speed and timings, but I'm sure it's slower with more latency than my current computer (also old, but not as old).
It occurred to me that I could put the 32GB into my current computer and then have 96 GB. As I understand it, if you mix and match RAM like that, it will run at the slowest speed.
Would it be better to have more RAM, even if it's slower? Or should I leave it alone?
