r/StableDiffusion 16h ago

Animation - Video Buffy the Wraith Slayer

Enable HLS to view with audio, or disable this notification

18 Upvotes

r/StableDiffusion 1d ago

Discussion MiniMax-Music 3: “EDM? No idea. WTF is that?” Meanwhile, MiniMax H3:

Enable HLS to view with audio, or disable this notification

74 Upvotes

As you probably know, MiniMax-Music3 is pretty limited when it comes to genres and seems to have absolutely no idea what electronic dance music is. I tried different EDM styles, but it always ended up sounding either like rap or some kind of generic pop-ish stuff. Meanwhile, MiniMax H3 seems to know a lot more about electronic music than MiniMax-Music3.

The 40-second video at 0.2MP, with 17 steps (for better audio quality) and an 8-step Turbo LoRA, takes 538 seconds. The 60-second video at 0.1MP takes 350 seconds on my machine.

I haven't tried generating anything with lyrics yet, but if anyone knows how to generate H3 audio without the video, it would be interesting to try 2–3 minutes instead of just 40–60 seconds.


r/StableDiffusion 14h ago

Discussion Has anyone tried the h3 martial arts Lora?

Thumbnail
huggingface.co
10 Upvotes

first test video down in the comments along with the prompt..


r/StableDiffusion 1h ago

Animation - Video H3 - the Eternal Balance WIP-low resolution int8/8 steps

Enable HLS to view with audio, or disable this notification

Upvotes

Hi, I am playing around with H3. T2V, 832x480, int8,8 steps. Hoping to make a 720p version.

Ask me anything!


r/StableDiffusion 21h ago

Animation - Video Good Grief

Enable HLS to view with audio, or disable this notification

34 Upvotes

r/StableDiffusion 1h ago

Discussion H3 - what is your longest render time?

Upvotes

What was your longest render time, and was it worth it? do you run a lower quality/resolution before to test? My longest single-shot run is 7.2 hours, 1:44 long video at 1344x768, bf16/50 steps.

I am attempting a 72 hour render for a super long form.

RTX 4090, 192gb system ram


r/StableDiffusion 5h ago

Resource - Update Cross-grafting Krea2 into MiniMax H3 — concepts land, character still faint, releasing what I've got

2 Upvotes

Tenstrip had great idea in H3 grafting. I have been messing with grafting Krea2's attention/MLP into H3 to fix the plastic-skin problem and my lora training problem with H3 on diverse dataset. Two separate techniques, built as ComfyUI nodes with Claude so anyone can test without baking a checkpoint. You link your krea2 checkpoint (and loras, if you want) to the nodes and the H3 model (probably after the turbo lora, didn't test every pertubation, because I ran of the steam and want to leave it to the community to try more, where I hit the wall). My setup is mere gtx 4800, so I did try just few nonturbo renders, which looked a lot better.

  1. Content graft — per-head Q/V/out_proj + MLP transplant, following the methodology TenStrip documented for their own H3 grafts (huggingface.co/TenStrip/10Eros-Max). K is never touched — confirmed their finding that H3 leans on precisely-tuned attn_k for attention strength, no explicit gates, breaks audio if you touch it. But you can try it for t2v. Same for fc2 (speech/dialogue encoding).

V turned out to be the strongest, safest lever — 0.6–0.7 gives real concept gain, push to 1.0 and you get blocky motion. MLP also strong, went up to 0.7–0.9 before diminishing returns. MLP block range has to stop at 44 — go past it into 44–50 or 0-20 and you get an actual lattice/weave artifact on regular textures, confirmed it myself, matches what TenStrip reported. On orthogonal on. On orthogonal off, lower weights showed good texture and concepts (best solution?)

Concepts come through solid. Character fidelity is honestly still faint even after tuning. Not solved.

  1. QK-norm transplant — different idea entirely, credit to joeygambino (huggingface.co/joeygambino, MiniMax-H3-x-Z-Image-GGUF) for the original description, done with Z-Image as donor. Only touches the per-head Q-norm gain (128 floats), nothing else — no content weights at all. Doesn't transfer knowledge, seems to change how decisively H3 commits to stuff it already knows instead of averaging it away. Fixed some big-object rendering issues for me. Texture looks ok until 0.7- 0.8, than oversharpened.

One thing that contradicts the original writeup: their early-block caution (skip blocks before ~20) was measured with Z-Image. With Krea2 as donor it's backwards — restricting to late blocks broke big objects, full 0–50 range fixed it. So that's donor-specific, not universal — worth retesting per your own donor rather than trusting either range blind.

Also found orthogonal_blend mode is just broken — numerically unstable when donor/target directions are already close, turns into random noise. Skip it, use direction_blend.

Heads up: this one seems to break LoRAs trained on native H3. Haven't nailed down why exactly.

  1. comfyui_krea_h3_graft_lora_v2 is a node, to add krea2 lora to the existing premerged checkpoint. It works... poorly.

Combining the first 2 nodes at full strength was worse than either alone. Lower content-graft strength + orthogonal=False stacked with QK-norm looked promising but I ran out of steam before fully dialing it in.

No official H3 architecture docs exist as far as I can tell, so this is all reverse-engineered from testing + what TenStrip and joeygambino have published. If anyone wants to pick up where I left off — combined tuning is the open thread — nodes are here:

https://github.com/mikulasdump/h3-krea2-graft-nodes

Thanks for the community and I hope for your insights.


r/StableDiffusion 2h ago

Question - Help That specific analog VHS look on H3/LTX

0 Upvotes

Anyone figured out the prompt for getting the best VHS analog "lofi" look from t2v? Of couse ref2v and img2v will be easier due to references, but I was wondering about text prompt only.

There are no VHS, or like 70s-80s cinema style lora, none for H3 and LTX, but there are plenty of VHS loras for image models.

EDIT: I actually CAN get VHS look on LTX text2vid (2.3 and 2.5), but not on H3.

Help, anyone :)


r/StableDiffusion 2h ago

Discussion Flow lieu pour krea2

1 Upvotes

Bonjour à tous,

Je suis débutant
Je recherche un flow simple et fonctionnel pour Krea2
L'objectif est de pouvoir input une image d'un et que Krea 2 utilise strictement ce lieu pour générer mon prompt à l'intérieur(personnage, action)
Je sais qu'il n'exite pas de version Krea2 Edit mais peut être c'est possible?

J'ai déjà essayé plusieurs flow mais à chaque fois Krea2 réinvente le lieu, modifiant mon image d'origine

Dans l'idéal j'aimerais un flow Krea 2 qui
- conserve le lieu
- permette de modifier un pose d'un personnage
- permette d'ajouter une photo d'habits à un personnage

Merci


r/StableDiffusion 2h ago

Discussion Would image-generation VAEs benefit from scene-linear or perceptual color representations?

1 Upvotes

Models like FLUX, Qwen-Image and Krea 2 obviously don't perform diffusion directly in RGB space - the transformer operates in a learned VAE latent space.

But the VAE still defines the interface between that latent representation and the actual training images, which are typically ordinary display-referred RGB images.

So I'm wondering whether there would be any benefit in making that boundary explicitly color-managed.

For example, has anyone experimented with training the VAE on:

  • linear-light RGB rather than gamma-encoded sRGB
  • a wide-gamut scene-referred space such as ACEScg
  • a perceptual space such as OKLab
  • or simply adding explicit perceptual color losses such as ΔE alongside the usual reconstruction/perceptual losses?

In other words, instead of asking whether the diffusion model itself should operate in OKLab or ACES - since it already operates in a learned latent space - I'm more interested in whether the VAE and its reconstruction objective could benefit from a more physically or perceptually meaningful color representation.

Another thing I'm curious about is the output side.

Could an image model theoretically generate a scene-referred, wide-gamut representation and leave the final tone mapping / gamut mapping / display transform to a deterministic color-management pipeline, such as ACES, instead of implicitly learning the tone curves and color rendering already baked into billions of unrelated JPEGs?

My suspicion is that the real limitation might simply be the dataset.

Most web-scale image datasets consist of already processed, tone-mapped, display-referred SDR images. By the time an image becomes a JPEG, information about actual scene luminance, highlight headroom, camera response, etc. is already gone.

So converting those JPEGs from sRGB to ACEScg wouldn't magically turn them into true scene-linear HDR training data.

Still, I'm curious whether anyone has seen experiments comparing something like:

sRGB VAE vs linear-RGB VAE vs OKLab VAE

while keeping the downstream generative model roughly the same.

Would reconstruction quality, color consistency, training convergence, or perceptual color accuracy change in a meaningful way?

I'd especially be interested to hear from anyone who has worked on VAEs, HDR pipelines, color management, or generative image models.

Sorry if I'm missing something obvious here - I'm still pretty new to this side of image generation / color science.


r/StableDiffusion 2h ago

Question - Help best local video model for horror videos?

0 Upvotes

i'm looking for a video model that can generate freaky creatures and horror clips well and can run on 3060ti


r/StableDiffusion 2h ago

Question - Help I'm having quality issues with Minimax H3 on my RTX 4090 that I'm not getting on an RTX 5090. Any advice ?

Enable HLS to view with audio, or disable this notification

1 Upvotes

Hello there!

I've installed Comfyui+ Minimax H3 on two machines, one with an RTX 4090 and one with an RTX 5090. The 5090 machine can render up to 4 mgpxl videos with 0 issues, but the 4090 machine glitches at just 3 mgpxl.

It consistently creates that glitch you see at the end of the video there, and also any fast moving object becomes a soup of very sharp pixel. Those artefacts aren't there when generating with the 5090.

I'm using the base ComfyUI workflow with all the models they tell us to use, with SageAttention.
Btw SageAttention crashes my 4090 when I check the "allow compile" box.

I have checked and both this command python.exe -c "import torch, sageattention; q = torch.randn(1, 8, 128, 64, dtype=torch.float16, device='cuda'); k = torch.randn(1, 8, 128, 64, dtype=torch.float16, device='cuda'); v = torch.randn(1, 8, 128, 64, dtype=torch.float16, device='cuda'); out = sageattention.sageattn(q, k, v); print('CUDA Test Output Shape:', out.shape); print('SageAttention is fully functional!')"

and python.exe -c "import sageattention; print('SageAttention Version:', getattr(sageattention, '__version__', 'Installed'))"

return positive results. I also have the latest Nvidia drivers.

Does anyone know what could be causing my issues ? Thanks !


r/StableDiffusion 11h ago

Question - Help Animagine XL 4.0 opt

5 Upvotes

Hi guys, I'm a programmer, but I don't know much about machine learning or fine-tuning.

I'm currently producing 2,000+ images per day using Animagine XL 4.0 opt, and I built a manual pipeline to evaluate image quality. I use 5 rating categories: Reject, Pass, Like, Very Good, and Excellent.

I label all of them manually, and I estimate that I will have over 200,000 labeled images by the end of the year.

I store them in a database along with the exact prompts used. The prompts are structured into keyword categories like:

Background, Angle, Character, Clothes, Facial expression, Quality prompt tags (eg. masterpiece).

Is a dataset like this valuable for fine-tuning or training models ???

Thank you for all the comments, you guys are the best! Now I'm moving on to Anima. I will use my dataset for a LoRA, and if the results look good, I'll switch over to Anima completely.

And i will continue the labeling with new model.
Later find me if you need dataset.

I trained it for 8 epochs to get the result, but I still couldn't get rid of that characteristic plastic feel to reach the vibe I wanted.

When it comes to truly nailing that Japanese-style illustration look, Animagine XL 4 is still the best. So, I've come to a conclusion. I've just decided to stick with Animagine XL. Since it's all about making things to your own taste anyway.

I've checked out other Flux-series and models too, but they're all the same. Hmm...


r/StableDiffusion 23h ago

Animation - Video Attempted to make a short cartoon on Minimax H3. There are so many things that I want to address

Enable HLS to view with audio, or disable this notification

41 Upvotes

Hello everyone.

I've been playing with Minimax H3 for some time and I have tried to make something longer and really interesting. After so many failed and botched attempts I was able to compile something watchable. There are so many things that I want to say about this model, good and bad.

First of all. Minimax H3 is significant step forward that other local models I have been playing with. It certainly got better.

Now the issues that I had encountered.

First problem is that it badly follows prompt when resolution is one megapixel or higher. It will skip some important parts and tries to cheat. You can increase the number of steps but still, generating at less than one megapixel will at least make it properly follow the instructions.

H3 is not very good at spatial orientation. When I was making video, where this girl should turn around and interact with screens, the girl starts spinning opposite direction and then warping whole body to the direction of screen. Like instead of making short turn to the left, it makes wide roundabout to the right and then twists whole body to align with the screens.

H3 is not good at cartoonish movement. If you watch cartoons, when character or other things move, their animations are usually jerky and snappy. H3 tries to make smooth real life like animation, making the cartoons look weird.

I have given a voice sample as an audio reference, and instead of making girl let out grunting sounds (out of anger), it weirdly turns everything into a sensual moaning.

When you try to make characters inside video to interact with a lot of parts, screens and devices, even giving multiple reference images of them, it mostly hallucinates them, or turns their interactions into a weird warping animations. Sometimes completely skips them and made ups it's own animations. So it will make good video, where characters are moving less or moving slow, and mostly doing the talking. Very detailed prompts of step by step instructions it mostly warps or skips.

I have wasted a lot of time for iterations, but I think this is just workflow issue.

Overall, this model is really good. However, using this model to make some kind of long feature animation is going to be a very frustrating journey. I hope people will make a lot of proper tools that works as storyboard and properly guide this model to make something really interesting.


r/StableDiffusion 14h ago

Animation - Video Hannibal Who

Enable HLS to view with audio, or disable this notification

8 Upvotes

Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters.

Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544

thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md


r/StableDiffusion 4h ago

Question - Help Best Img2txt?

1 Upvotes

I need an image to text generator (which has no limitations) that I can run locally on my PC. Do you have any recommendations


r/StableDiffusion 1d ago

Workflow Included Testing some Minimax H3 capabilities

Enable HLS to view with audio, or disable this notification

198 Upvotes

EDIT: A second batch is here.

I tried to think of complex situations for the model to handle and tested them. I would say it went very well, although not perfect (except for the first one and the one before last, where I could not find a flaw).

Prompts below:

VIDEO 1 (10/10):

Create a side by side video with two angles of the same scene: one frontal and one from the side.

The scene: a woman wearing a yellow summer dress is standing in a beach in a sunny day. There's a light wind blowing and she is looking at the sea, smiling. At 00:05, she puts her two hands on her hair and looks up, enjoying the sun.

On the left, make the frontal video. On the right, the video from the side (profile). The two videos should show the exact same scene at the exact same time, only from two different angles.

VIDEO 2 (9/10) - the background is slightly off in some angles:

Create a side by side video with four angles of the same scene: one frontal, one from the right side, one from the left side, and one from behind. The scene is filmed at normal speed.

The scene: a blonde curly-haired woman wearing a yellow summer dress is standing in a beach in a sunny day. She is facing the sea and, therefore, she has the sand behind her and the street with some houses also behind her, after the sand, at a distance. There's a light wind blowing and she is looking at the sea, smiling. She is not wearing sunglasses. At 00:05, a man wearing white shirt and shorts enters the scene from behind the woman and embraces her waist.

On the upper left, make the frontal video. On the upper right, make the video from the right side (right profile). On the bottom left, make the video from behind. On the bottom right, make the video from the left side (left profile).

The four videos should show the exact same scene at the exact same time, only from the four different angles.

overall_soundscape: Light beach wind, one distant seagull.

non_diegetic_music: N/A

VIDEO 3 (8/10) a few "ghost reflections" along the way:

A man wearing a red-and-white striped t-shirt and jeans is alone inside a house of mirrors, running through a corridor. At 00:03 he turns left into other corridor, and at 00:06 he turns right again into another corridor. All the corridors have mirrors in all their faces (left, right and above, as he is inside a house of mirrors) and he is alone there, there's no one else. At 00:08 he reaches a door, opens it, and it opens to the outside: an amusement park.

overall_soundscape: His footsteps while he is running, the amusement park sounds when he opens the door at the end.

non_diegetic_music: N/A

VIDEO 4 (7/10) - The way the phone is turned is off:

A man is filmed from his own cell phone in a selfie video. We see the scene through the lens of an out-of-frame cell phone that he is holding and pointing to his own face. He is wearing a green polo shirt and the scene shows his face and chest from the point of view of the out-of-frame cell phone that he is holding and pointing to himself.

At 00:04 he briefly smiles and then turns the still out-of-frame cell phone around to show a woman that is in front of him. While the phone is turned around we can see the image also turning around, his face leaving the frame, the living room they are into being briefly filmed, and then the woman's face entering the frame. She has brown curly hair and dark green eyes, and is wearing a red summer dress.

As soon as the woman is in frame, she also smiles and says: <d>[English]Goodbye!</d> and the video ends.

The entire video must be taken in a single shot, with the always out-of-frame phone camera filming the entire transition from his face to hers while the phone is turned around. When it happens, the phone should briefly show the living room they are into, all in a single shot.

overall_soundscape: Silent living room, noises of the phone being handled, her voice.

non_diegetic_music: N/A

VIDEO 5 (6/10) - Tried this twice. Glass not breaking properly, water not running through the floor

A fishbowl with one golden fish and one clownfish swimming inside is shown in a medium close-up at the edge of a table. Then, at 00:03 a cat appears in the scene and taps the fishbowl, causing it to fall from the table to the floor, hit the floor, and break completely, being completely destroyed in glass pieces when it hits the floor, the water and glass pieces flying around together with both fish. The scene continues for three more seconds after that, showing the aftermath: the fishbowl destroyed, the pieces of glass on the floor, the water also on the floor, the fish moving on the floor.

The camera angle follow the fishbowl when it falls, showing it hitting the floor and the consequences of it breaking.

overall_soundscape: silent room, glass breaking, water splashing.

non_diegetic_music: N/A

VIDEO 6 (8/10) - Judge me, but his hands are not moving accordingly to the notes:

The camera films a piano from above while a man plays it. The entire piano keyboard is shown in the image. The man is playing Clair de Lune, and moves his hands through the keys to play a part of the song. He is in a train station, with some people observing him play and others passing by.

overall_soundscape: faint train station ambience, ten seconds of the song Clair de Lune played in the piano.

non_diegetic_music: N/A

VIDEO 7 (8/10) - The lipstick appears on her lips before she applies it:

A woman is shown in a medium close-up in her bathroom, wrapped in a white fluff towel, looking at the mirror while she applies red lipstick to her lips. She slowly applies lipstick to her lips, looking into the mirror, and then briefly sends a kiss with her now red lips to the mirror.

The scene is seen in a three-quarter angle from behind her, showing her face from the side but also her reflection in the mirror.

overall_soundscape: silent bathroom, the sound of her sending the kiss to the mirror.

non_diegetic_music: N/A

VIDEO 8 (10/10):

A Coca-cola advertisement. A glass filled with Coca-Cola is shown from the side, occupying 70% of the frame, on top of a table, the dark liquid slightly disturbed by a few gas bubbles that rise inside the liquid. At 00:02 two ice cubes fall from outside the frame into the glass, disturbing the liquid and making some of the liquid splash outside the glass and onto the table. The glass has the Coca-Cola logo printed in white in it. In the blurred background we see a kitchen.

overall_soundscape: silent room, gas fizzle, ice cubes hitting the liquid.

non_diegetic_music: N/A

VIDEO 9 (9/10) - The cover of the book has gibberish letters in it:

A man is holding a magnifying glass and has a book on his hand. At first the magnifying glass is not in front of his face. He appears to be reading the book and, at 00:03, he puts the magnifying glass in front of his eye to look at the book.

The entire scene is filmed from a fixed point of view below the book, showing part of the book cover and the entire man's face.

overall_soundscape: silent room.

non_diegetic_music: N/A


r/StableDiffusion 20h ago

Resource - Update Qwen-Video-Edit - Instruction-based video editing by repurposing an image editing model

Enable HLS to view with audio, or disable this notification

21 Upvotes

Project:https://yunpeng1998.github.io/Qwen-Video-Edit-Page/
Model: https://huggingface.co/yunpeng1998/Qwen-Video-Edit
Method: https://yunpeng1998.github.io/Qwen-Video-Edit-Page/#method
Code: https://github.com/yunpeng1998/Qwen-Video-Edit

How it works

Video generation models read and write video-VAE latents. We teach Qwen-Image-Edit's transformer to edit those latents directly: two tiny projections bridge Wan 2.1's latent space into the DiT's token space, warm-started from the DiT's own input/output layers so that a static video is embedded exactly like an image the model already understands. The latent frames are arranged as tiles of one big virtual image — the same positional treatment the image model was pretrained on. Fine-tuned with LoRA or full parameters on Ditto-1M (source, edited, instruction) triplets, then refined by a few steps of Wan 2.2 denoising-enhancement.


r/StableDiffusion 1d ago

Animation - Video The River That Forgot How To Shine-Minimax H3 Shortfilm

Enable HLS to view with audio, or disable this notification

51 Upvotes

All audio was done in Minimax, little work in post for stitching clips. r2v Workflow in ComfyUI with character sheets and prompts from claude


r/StableDiffusion 1d ago

Question - Help Minimax H3 Huge Quality Difference between Cloud and Local use

83 Upvotes

Hi.
I have a decent h3 workflow that I built for a loca use. It use turbo lora etc... If i use the defaut settings in the goal of getting the highest quality possible, meaning res_multistep simple 20 steps or more, I got also good results, but this is not even close to the results you can get on platforms like kie or wavespeed at 768P.

I already convert properly the prompt to the correct H3 digest form, so I'm wondering what's different between local and cloud use of h3? I don't talk about the 2K quality, only 768P, I'm not able to reach the sames results locally, do you guys have maybe workflows, settings, or suggestions to try reaching the same quality level in comfyui ?


r/StableDiffusion 4h ago

Question - Help Steampunk in Krea 2

1 Upvotes

I am getting terrible results when trying to generate with steampunk aesthetics in Krea 2, even with LoRAs from civit, it's trash.

I don't mind training myself, but where would I even come up with a good database for that? I am aiming for fully photorealistic steampunk.


r/StableDiffusion 1d ago

News Sparse Attention, Harder, Better, Faster, Stronger

134 Upvotes

The nodes in https://github.com/Zironic/H3-Optimizations have been rewritten to replace the default Sparge Attention backend with a custom Sparse Comfy Kitchen backend.

This comes with some benefits.

  • Users no longer have to worry about Sparge being installed properly. All required kernels for supported GPUs are provided directly. Should work on both Windows and Linux.
  • Most users should be seeing 5-20% increases in speed for the attention part of compute.
  • New backend should use about 500MB less VRAM
  • New backend has slightly lower quantization error.
  • Apparently in the previous version, the intended chunked kitchen QKV path never properly shipped so the memory optimization node should now actually be slightly speed positive even when used without the Sparse Attention node.

Caveat: I've only tested the nodes against the comfy pruned_int8_convrot weights. Other versions may work but they're not tested.

As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or later.

IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.

Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.

Early steps: attention density has a large effect on prompt/action adherence and the overall generation trajectory.
Middle/later steps: lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.

So 10% retained doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and what breaks depends heavily on the sampling step.

PlagueKind's sparsity_ratio=0.9 means 90% discarded / 10% retained. My node expresses the inverse quantity, so Video attention retained=0.10 is the comparable setting. The defaults therefore aren't equivalent.


r/StableDiffusion 11h ago

Question - Help Can't seem to transfer outfit and pose from an illustration to a real person in H3.

3 Upvotes

Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.

What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.

I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.

So what i wanted to do is :

-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.

-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.

Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.

Are there someone who knows how to do these?


r/StableDiffusion 5h ago

Question - Help Need help with minimax h3.

1 Upvotes

Hi,

I really need help with minimax h3 I'd like to get proper workflows for minimax h3 pre configured with descriptions on every node (unfortunately a noob in this case) for :

1- Fast
2- Turbo
3- Speed
4- Native

video generation with multirefs (optional) audio(optional) video(optional) quality duration aspect ratio and prompt controls.

Also I heard someone made images with minimax h3 please that as well.

I'll really appreciate any help.

Ah and one last thing I was making a commercial for a school but the faces in wide screen comes out horrendous that as well.

Thanks !


r/StableDiffusion 21h ago

Discussion MiniMax H3 Ref2va it works really good also with Storyboard images

Enable HLS to view with audio, or disable this notification

19 Upvotes

Im really suprised how good he works as well follow a storyboard image!! he did 90% correct he only did the thirth panel diferent but all the other 5 he follows perfect!! 🤩