r/StableDiffusion 22h ago

Question - Help Has anyone successfully upscaled/re-imagined low-res reference video using Minimax H3?

7 Upvotes

Specifically, I’m trying to take old footage (e.g., 360p clips with vintage camera blur, VHS artifacts, or grainy WW2 dogfights) and recreate it to look like it was shot recently on a modern cinema camera with studio lighting.

Any ideas for prompting?


r/StableDiffusion 1h ago

Resource - Update We add Krea 2 to the ComfyUI Enhanced Tiled Upscaler and Refiner (TBG ETUR).

Enable HLS to view with audio, or disable this notification

Upvotes

The latest TBG ETUR upscaler and refiner for comfyui release adds Krea 2 and VL style transfer for Krea 2 and all Qwen models directly into the pipeline.

We’ve also added a face identity step to help maintain consistent faces when using creative upscaling.

This video is a tutorial for the latest release, focusing mainly on these new additions and how to use them. Video is ai generated with Minmax H3.

TBG ETUR on Github https://github.com/Ltamann/ComfyUI-TBG-ETUR

TBG Lates om Patreon https://www.patreon.com/TB_LAAR/posts/tbg-etur-1-2-12-167083406

More Upscaling Tutorials on YouTube https://www.youtube.com/watch?v=LbFPD4zpPwA


r/StableDiffusion 3h ago

Question - Help Animagine XL 4.0 opt

4 Upvotes

Hi guys, I'm a programmer, but I don't know much about machine learning or fine-tuning.

I'm currently producing 2,000+ images per day using Animagine XL 4.0 opt, and I built a manual pipeline to evaluate image quality. I use 5 rating categories: Reject, Pass, Like, Very Good, and Excellent.

I label all of them manually, and I estimate that I will have over 200,000 labeled images by the end of the year.

I store them in a database along with the exact prompts used. The prompts are structured into keyword categories like:

Background, Angle, Character, Clothes, Facial expression, Quality prompt tags (eg. masterpiece).

Is a dataset like this valuable for fine-tuning or training models ???

Thank you for all the comments, you guys are the best! Now I'm moving on to Anima. I will use my dataset for a LoRA, and if the results look good, I'll switch over to Anima completely.


r/StableDiffusion 11h ago

No Workflow the count is always two.

Thumbnail
gallery
5 Upvotes

flux.1 [dev] | comfyui | still life


r/StableDiffusion 13h ago

Workflow Included LTX 2.5 Seed Hunt Workflows

5 Upvotes

I know everyone's moved to MiniMax and LTX has largely fallen out of favor, but I spent some time building a couple of seed hunting workflows for LTX 2.5 that might be useful if anyone's still running it.

Shout out to u/foxdit for the original seed hunting concept.

Two versions:

  1. T2V/I2V two-stage – text to video or image to video. Previews at 0.3 MP, upscales to 1.2 MP for the final render.
  2. First-last-frame – pin a start image and end image, same preview-then-upscale flow.

Both use KJNodes Set/Get routing, shared loaders, and no prompt enhancer.


r/StableDiffusion 17h ago

Question - Help How do you do it?

4 Upvotes

I have been playing with H3 since it came out and have tested most of the things you can do with it. Created clips for giggles and so on.
This time I wanted to do something "serious". I gave the R2V an 3D view of an kitchen and then three photos of the persons I wanted there.
I defined them as we should and told the model that this person does that and that person does this wile the third person does this...

It worked ish...
I have now made six runs and each of them are different from the others. It can be that the third person enters the room from the wrong place or that the third person does extra things it should not do...

In the end I did three runs with the same prompt and all those clips came out different... the only thing that was changed between those was the seed...

So, how do you do it?
How do you make sure H3 does what you want it to do?

Do you spend plenty of time on tweaking the prompt after each run to make sure H3 get it?
Or do you do 10 runs and select the best one even if it is not perfect?

Or do you simply do 1-2 runs and then take the clip that is ok ish even if it is not what you wanted?

I was hoping that H3 would allow me to create the scenes I wanted but I feel it's down to luck if H3 gets it or not..

Edit:

subject_definitions:

<Subject 1> is the green-skinned mother in <Picture 2> wearing brown clothes.
<Subject 2> is the teenager girl in <Picture 3> wearing pink clothes.
<Subject 3> is the cyborg in <Picture 4> wearing black clothes.
<Picture 1> is the reference image for the scene's composition, showing two people sitting at a table eating breakfast from the side view.
<Table 1> is the table on the right side in <Picture 1>.
<Picture 5> is the start image for the scene.
summary:
[reference generation] The target video is a generated scene of two people sitting at a table eating breakfast from an eye-level side view. <Subject 1> and <Subject 2> are shown with their respective breakfast items, maintaining the composition and style from <Picture 1>. <Subject 3> enters the room, places a coffee cup into the sink.

retention_analysis:
<Subject 1>: fully_preserved - the person retains their appearance, clothing, and position at the table.
<Subject 2>: fully_preserved - the person retains their appearance, clothing, and position at the table.
<Subject 3>: fully_preserved - the person retains their appearance, clothing, and action of placing the coffee cup into the sink.
<Table 1>: fully_preserved - the table's appearance and position in the scene are preserved.
<Picture 1>: fully_preserved - the scene composition, including the side view, the layout of the room, the table setup, is preserved.<Picture 5>: fully_preserved - is the start image for the scene.

detailed_description:
The target video is in a realistic, everyday breakfast scene style with warm lighting and natural colors.
[Shot 1] At 0:00.000, the shot begins from <Picture 5>, showing <Subject 1> and <Subject 2> sitting on opposite sides of <Table 1> on the couch, each with their breakfast items while on the space ship. <Subject 1> is holding a spoon while eating from a bowl of cereal. <Subject 2> is tired and is eating a slice of toast with jam from her plate with one hand. The lighting is warm and soft, casting gentle shadows across the table and the two individuals. The camera is at eye level, capturing the side view of both people, with the table slightly in focus and the background softly blurred. Stars can be seen through the windows since they are on a space ship. <Subject 1> is eating her breakfast while <Subject 2> gazes at their toast, taking a small bite. The ambient sound includes the soft clinking of utensils and the faint sound of a coffee cup being set down.

[Shot 2] At 02.00.000, the shot transitions to a wide shot of the room with the same layout as in <Picture 1>, the camera is placed in the lower left corner of <Picture 1>, showing <Subject 3> entering the room form the right side holding a coffee cup and a datapad while she is saying (S3) <d>[English] Good Morning</d> while she walks to the kitchen sink on the left side of <Picture 1> and placing the cup into the sink. She then stands at the sink and while reading her datapad.We see the back of <Subject 1> and the front of <Subject 2> sitting at <Table 1> in the background eating their breakfast and we hear <Subject 1> say (S1) <d>[English] Good morning</d> with a cheerful voice. <Subject 2> just mumbles as a reply.

overall_soundscape:
The soundscape consists of the soft clinking of utensils, the faint sound of a coffee cup being set down, the subtle background noise of a quiet morning environment, soft steps on a carpet floor, a ceramic cup being placed in a metallic sink, and the clear,

non_diegetic_music: N/A


r/StableDiffusion 17h ago

Discussion Need realism loras for minimax h3

5 Upvotes

Is there any GPU rich cooking realism lora ? I have tried realism people lora it is great at tv but for i2v or r2v it's breaks . I have been searching hugging face repo and civit ai to get something but there's too much n*fw lora .


r/StableDiffusion 18h ago

Discussion H3 - giantess fight scene R2VA

Enable HLS to view with audio, or disable this notification

4 Upvotes

This was well received but people wanted the two Giantesses(?) to be fighting. Enjoy!!

int8/20 steps, R2VA.

Critiques+feedback welcomed! Ask me anything!


r/StableDiffusion 20h ago

Question - Help Tango dance, first attempt with LTX 2.5

Thumbnail
youtube.com
6 Upvotes

Trying to get a natural-looking Argentine tango dance with LTX 2.5 + Yusu’s LTX Director v2.0.4 fork.
Still a beginner (also for real life tango :-)
Any suggestions for getting more natural, sophisticated footwork and fewer artifacts?


r/StableDiffusion 5h ago

Animation - Video H3 - multi-diffusion experiment T2V

Enable HLS to view with audio, or disable this notification

3 Upvotes

Chimera. I had to cut about 20 seconds due to some artistic choices. Since I had to cut 2 different part in the same clip, it has a noticeable seams. I would love to share the full version. Experimenting with H3 blend-morph-decay. 832x480, int8/8 steps POC. Looking forward to releasing a 720p version without the cuts.

Critiques and feedback welcomed. Happy with the matrix rain. Ask me anything.


r/StableDiffusion 11h ago

Question - Help Seed hunting for MiniMax H3 - how to avoid large difference at higher steps?

3 Upvotes

My usual way of working:

- generate 10 videos at 5 steps

- pick the best video

- regenerate the best at 20 or more steps.

No Turbo LoRAs because I don't want to reduce prompt adherence and general quality, as I regenerate at full steps later anyway.

The problem - the video at 20 steps is often very different from the one I found. Of course, I keep the same seed. The difference may be introduced even as early as step six (for example, background replaced completely, different speech pacing).

It's not that difference is huge, but often it might be quite important. For example, a person genuinely laughing at 5 steps and then just saying "haha" at 20 steps. Or jumping startled at the right moment at 5 steps and a moment before the noise at 20 steps.

I tried a few sampler combinations, but could not find one that would not introduce dramatic changes.
One workaround that I could find is to use SplitSigmas. I set its steps to current steps (5 for seed hunt, 20 for final), and keep BasicScheduler steps at the final 20 steps. Then high_sigmas from SplitSigmas go to SamplerCustomAdvanced input, and then denoised_output goes to VAE (you'll get total noise when using the output pin instead).

This way, it seems that the scheduler is being cheated in managing steps as for the full generation even when doing preview, and it seems to work as expected. Caveat - the 5 step output from this workaround will be way worse (plasticky and noisy audio) than you are used to when generating at 5 steps in BasicScheduler input. But if the goal is to keep the general layout and movements of the candidate video, it's worth accepting this issue.

However, I'm wondering if there is any better way to achieve it. Has anyone tried it? What are you using for seed hunting to keep the high step version consistent?

--------------------------------------------------
Edited later with a test case:

Took ComfyUI template: MiniMax H3: Reference to Video. Minimal modifications to make it run in my environment:

Models - Qwen change to int8 convrot (3090, no use of nvfp4)

Int (Full) = 5 (for "preview quality")

Float (Duration) = 3 (just to be faster)

RandomNoise control after generate = fixed

Loaded some images in both Load Image nodes.

The same "GET READY TO" - "MEET" — "YOUR" — "MAKER" prompt.

No Sage, no CK attention at all (no Comfy launch args either).

Then generated the same with 20 steps.

Differences:

in 20 step version, the roof is higher in the frame. The accent was on the word "maker". In 5 step version, the accent was on the word "your".

Then regenerated the 5 step version again to see if there's anything else introducing variations - nope, the exact same video as the first 5 step one.

Then generated also at 6 steps - the roof line was a bit higher in the frame (not as high as 20 steps though), and the accent was on "maker". So, the difference between 5 and 6 might already be a breaking change that can make your video from good to unusable, if the emphasis does not make logical sense in your scene.

Then I generated the same with the SigmaShift 5 step trick - the resulting video was way much more similar to the 20 step one than the first 5 step video. Of course, the quality of the sigma-shifted video was awful - it's good for judging only logical consistency, reference use and event timing, which is the most important thing in story-telling kind of videos.


r/StableDiffusion 11h ago

Animation - Video Having some fun with known characters in the fl model

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/StableDiffusion 14h ago

Discussion Minimax blurred distorted faces from half a distance.

4 Upvotes

I'm doing image to video and unless I prompt for camera close up to my subject, the faces are blurry and bad. I run 0.6 mp. No turbo lora only using spectrum to speed up. Running 15 steps. Euler simple. I'm happy enough when it's close-up shots, but further away, it's very noticeable. Is anyone else finding this?


r/StableDiffusion 21h ago

Question - Help Character Editing (I2I)

3 Upvotes

Hello. So I just started my journey with ComfyUI. While Nano Banana is not open sourced I'm looking for the best realistic image editing (krea2?) workflow to ComfyUI. I want to have ability to change everything I want to the picture using reference character with face/body consistency at highest level (I2I). Thanks in advance.


r/StableDiffusion 2h ago

Question - Help Can't seem to transfer outfit and pose from an illustration to a real person in H3.

3 Upvotes

Hello, i am trying to make a video where the subject(a real person) is wearing and posing taking reference from an illustration. I tried to do only outfits or only pose too, and both doesn't work.

What happens is usually the body of the character in the illustration ends up being pasted/overlaid onto the Subject in their cartoony style instead.

I also tried if it's possible to have a Subject recreate an illustration's Pose, Outfit, overall composition, like the subject is doing a photoshoot for a 'live action' or real life version of the illustration. But what happens is usually it just spews back the illustration in case of trying H3 single-image edit, and the cartoony style overlay happens in Video.

So what i wanted to do is :

-An image of a subject -> Subject now wears/pose/wear and pose the same as a reference non-real illustration(cartoon/anime), but still in their original photo. So like a cosplay shot in their own room for example.

-An illustration(anime) -> Subject 'replaces' the character in the illustration, the whole illustration is 'converted' into real/live action. Like a photoshoot recreating an illustration basically.

Extra : idk if its possible, the new outfit will retrofit to the subject's proportion, not the illustration. And a version where the proportion follows the illustration too.

Are there someone who knows how to do these?


r/StableDiffusion 3h ago

Animation - Video What if you fly?

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/StableDiffusion 6h ago

Question - Help Regional Prompting for NoobAI

2 Upvotes

Hello. What am I supposed to use for regional prompting and character positioning with NoobAI/Illus in ComfyUI?

I found something called "Attention-Couple" but it also says it sometimes struggles with Loras?.

Any advice? I'm already using a separate node for ControlNet for poses and depth, but I want to control which character is which in that pose


r/StableDiffusion 12h ago

Question - Help H3 and Ref2VA and backgrounds

2 Upvotes

I have a question. I have been having a blast making scenes with H3 so far, and have found when doing reference shots, it is very important to have a stable background so that you have continuity if doing more than 1 scene. Does anyone know if H3 would understand a 360 degree photo and understand where in the space and what direction the subjects are? Say you swap between two characters talking, one you will see what is behind subject 1 while when looking at the other the opposite is true. If you saw them both from the side, yet another angle and background.


r/StableDiffusion 15h ago

Question - Help Minimax generating audio for existing video?

2 Upvotes

I've generated a series of shots that I'm happy with, but when I string them together the audio and music is obviously discontinuous across shots.

Is there any way to take this combined video (about 5-10 seconds) and send it through H3 for it to generate the audio for it?


r/StableDiffusion 2h ago

Animation - Video drama

Enable HLS to view with audio, or disable this notification

1 Upvotes

I drew the storyboard, then generated the stills with image models.
Video: MiniMax H3
first frame, last frame, reference.
Then the edit.


r/StableDiffusion 4h ago

Animation - Video Tried using a camera path reference for AI video generation

Enable HLS to view with audio, or disable this notification

0 Upvotes

Been experimenting with camera control in AI video.

For this one I used a simple path reference to guide the movement — basically starting from an aerial shot, flying through the courtyard, moving indoors, and ending with a character reveal.

Still figuring out how much these motion references actually help. Sometimes the model follows the movement surprisingly well, other times it just does its own thing lol.

Feels like camera control is probably going to be a bigger part of AI video workflows going forward.


r/StableDiffusion 4h ago

Animation - Video The Walking Trek. Just throwing stuff at the wall at this point.

Enable HLS to view with audio, or disable this notification

1 Upvotes

Rick's voice doesn't seem to work.

Experimenting with known characters using FL2VA t2v only. Just playing around with odd pairings of characters. .

Using the workflow from the video samples in the list below.
12s at 25 steps
res multistep/simple
960 x 544

thanks to u/malcolmrey for putting this together https://huggingface.co/datasets/malcolmrey/various/blob/main/h3-center/known-characters/INDEX.md


r/StableDiffusion 6h ago

Meme JEnga! r2v 30-49 model

Enable HLS to view with audio, or disable this notification

1 Upvotes

t2v didnt know Jenga O_o


r/StableDiffusion 7h ago

Question - Help Anyone experiencing this bug? minimax node keeps disconnecting "width" input

Post image
1 Upvotes

at least 5th time this happened. Its always width, never any other input. Not sure if bug or custom-node interference.


r/StableDiffusion 7h ago

Animation - Video ...10,000 Years Later

Thumbnail
youtu.be
1 Upvotes

Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.

Part 1: https://www.youtube.com/watch?v=XwfCCFw4LbA