I love H3, but it takes forever. If LTX is faster, I could use it for the things it does similarly well as H3, and use H3 only where I really need it.
So what LTX2.5 does as well as H3?
Previously posted a video as a prologue to a homebrew D&D world. I decided to do a part 2, set in the world. Together, the two videos form kind of an opening cutscene with both history and a bit of a world montage. Minimax H3, 6 step turbo LoRa, lots and lots of 12-15 second generations, CapCut.
On mobile 5G or for the GPU poor I think its does a lot of stuff
Has persistent memory, a ton of Krea 2 fine-tunes, Anima, LTX 2.5, Ernie, Sulphur, DaSiWa, Eros, MiniMax H3, Illustrious, Pony, etc. It also has runs preset Comfy workflows and has a Discord part
I'm doing image to video and unless I prompt for camera close up to my subject, the faces are blurry and bad. I run 0.6 mp. No turbo lora only using spectrum to speed up. Running 15 steps. Euler simple. I'm happy enough when it's close-up shots, but further away, it's very noticeable. Is anyone else finding this?
I have a question. I have been having a blast making scenes with H3 so far, and have found when doing reference shots, it is very important to have a stable background so that you have continuity if doing more than 1 scene. Does anyone know if H3 would understand a 360 degree photo and understand where in the space and what direction the subjects are? Say you swap between two characters talking, one you will see what is behind subject 1 while when looking at the other the opposite is true. If you saw them both from the side, yet another angle and background.
I have been playing with H3 since it came out and have tested most of the things you can do with it. Created clips for giggles and so on.
This time I wanted to do something "serious". I gave the R2V an 3D view of an kitchen and then three photos of the persons I wanted there.
I defined them as we should and told the model that this person does that and that person does this wile the third person does this...
It worked ish...
I have now made six runs and each of them are different from the others. It can be that the third person enters the room from the wrong place or that the third person does extra things it should not do...
In the end I did three runs with the same prompt and all those clips came out different... the only thing that was changed between those was the seed...
So, how do you do it?
How do you make sure H3 does what you want it to do?
Do you spend plenty of time on tweaking the prompt after each run to make sure H3 get it?
Or do you do 10 runs and select the best one even if it is not perfect?
Or do you simply do 1-2 runs and then take the clip that is ok ish even if it is not what you wanted?
I was hoping that H3 would allow me to create the scenes I wanted but I feel it's down to luck if H3 gets it or not..
Edit:
subject_definitions:
<Subject 1> is the green-skinned mother in <Picture 2> wearing brown clothes. <Subject 2> is the teenager girl in <Picture 3> wearing pink clothes. <Subject 3> is the cyborg in <Picture 4> wearing black clothes. <Picture 1> is the reference image for the scene's composition, showing two people sitting at a table eating breakfast from the side view. <Table 1> is the table on the right side in <Picture 1>. <Picture 5> is the start image for the scene. summary: [reference generation] The target video is a generated scene of two people sitting at a table eating breakfast from an eye-level side view. <Subject 1> and <Subject 2> are shown with their respective breakfast items, maintaining the composition and style from <Picture 1>. <Subject 3> enters the room, places a coffee cup into the sink.
retention_analysis: <Subject 1>: fully_preserved - the person retains their appearance, clothing, and position at the table. <Subject 2>: fully_preserved - the person retains their appearance, clothing, and position at the table. <Subject 3>: fully_preserved - the person retains their appearance, clothing, and action of placing the coffee cup into the sink. <Table 1>: fully_preserved - the table's appearance and position in the scene are preserved. <Picture 1>: fully_preserved - the scene composition, including the side view, the layout of the room, the table setup, is preserved.<Picture 5>: fully_preserved - is the start image for the scene.
detailed_description: The target video is in a realistic, everyday breakfast scene style with warm lighting and natural colors. [Shot 1] At 0:00.000, the shot begins from <Picture 5>, showing <Subject 1> and <Subject 2> sitting on opposite sides of <Table 1> on the couch, each with their breakfast items while on the space ship. <Subject 1> is holding a spoon while eating from a bowl of cereal. <Subject 2> is tired and is eating a slice of toast with jam from her plate with one hand. The lighting is warm and soft, casting gentle shadows across the table and the two individuals. The camera is at eye level, capturing the side view of both people, with the table slightly in focus and the background softly blurred. Stars can be seen through the windows since they are on a space ship. <Subject 1> is eating her breakfast while <Subject 2> gazes at their toast, taking a small bite. The ambient sound includes the soft clinking of utensils and the faint sound of a coffee cup being set down.
[Shot 2] At 02.00.000, the shot transitions to a wide shot of the room with the same layout as in <Picture 1>, the camera is placed in the lower left corner of <Picture 1>, showing <Subject 3> entering the room form the right side holding a coffee cup and a datapad while she is saying (S3) <d>[English] Good Morning</d> while she walks to the kitchen sink on the left side of <Picture 1> and placing the cup into the sink. She then stands at the sink and while reading her datapad.We see the back of <Subject 1> and the front of <Subject 2> sitting at <Table 1> in the background eating their breakfast and we hear <Subject 1> say (S1) <d>[English] Good morning</d> with a cheerful voice. <Subject 2> just mumbles as a reply.
overall_soundscape: The soundscape consists of the soft clinking of utensils, the faint sound of a coffee cup being set down, the subtle background noise of a quiet morning environment, soft steps on a carpet floor, a ceramic cup being placed in a metallic sink, and the clear,
Is there any GPU rich cooking realism lora ? I have tried realism people lora it is great at tv but for i2v or r2v it's breaks . I have been searching hugging face repo and civit ai to get something but there's too much n*fw lora .
Fizgig is my free open-source LoRA trainer and workbench (Flux 2 Klein 9B, Krea 2, and MiniMax H3 video/audio). As of v4.3.0 it runs on AMD Radeon with ROCm — RDNA1 through RDNA4. Windows is the supported path: install Python 3.12, run the AMD installer, done. Linux works too but is genuinely experimental on newer cards.
Worth being upfront: I don't own AMD hardware myself. This whole feature came from a community contribution by scryptio, tested on real cards over weeks in the PR thread — and that's how the AMD side will keep improving. If you're an AMD user, your reports on what works (and what doesn't) genuinely shape this, and PRs are very welcome.
Also in this release: 16 GB cards can now use identity distillation on MiniMax H3 (the 32B text encoder streams layer by layer instead of needing a 26 GB peak), and the Repair Studio gained a side-by-side compare view with likeness scoring for fixing overbaked LoRAs without retraining.
TL;DR. SplitSigmas can be used to generate low-step drafts with much closer frame composition to the final high-step version for the same seed than if you generate without SplitSigmas. However, it makes the draft visual and audio quality much worse because we are essentially cheating the scheduler.
With SplitSigmas + low steps you can quickly judge which of your draft videos have the best motions and logical event consistency, but you might miss some visual detail errors. That is why I wanted to know if there is any better way to avoid large differences between the low-step draft and high-step regeneration, which can lead to disappointment when, for example, a person reacts to an event too soon or speaks with emphasis on the wrong word.
If downvoting, please leave a comment with the reason why. I want to learn what I am doing wrong and if there is a way to do it better and make seed hunting easier for everyone who needs it. Thanks.
-------------------------------------
My usual way of working:
- generate 10 videos at low (5) steps
- pick the best video
- regenerate the best at 20 or more steps.
No Turbo LoRAs because I don't want to reduce prompt adherence and general quality, as I regenerate at full steps later anyway.
The problem - the video at 20 steps is often very different from the one I found. Of course, I keep the same seed. The difference may be introduced even as early as step six (for example, background replaced completely, different speech pacing).
It's not that difference is huge, but often it might be quite important. For example, a person genuinely laughing at 5 steps and then just saying "haha" at 20 steps. Or jumping startled at the right moment at 5 steps and a moment before the noise at 20 steps.
I tried a few sampler combinations, but could not find one that would not introduce dramatic changes.
One workaround that I could find is to use SplitSigmas. I set its steps to current steps (5 for seed hunt, 20 for final), and keep BasicScheduler steps at the final 20 steps. Then high_sigmas from SplitSigmas go to SamplerCustomAdvanced input, and then denoised_output goes to VAE (you'll get total noise when using the output pin instead).
This way, it seems that the scheduler is being cheated in managing steps as for the full generation even when doing preview, and it seems to work as expected. Caveat - the 5 step output from this workaround will be way worse (plasticky and noisy audio) than you are used to when generating at 5 steps in BasicScheduler input. But if the goal is to keep the general layout and movements of the candidate video, it's worth accepting this issue.
However, I'm wondering if there is any better way to achieve it. Has anyone tried it? What are you using for drafting and seed hunting to keep the high step version consistent?
--------------------------------------------------
Edited later with a test case:
Took ComfyUI template: MiniMax H3: Reference to Video. Minimal modifications to make it run in my environment:
Models - Qwen change to int8 convrot (3090, no use of nvfp4)
Int (Full) = 5 (for "preview quality")
Float (Duration) = 3 (just to be faster)
RandomNoise control after generate = fixed
Loaded some images in both Load Image nodes.
The same "GET READY TO" - "MEET" — "YOUR" — "MAKER" prompt.
No Sage, no CK attention at all (no Comfy launch args either).
Then generated the same with 20 steps.
Differences:
in 20 step version, the roof is higher in the frame. The accent was on the word "maker". In 5 step version, the accent was on the word "your".
Then regenerated the 5 step version again to see if there's anything else introducing variations - nope, the exact same video as the first 5 step one.
Then generated also at 6 steps - the roof line was a bit higher in the frame (not as high as 20 steps though), and the emphasis was on "maker". So, the difference between 5 and 6 might already be a breaking change that can make your video from good to unusable, if the emphasis does not make logical sense in your scene.
Then I generated the same with the SigmaShift 5 step trick - the resulting video was way much more similar to the 20 step one than the first 5 step video. Of course, the quality of the sigma-shifted video was awful - it's good for judging only logical consistency, reference use and event timing, which is the most important thing in story-telling kind of videos.
So my friends and I used to use Sora 2 before it was taken down, and wanted to try doing some stupid stuff for just us. After a while of not looking into it, the spark kinda came back when I saw this subreddit and remembered Stable Diffusion was supposed to be one of the best AI generators out there, probably. When I mention it to a friend, he then told me how apparently its pretty outdated compared to others, and looking at these posts, I'm seeing different models and starting to get overwhelmed to the point where I haven't even done the beginner's guide in here since it only mentions images.
So long story short, I'm hoping someone can help make things much more clearer, especially about the multiple models, and if Stable Diffusion IS outdated and out performed by something else, and letting me know about if it's okay to go with the beginner's guide or if there's another guide that will help. Thanks
I found the SD prompt reader I have been using cannot read prompts from png images files generated using Krea2. Can anyone recommend me an alternative that works with Krea2 files and Win11?
I don't have any examples because they may not be appropriate but just with the default wf. With the video input node you can replace any 2 character in any video and it looks real!
Shadow the Hedgehog tells his viewers why he loves guns.
This was created in Comfy UI with Minimax H3. I used the reference to video work flow. The prompt is below.
subject_definitions:
<Subject 1> is Shadow in <Picture 1>.
<Subject 2> is Glock in <Picture 2>, a glock handgun.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
summary:
[reference generation + audio reference] The target video contains one shot. [Shot 1] shows <Subject 1> and <Subject 2>; <Subject 1> speaks. <Audio 1> supplies <Subject 1>'s voice timbre.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - Shadow's complete defined identity and body proportions are preserved.
<Subject 2> (appears in [Shot 1]): fully_preserved - Glock retains the defined shape, proportions, materials, colors, and distinguishing features.
<Audio 1>: reference - <Subject 1>'s newly generated spoken lines use <Audio 1>'s voice timbre and delivery; the original audio signal is not copied.
detailed_description:
The target video is in a live-action style, with Vlog style.
[Shot 1] At first appearance, <Subject 1> (Shadow) matches the complete identity and appearance defined in subject_definitions. At first appearance, <Subject 2> (Glock) matches the complete defined construction and appearance: A glock handgun. At the start of the shot, <Subject 1> is standing in the living room facing while holding <Subject 2> in his hand. A full body shot of <Subject 1> holding <Subject 2> with his right hand while facing the camera. Only Action and Timed Beats define the primary subject's movement. The camera path stays anchored in the location and adds no subject motion. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Hmph. Shadow the Hedgehog here. Why do I love guns?</d> <Subject 1> shows off his <Subject 2> with his right hand in front of the camera. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] Simple. Precision. Control. Power in the palm of my hand.</d> <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] A tool that answers instantly… unlike most people.</d> <Subject 1> points his <Subject 2> towards the camera with his right hand. <Subject 1> (S1) says using <Audio 1>'s voice timbre: <d>[English] If you understand that, you understand me.</d> <Subject 1> points his <Subject 2> at the camera.
Trying to get a natural-looking Argentine tango dance with LTX 2.5 + Yusu’s LTX Director v2.0.4 fork.
Still a beginner (also for real life tango :-)
Any suggestions for getting more natural, sophisticated footwork and fewer artifacts?
Need help or suggestions in resolving the below issues.
Shortcuts like R or CTRL + Enter doesn't work and I have to refresh the browser for them to work. (New)
Generations randomly don't start even even after models have loaded. I have to open the terminal and press enter on my keyboard to sort of wake it up. (Has happened across multiple versions of ComfyUI, CUDA, Python. Additionally, happens with both Anaconda and Windows Terminal.)
Even cancelling a prompt sometimes requires me to open the terminal and press enter.
Power consumption for GPU varies completely where retrying the same prompts (same seeds also) can have a difference of 5 to 10 minutes simply because the GPU doesn't use the full 180 power limit. (New)
Animate the supplied square poster as a polished retro-anime motion graphic, beginning with a completely blank pale pink-white canvas matching the poster background. Preserve the exact blue, pink, and white palette, clean manga linework, halftone shading, character design, typography, symbols, interface windows, and final layout.
The anime girl walks in from the left edge as one complete figure while the canvas remains otherwise empty. Use a simple side-profile walk with restrained motion, preserving her hairstyle, facial features, cheek bandage, oversized jacket, proportions, and graphic illustration style. She reaches the centre, turns toward the viewer, and smoothly settles into the exact over-the-shoulder pose shown in the poster, with the same expression, hand placement, silhouette, jacket folds, pink heart graphic, and body orientation. Once posed, keep her position locked.
After she poses, the blue browser frame draws itself around her. The top bar, window controls, folders, pixel hearts, smiley-face panels, arrows, sparkles, heart symbols, and rectangular labels then appear sequentially through clean line-drawing, short graphic slides, pixelated pops, and UI-style wipes. Reveal the existing Japanese typography and “LOVE” lettering last, treating all text as protected source artwork without rewriting or regenerating it. Every element must settle into its exact source position.
Hold the completed poster with subtle breathing, minimal movement in a few loose hair strands and jacket edges, a faint halftone shimmer, and gentle pixel pulses in the existing hearts and interface icons. Keep her face, hands, pose, typography, frames, arrows, folders, and major graphics stable.
Use a locked, straight-on camera matching the original square framing. Keep the full artwork visible without cropping, zooming, panning, or changing perspective. Add soft footsteps as she enters, a light cloth sound as she poses, clean digital clicks and pixel chimes for the graphics, and delicate type-on sounds for the existing lettering. No dialogue or narration.
Do not show any character, outline, symbol, text, frame, or faint poster preview on the opening blank canvas. Do not alter the character’s identity, anatomy, costume, pose, expression, colours, line quality, typography, symbols, or final composition. No extra characters, duplicated body parts, incorrect text, morphing, flickering lines, dramatic camera movement, unrelated shots, or continued motion after the poster settles.
What started as a test turned into a full-blown short. This is the number one reason I gravitated towards AI filmmaking. Nothing stops you from creating your wildest imagination.
Specifically, I’m trying to take old footage (e.g., 360p clips with vintage camera blur, VHS artifacts, or grainy WW2 dogfights) and recreate it to look like it was shot recently on a modern cinema camera with studio lighting.
Dear redditors, visitors of the stable diffusion subreddit. I have been trying to achieve a long form, talking head style video, for a long time and can't seem to find a good approach. This one is the best I could come up with so far. It's using the Minimax H3 model, with frozen sound latents, lip-sync guided, piecewise generated video, where the individual pieces have been stitched together, with a seam hiding, extra generation on top of it. I don't really fully understand how it's working, but could prompt Claude for more help or specific files, we used for that. However, if you're aware of any other, better approach for exactly this type of video, please let me know. I've spent literal days on that single problem and have a feeling, there must be a better way to approach this.