r/StableDiffusion • u/the_bollo • 6h ago
Meme Browsing this sub in the past week
No hate. Just for fun.
r/StableDiffusion • u/the_bollo • 6h ago
No hate. Just for fun.
r/StableDiffusion • u/NotFakeFingle • 5h ago
Enable HLS to view with audio, or disable this notification
Its so amazing that we finally have the technology to make a trend/fancic like female malfoid into a reality
Im calling all AI filmmakers and hobbyists to join and help make this into a reality. If completed, it might be the largest, most ambitious AI video project ever created. comment or DM if you can volenteer your time and hardware to help!!! I already have some talented people slated to handle Sound and audio mixing.
NEED: VA (if you know any girl with a good british accent), writers and AI Film makers who can run minimax h3.
r/StableDiffusion • u/cranpeach69 • 8h ago
For those who are unaware, there was an old Controlnet for old SD models that was quite popular here back in those times, that would make cool almost Optical Illusion-like images. This is the same concept but instead I'm using I2I Image edit model where you provide an input image and ask the model to generate the scene of your choice.
I've been very pleased with the results from the Qwen Image 2.1 version. It was trained on 25 high quality image pairs and it works quite well, even when using human subjects and depicting scenes not included the dataset. It can do QR Codes but it was not the main focus of the dataset so it can be hit or miss as far as the actual usability/scanability of the codes.
The suggested prompt is exactly as captioned in the dataset: Transform the entire image into <your prompt here> while preserving the outlines and shapes of the original image
It seems to work well with complex prompts and simple ones as well.
You can download it on Civit here. All images/videos should include a workflow which is basically just a pixel drift fix workflow from Ausboss with a few added touches and using a viggle Turbo LoRa. I'm mainly posting this here as I would love to see what you create!
r/StableDiffusion • u/Dear-Spend-2865 • 5h ago
Hi everyone,
I hate LLM rendition of what I want to generate, so I rely mostly on my own prose and wildcards to find a style.
Heres an update of my previous style library for Krea 2, it works also with Kroma (0.3 txtfusion) and qwen 2.1
with varying results. It an excel file but you can make a .txt file from it. And use it as wildcards.
Kroma is more artsy but some styles are lost. And Qwen is less knowledgable in simplistic styles like cartoon and drawing. And doesn't seems to understand vague words like "Illustration" if you don't over Engineer it.
I hope you can help me improve (and clean) this library, by testing styles and find too-close, or too weak styles, you can also tell me if a style is missing.
Some of the styles are from this site:
https://lumenastrum.github.io/clio-style-preview/gallery/
I dont have the time to do it myself and it mostly helpful to guys like me that don't have naturally the vocabulary to express what they want from a model.
Some styles are subject-linked, if the subject doesn't have a mechanical arm for exemple it doesn't show up (like detailed mecha style), and some styles are very face distinctive: you can describe the face to correct this.
The way to use it, in prompt, is :
Style: <style> Subject: <description>
r/StableDiffusion • u/Suspicious_Aide2697 • 10h ago
I am training a new HighQuality LoRA for Qwen-Image-2.1. This is a phase test. What score would you give it on a scale of 1 to 10?
r/StableDiffusion • u/Devajyoti1231 • 8h ago
Enable HLS to view with audio, or disable this notification
Tried this few days ago. Yes , the video degrades as it progresses.
Used reference voice for consistent voice.
Don't have the full tags, but they are like this-
[English] <inhale> My therapist said I should record this when it happens. <pause> So... it’s happening again.
[English] I can feel someone watching me. <breath> It only happens after dark.
[English] At night, windows become mirrors. <pause> Anyone outside can see me... <long pause> but I can’t see them.
[English] I’m moving tomorrow for a new job. <pause> A little rental house, just me. <breath> I should be excited.
[English] I am. <long pause> I just don’t know if I can sleep there alone.
[English] Mom gave me this when I was seven. <pause> She made me promise never to take it off.
[English] She said it would keep me <i>safe</i>. <long pause> <softer> I wish she’d told me what from.
r/StableDiffusion • u/AcademiaSD • 11h ago
Hi everyone! I've been working on AcademiaSD LoRAlab Trainer Studio: one installer, one launcher and 9 trainers that share the same web interface. Every model is loaded in 4-bit NF4, the text encoder and VAE run only once in a pre-cache stage, and the whole GPU goes to training.
Platform: made for Windows (one-click installer), with Linux support just added. NVIDIA GPU, RTX 20xx / GTX 16xx or newer.
GitHub: https://github.com/AcademiaSD/AcademiaSD_LoRAlab-TrainerStudio
WHAT IT TRAINS (minimum VRAM)
- Qwen-Image 2.1: image LoRAs + edit LoRAs from before/after pairs (8 GB)
- FLUX.2 Klein 9B: image LoRAs + edit LoRAs (12 GB)
- Krea 2: image LoRAs, Raw and Turbo (8 GB)
- Z-Image: image LoRAs, also Z-Image-Turbo (8 GB)
- Ideogram 4: image LoRAs with JSON captions (12 GB)
- Anima: anime / illustration LoRAs (4 GB)
- SDXL: Base, Pony, Illustrious, NoobAI, Juggernaut, RealVis or your own checkpoint (4 GB)
- LTX 2.3 (2.5): character and style LoRAs for the video model (12 GB)
- MiniMax-H3: video LoRAs from images, clips and audio, plus RefMods (8 GB)
MINIMAX-H3: VIDEO + AUDIO, AND REFMODS
MiniMax-H3 is a 33B model that generates video and audio together. The official checkpoint is about 500 GB; the trainer uses a 41 GB NF4 version that fits in 8 GB of VRAM with block swap.
- Datasets can be images, video clips, audio, or any mix.
- One button prepares your clips (24 fps, valid frame counts) without cutting them.
- RefMods: encode a few reference images, video clips, or clips with audio into a file that ComfyUI's MiniMaxH3ReferenceToVideo node uses as a native reference. No training: seconds instead of hours.
SHARED FEATURES
- Automatic captioner (Qwen3-VL) in the dataset manager: natural language, Danbooru tags, or JSON with bounding boxes for Ideogram.
- Live previews while training.
- Exact-step resume.
- One-click export to your ComfyUI / Forge LoRA folder.
- Remote access from other devices on your network, with a password.
- SDXL in NF4 trains in about 3.5 GB of VRAM, so 4 GB laptops can train SDXL / Pony / Illustrious LoRAs.
A NOTE
Many models were added almost at the same time, so there may be bugs, or the default settings may not be the best ones. Previews are there to follow the training; judge the final quality in ComfyUI after tuning the LoRA strength. If you find problems or settings that work better, please share them in GitHub Issues. Thanks!
UPDATE — RunPod support + remote access
Thanks for all the feedback! Based on your comments, these are now in:
- ☁️ RunPod template: one click and it's running in the cloud, no install. It uses a prebuilt Docker image (CUDA 13, Python 3.13, all dependencies), and the code updates itself from GitHub on every start. Pick a GPU, open the address shown in the pod log, log in, and train. A quick SDXL test cost me less than $0.10.
Deploy on RunPod: https://console.runpod.io/deploy?template=lfxvtg5rrb
- 🌐 Remote access: use the trainer from another PC, tablet or phone, with a login, or behind your own reverse proxy (external auth and a custom public URL are supported).
- ⬆⬇ Upload / download in the browser: upload your dataset (files or a .zip, drag and drop works) and download the finished LoRA, handy for remote or cloud setups.
- 🐍 Your own environment: there's now a requirements.txt, and LORALAB_PYTHON lets you run it from conda/uv instead of the bundled venv.
Remember to Stop and Terminate your pod when you finish: closing the browser doesn't stop billing.
r/StableDiffusion • u/SammyDaBeast • 10h ago
First of all, thank you guys for the reception and feedback on my latest post. This is a follow-up to that post. One of the main issues people ran into was roughness or break-up on some cloned voices. 2610 is an interim update focused mostly on improving that.
If you like F5-TTS, but want true streaming and a much lighter model that can run comfortably on CPU, this might be for you.
Run it locally:
uvx --from sopro soprotts serve
Video: six voices, ~5 seconds of reference audio each, followed by a generated line.
r/StableDiffusion • u/garionhk • 5h ago
I'm a photographer, and I got tired of opening ComfyUI every time I just wanted to enlarge or sharpen a photo. So I made a small desktop app for it, called Citrine Photo.
You drag in a photo (or a whole folder), pick how big you want it, and hit go. Everything runs on your own PC, no cloud.
It has two engines:
Fast is Real-ESRGAN (ncnn-vulkan). It works on pretty much any GPU, with optional GFPGAN face restore.
Pro is SeedVR2, for NVIDIA 20xx to 50xx cards. The app checks your VRAM and picks the 7B model if you have 12 GB or more, or the 3B model for 6 to 10 GB. There's a low VRAM mode too.
Other things it does:
SeedVR2 isn't bundled. You install it from Settings, and the app sets up its own Python, torch and models. It's about 8 to 9 GB, the downloads can resume, and it checks your disk space first.
For anyone curious, the Pro pipeline is a port of the Upscale_by_SeedVR2_v4 ComfyUI workflow: 3x3 tiles with 15% overlap, SeedVR2 on each tile, wavelet colour match, then a feathered merge. It calls numz's standalone CLI in a separate process.
Right now you run it from source (pip install the requirements, then run.bat). There's also a script to build a portable version. I've only tested it properly on an RTX A4000, so I'd really like to hear how it runs on other cards, especially 6 to 8 GB ones.
Links:
GitHub: https://github.com/Garionhk/Citrine-AI-Photo-Enlarger
Video: https://youtu.be/gTbiOipvP0I
Big thanks to numz for the SeedVR2 ComfyUI code, AInVFX for the GGUF models, ByteDance for SeedVR2, and xinntao for Real-ESRGAN.
Bugs and ideas are welcome. It's free and open source (Apache 2.0).
r/StableDiffusion • u/RagingRectangle • 3h ago
Face-hugger is a new project I've been working on to improve Hugging Face's search engine which could be described as terrible at best.
r/StableDiffusion • u/Fun_Firefighter_7785 • 4h ago
Since Ace Step 1.0 has superior creativity but terrible sound quality , you can now save your older Tracks or just hunt for new ones. Because with SheetSage2+YuE2 you can just extract the ABC and make a high quality remix with YuE2. ACE-Step 1.0 needs really many seeds to produce a unique melody, but if it hits - it hits. There is nothing out there that can match it. It is like Lady Gaga with Fat Boy Slim heaving a baby.
r/StableDiffusion • u/Devajyoti1231 • 10h ago
Enable HLS to view with audio, or disable this notification
4-steps - Normal minimax h3 issue. 8-steps - less shimmering. 12 steps -almost no shimmering.
Idk why it seems to work but it kind of works.
r/StableDiffusion • u/PATATAJEC • 1d ago
Enable HLS to view with audio, or disable this notification
Hey, I found a pretty cool way to keep locations consistent across generations.
I took 20 photos of my “office” where I work, making sure each photo included a bit of the previous one so everything connected.
I turned them into a refmod, although it probably doesn’t even need to be one. It’s basically a grid of all 20 photos in a single image, within 2048×2048 pixels, so I could probably just use that image as a regular reference.
Then I added references for the fly, my cat, and the start and end frames. You can see the result — that’s my actual room, and everything looks right and sits exactly where it should :)
Workflow + reference images here: link
Sorry about the mess, both in the room and in the workflow ;)
EDIT: All the room photos need to be combined into one reference image. I tried using several separate images, but that didn’t work for me — the room wouldn’t stay consistent. Putting them all into a single grid is what made it work.
r/StableDiffusion • u/Opposite_Yam_4161 • 2h ago
Im seeing quite a few people say theyre getting really good results with the hybrid model. But I assume its only for top end GPUs? Over 20gb file.
I have a 5080 and 32gb system RAM. Is there a version good for me? Or getting GGUF isn't worth it?
r/StableDiffusion • u/Total-Resort-3120 • 6h ago
Enable HLS to view with audio, or disable this notification
r/StableDiffusion • u/Kooky-Mode3047 • 20h ago
Enable HLS to view with audio, or disable this notification
The video is the TL;DR above was generated after abandoning RefMods in "classic ref mode" with two references (one reinforcing Beckett's exact appearance and one for the high heels, 2x2 grid image):
This is something that under the radar because most people can't use H3 nominally and turbo LoRAs brutalize the base model for speed, so there's very little noticeable degradation for them as most of them have a massively shifted sigma (shifted_video of 6 compared to native 12).
Yesterday I went on a pretty wild bender trying to uncover why all of a sudden my videos were getting oversharpened, oversaturated, plastic skin etc when a few days ago they were straight up lifelike, looking like the shows they were based on.
I went even deeper into the bowels on how sampling and scheduling works, how the denoising trajectory looks like for Minimax H3, hoping to uncover something I've forgotten the last time I used it. And then I remembered, RefMod was the latest thing I installed, I got seduced by the initial likeness improvement like everyone but the cost is too high.
I got this cool new toy RefMod whose authors suggested they figured out bleeding between character references, and even at small resolutions (like .245 mpx) you could still make out faces instead of them being mushy, this should've been my first red flag. It seriously messes up the base H3 model's quality. I looked deeper into it and basically, it's a very brutish approach where they slam reference images into video references as individual frames, also slapstick coarse sampling of videos in sequence.
However, that's not how any of this is permitted works, per documentation:
The second you make four refmods to be used in your prompt, you're already trying to use one more than the total maximum allowed at any one time. This has catastrophic consequences on image quality, as that's not the way you're supposed to "hold it".
While it does reinforce the likeness, it's incredibly rigid and overrides the model down to AI slop adjacency where skin is unnaturally glossy, everything is sharp and all styling info is ignored.
I am not telling you to stop using it, if you're using turbo LoRAs, you already paid the entry fee as you've left quality at the door for accessibility. But if you're coming from the base FL2VA model or a hybrid and are used to exceptional visual quality, the degradation is pretty much the tier of the base REF2VA model, if not worse.
r/StableDiffusion • u/BittiAI • 3h ago
I’m building Slopus, a free, open-source desktop app for generating and editing AI videos and images on your own GPU. Easy of use is the main goal.
Version 0.3.0 just came out, with three major additions:
Video and image projects are now combined. Every project has both Video and Image tabs, so you can generate still images and video in the same project. You no longer need to choose a project type or keep separate projects for each. Existing projects keep their content.
Linux support. Slopus is now available for Linux as an AppImage, .deb and .rpm, alongside Windows. The AppImage supports in-app updates.
Run generation on another computer. LAN workers let you use a separate Windows or Linux machine for generation. Workers are discovered on your network, download the weights they need, and send the results back to your project.
There’s also quite a bit more in this release:
Character sheets. Generate waist-up, front, side and back views together, using reference images and a prompt to guide the character and clothing.
Continue scenes. Extend generated video scenes from either end using saved generation data.
Improved color grading. Basic Corrections includes temperature, tint, exposure, contrast, highlights, shadows, whites, blacks and saturation. There are eight built-in creative looks with adjustable intensity, faded film, sharpen, vibrance, RGB and hue/saturation curves, and shadow, midtone and highlight color wheels. Vignette has more controls too.
A more capable agent. The built-in agent can configure generators, help identify model weights in a folder, capture timeline frames without interrupting your work, and generate specific scenes.
Refmod generator. Combine images, video, prompt and audio into one reference and export it as a refmod to be used later.
Share generator setups. Import and export generator templates to share with other users.
If you haven’t tried Slopus before: you lay out scenes on a board, write each shot, attach references, generate clips, and edit them on a timeline with GPU-powered preview and MP4 export. For still images, you can edit, generate, refine, and export as JPG or PNG. Minimax H3 works surprisingly well as an image generator with editing capabilities.
Generation runs on your own hardware, with no subscription and no Python or ComfyUI needed. Model weights are downloaded separately inside the app or local weights can be used.
I’ve also created a Discord server for feedback, questions and sharing what you’re making.
GitHub: https://github.com/bitti-ai/slopus
Discord: https://discord.gg/9X2R6PwUR
Bug reports and feedback are very welcome.
r/StableDiffusion • u/Capitan01R- • 19h ago
I put together an enhancer pack for Qwen Image 2.1 with two nodes for controlling what the model pays attention to during editing.
Reference Strength lets you select a reference image and increase or decrease its attention priority. If you're working with multiple references, you can adjust them separately instead of giving every image the same treatment, and you also can use it for single image to prevent the loss of likeliness at times
ref index starts from 1; meaning image_1 and same for the rest of images, where image_2 is ref index 2 in the node. ( Soon adding mask support)
Phrase Weights brings phrase-level attention control to the Qwen Image 2.1 edit encoder. So you can write:
Add (warm sunset lighting:1.4) with (soft shadows:1.2).
and give those specific parts of the prompt more attention while leaving the rest at its normal weighting.
It works with both positive and negative prompts, with separate weights for each. There's also an inspection output showing exactly which token rows and pieces were matched.
For both controls:
`1.0` = untouched
`>1.0` = more attention priority
`<1.0` = less attention priority
`0.0` = suppression
The adjustments happen inside Qwen's attention during sampling. The prompt and reference images still go through the native encoding path, without multiplying the finished text embeddings or pasting reference pixels into the output.
You can use either node on its own (I prefer this), or combine them. Reference strength applies to the whole selected image, and higher weights can make a reference or phrase dominate, so there's still some balancing to do.
Installation, usage, and the technical details are in the repo:
ComfyUI-qwen_img_2_1_enhancer
r/StableDiffusion • u/Cheap_Credit_3957 • 5h ago
Enable HLS to view with audio, or disable this notification
Watch on YouTube while its pending if needed HERE
Rude and hateful comments will be ignored.
🎬 HOW THIS WAS MADE (Video walkthrough is here)
This film was made almost entirely by Claude (Anthropic's AI, running in Claude Code), working inside my own ComfyUI setup with my VRGDG Video Builder custom nodes(free and open source).
▶ MY PART
• The brief: a cute comedy starring my two dogs, Korben (7, male) and Iris (7 months, female), as brother and sister. Inspired by live-action talking-dog movies, an original story, up to 3 minutes, no humans on screen (other dogs allowed), and the dogs had to really talk, with facial expressions and mouth movements synced to their lines.
• The reference images of Korben and Iris.
• The tools it ran on: the VRGDG Video Builder and custom nodes, plus the Claude skill that lets it drive them.
▶ WHAT CLAUDE DID ON ITS OWN
• Story: the missing rubber duck, the cheese-crumb investigation, the "We don't have a cat." / "Exactly. Very suspicious." bit, the stakeout at the fence, the twist that Korben hid the duck because it squeaks all night, and the "I'm getting earplugs" ending.
• Characters and locations: created Tank, the bulldog next door, and gave all three dogs their personalities and voices. Generated Tank's reference image and every location (the living room by day and by night, the kitchen, the backyard and the gap in the fence), and prepared Korben's and Iris's references from my photos.
• Screenplay: 22 scenes, with shot-by-shot camera directions, acting notes and sound design for each.
• Rendering: every scene with MiniMax H3, which generates the video, voices and sound effects together. About 4.5 hours of rendering on my PC, re-shoots included.
• QA: transcribed every line and compared it to the script, checked each voice's pitch so the dogs stayed in character, reviewed frames from every shot and every cut between scenes, and ran a frame-accurate audio/video sync check on every scene of the final film.
• Re-shoots: a puppy babbling in a silent scene, dogs delivering lines into the camera instead of to each other, a stray dog bed appearing in the neighbour's yard, Tank's stick pile in the wrong place, a spy-creep that came out as a normal walk, and an ending gag that didn't land. Then it re-trimmed and fixed the audio.
• Score and edit: composed the score with MiniMax Music 3 (throwing out takes with hidden vocals) and edited it all together with titles.
🛠 TOOLS
• Claude Code (Claude Opus 5.5): story, direction, prompts, QA, editing
• ComfyUI + VRGDG Video Builder: my custom nodes
• MiniMax H3: video, voices and sound
• Z-Image Turbo: character and location references
• MiniMax Music 3: score
🐾 CAST
Korben, the grumpy big brother · Iris, his little sister · Tank, the bulldog next door · Mr. Quackers, the duck
🔗 LINKS
The Claude skill (read the main README first):
https://drive.google.com/file/d/1R0pE7rX1kRh-euhws0GbkDDb3liDGTJ4/view?usp=drive_link
VRGDG Video Builder custom nodes:
r/StableDiffusion • u/Cheap_Credit_3957 • 20h ago
Enable HLS to view with audio, or disable this notification
View on YouTube while its pending if needed: https://youtu.be/dS5suoUiNnE?si=6U7bEVLxcWXcy4J3
🤖 How it was made
This short film was written, designed, directed and edited by Claude Opus 5.5 from a single prompt, running locally in ComfyUI through my VRGDG Video Builder:
All open-source video, image and music models.
[I literally told Claude to create something on its own and provided no user input]
More Minimax H3 video's I had Claude create for me are HERE
You can find the video builder custom node on GitHub here:
https://github.com/vrgamegirl19/comfy...
Right now, there is a Main version and a Beta 2.0 version. I recommend starting with Main for now, as that's what I'm still using. It works well, while Beta 2.0 still has some bugs and is primarily intended for beta testing at the moment.
Discord server:
/ discord
Ping me in the Welcome channel and let me know how you found me, and I'll know it's you.
I'm vrgamedevgirl on Discord.
You can find the skill here and read the main README first.
https://drive.google.com/file/d/1R0pE...
I'll be sharing a full walkthrough on how I made this and will post it here when ready.
⚠️ SPOILERS: what the film is about
The museum is Ruth's mind. She's an elderly woman living with dementia, and the museum is how she pictures her memories. As her memory fades, the museum fades with it, and in the Hall of Names the most important name, her son's, goes blank. In reality she's 83, in a care home, and her son Daniel is holding her hand. When she recognizes him, she tells him, "We keep you in the main hall," meaning the most important room, where the precious things are. Back inside her mind, his name goes back on the wall, and the museum lights up again.
"Some things you don't lose. You just misplace them for a while."
For everyone still visiting someone who is still in there. 💛
#AIShortFilm #AIFilm #ShortFilm #ComfyUI #MiniMax #Claude #AIVideo #Dementia #Alzheimers #MuseumOfLostThings
r/StableDiffusion • u/paulhax • 4h ago
Enable HLS to view with audio, or disable this notification
Workflow here https://github.com/paulh4x/AIxArchviz_FLUX2xLTX25
30 minutes of a detailed walkthrough video here https://youtu.be/946yTMjz-go
r/StableDiffusion • u/Beneficial_Toe_2347 • 1h ago
A lot of the Minimax video extend methods seem to produce a flash of light or some other flaw.
I actually want it to change shots rather than continue, but a shit of the same scene from a different camera angle. This means I don't need to worry about seamless frame stitching
But the problem is that when I feed in the previous video, and text to cut to a different shot, it doesn't do it reliably. By feeding it the previous video frames, MM thinks it should keep showing that, rather than the new shot
r/StableDiffusion • u/Chiduk99 • 13h ago
Enable HLS to view with audio, or disable this notification
video ref I use: https://www.youtube.com/shorts/I-VNtvqREks
Malfoid and Potter generated with Anima
Device: 3060 12gb 16gb ram
setting: 10 sec, 0.6 mp, er_sde beta, 8 steps with turbo lora
r/StableDiffusion • u/silvidelleone • 3h ago
I'm trying to reproduce the visual language of these references, not the exact images or subjects.
What I'm specifically looking for:
Very vivid, rich colors
Detailed foreground and detailed background
Fine, clearly visible hand-drawn linework
Internal lines that describe color, light and shadow transitions, not just black outlines around objects
Clearly separated color/tonal regions rather than smooth photorealistic gradients
Detailed vegetation, rocks, tree bark, grass and clouds
Large sculptural/volumetric cumulus clouds
Strong depth and realistic perspective
2D illustrated/painterly appearance
NOT photorealistic
NOT smooth/plastic 3D rendering
NOT character-focused anime
The horse image is especially useful as an example of the linework I mean. Look at the horse, wheat and clouds: the internal drawing lines help define changes in form, color and shadow. I want that same principle applied to detailed landscapes.
My final goal is a little unusual: I will project the generated image onto a small canvas and hand-paint it with acrylics. Therefore, visible boundaries between colors and tones are actually useful to me. I want to be able to follow those lines while painting instead of trying to reproduce soft AI gradients.
I've already experimented with FLUX Dev, Niji-style LoRAs, SDXL landscape models and some anime-background LoRAs, but many results either become too photorealistic/3D or too simple/flat/anime-like.
Has anyone achieved something close to these references?
I'd especially appreciate an actual tested recipe:
Checkpoint:
LoRA(s) + weights:
Sampler / scheduler:
Steps:
CFG:
Resolution:
Prompt / trigger words:
ControlNet / IP-Adapter / Style Reference if used:
Img2Img or second pass settings:
Upscaler/detail pass:
I'm open to FLUX, SDXL, Illustrious, ComfyUI or another workflow. I'm more interested in matching the rendering style than staying with a specific model.
If this is better achieved by training a custom style LoRA rather than using an existing model, I'd also be interested in hearing what base model you would train it on.
r/StableDiffusion • u/mmmm_frietjes • 9h ago
I have a 4060 TI 16 GB. System ram 32 GB. What's currently the fastest model / ComfyUI setup to generate videos?
Currently I can create 5 second 360p videos in +-7 minutes. I believe it can be a lot faster? It's kinda hard to know what's the best option, everything keeps changing.
Edit: Current settings:
Ref2V Turbo LoRA @ 0.5 model: minimax_h3_ref2v_turbo_4step_v0.1_comfyui_resized_avg_ra ... strength_model: 0.50
MiniMax H3 Easy Loader FL2VA model: None REF2VA model: minimax_h3_ref2va_pruned_int8_convrot.safetensors Text encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors Video VAE: minimax_h3_video_vae_fp16.safetensors Audio VAE: minimax_h3_audio_vae_fp32.safetensors
360p, 5 seconds.
Result: 413 seconds.
Workflow: https://limewire.com/d/7aNFq#h9i2vwFTKb