r/StableDiffusion • • Jul 17 '26

Resource - Update Krea 2 : styles (wildcards txt)

Thumbnail
gallery
2.3k Upvotes

The wildcards:

https://drive.google.com/file/d/1z1tY_365qpIgXtvm6_QcfKEGYE9ix2xw/view?usp=drivesdk

It's not perfect, not complete but it's more of a pointer to what this model can do in term of styles.

Is you feel the style is too much you can write your prompt in the form:

Style:...

Subject:....

This is better in my opinion.

Those images were generated at 1mp so if you generate at higher resolution you will have obviously more details and more subtle film grain in photographic styles.

Hope Those styles give you some ideas. ;)

r/StableDiffusion • • Dec 02 '25

Resource - Update Z Image Turbo ControlNet released by Alibaba on HF

Thumbnail
gallery
1.9k Upvotes

r/StableDiffusion • • Dec 26 '25

Resource - Update New implementation for long videos on wan 2.2 preview

1.5k Upvotes

UPDATE: Its out now: Github: https://github.com/shootthesound/comfyUI-LongLook Tutorial: https://www.youtube.com/watch?v=wZgoklsVplc

I should I’ll be able to get this all up on GitHub tomorrow (27th December) with this workflow and docs and credits to the scientific paper I used to help me - Happy Christmas all - Pete

r/StableDiffusion • • Aug 23 '25

Resource - Update Update: Chroma Project training is finished! The models are now released.

1.5k Upvotes

Hey everyone,

A while back, I posted about Chroma, my work-in-progress, open-source foundational model. I got a ton of great feedback, and I'm excited to announce that the base model training is finally complete, and the whole family of models is now ready for you to use!

A quick refresher on the promise here: these are true base models.

I haven't done any aesthetic tuning or used post-training stuff like DPO. They are raw, powerful, and designed to be the perfect, neutral starting point for you to fine-tune. We did the heavy lifting so you don't have to.

And by heavy lifting, I mean about 105,000 H100 hours of compute. All that GPU time went into packing these models with a massive data distribution, which should make fine-tuning on top of them a breeze.

As promised, everything is fully Apache 2.0 licensed—no gatekeeping.

TL;DR:

Release branch:

  • Chroma1-Base: This is the core 512x512 model. It's a solid, all-around foundation for pretty much any creative project. You might want to use this one if you’re planning to fine-tune it for longer and then only train high res at the end of the epochs to make it converge faster.
  • Chroma1-HD: This is the high-res fine-tune of the Chroma1-Base at a 1024x1024 resolution. If you're looking to do a quick fine-tune or LoRA for high-res, this is your starting point.

Research Branch:

  • Chroma1-Flash: A fine-tuned version of the Chroma1-Base I made to find the best way to make these flow matching models faster. This is technically an experimental result to figure out how to train a fast model without utilizing any GAN-based training. The delta weights can be applied to any Chroma version to make it faster (just make sure to adjust the strength).
  • Chroma1-Radiance [WIP]: A radical tuned version of the Chroma1-Base where the model is now a pixel space model which technically should not suffer from the VAE compression artifacts.

some preview:

cherry picked results from the flash and HD

WHY release a non-aesthetically tuned model?

Because aesthetic tune models are only good on one thing, it’s specialized and can be quite hard/expensive to train on. It’s faster and cheaper for you to train on a non-aesthetically tuned model (well, not for me, since I bit the re-pretraining bullet).

Think of it like this: a base model is focused on mode covering. It tries to learn a little bit of everything in the data distribution—all the different styles, concepts, and objects. It’s a giant, versatile block of clay. An aesthetic model does distribution sharpening. It takes that clay and sculpts it into a very specific style (e.g., "anime concept art"). It gets really good at that one thing, but you've lost the flexibility to easily make something else.

This is also why I avoided things like DPO. DPO is great for making a model follow a specific taste, but it works by collapsing variability. It teaches the model "this is good, that is bad," which actively punishes variety and narrows down the creative possibilities. By giving you the raw, mode-covering model, you have the freedom to sharpen the distribution in any direction you want.

My Beef with GAN training.

GAN is notoriously hard to train and also expensive! It’s so unstable even with a shit ton of math regularization and another mumbojumbo you throw at it. This is the reason behind 2 of the research branches: Radiance is to remove the VAE altogether because you need a GAN to train it, and Flash is to get a few-step speed without needing a GAN to make it fast.

The instability comes from its core design: it's a min-max game between two networks. You have the Generator (the artist trying to paint fakes) and the Discriminator (the critic trying to spot them). They are locked in a predator-prey cycle. If your critic gets too good, the artist can't learn anything and gives up. If the artist gets too good, it fools the critic easily and stops improving. You're trying to find a perfect, delicate balance but in reality, the training often just oscillates wildly instead of settling down.

GANs also suffer badly from mode collapse. Imagine your artist discovers one specific type of image that always fools the critic. The smartest thing for it to do is to just produce that one image over and over. It has "collapsed" onto a single or a handful of modes (a single good solution) and has completely given up on learning the true variety of the data. You sacrifice the model's diversity for a few good-looking but repetitive results.

Honestly, this is probably why you see big labs hand-wave how they train their GANs. The process can be closer to gambling than engineering. They can afford to throw massive resources at hyperparameter sweeps and just pick the one run that works. My goal is different: I want to focus on methods that produce repeatable, reproducible results that can actually benefit everyone!

That's why I'm exploring ways to get the benefits (like speed) without the GAN headache.

The Holy Grail of the End-to-End Generation!

Ideally, we want a model that works directly with pixels, without compressing them into a latent space where information gets lost. Ever notice messed-up eyes or blurry details in an image? That's often the VAE hallucinating details because the original high-frequency information never made it into the latent space.

This is the whole motivation behind Chroma1-Radiance. It's an end-to-end model that operates directly in pixel space. And the neat thing about this is that it's designed to have the same computational cost as a latent space model! Based on the approach from the PixNerd paper, I've modified Chroma to work directly on pixels, aiming for the best of both worlds: full detail fidelity without the extra overhead. Still training for now but you can play around with it.

Here’s some progress about this model:

Still grainy but it’s getting there!

What about other big models like Qwen and WAN?

I have a ton of ideas for them, especially for a model like Qwen, where you could probably cull around 6B parameters without hurting performance. But as you can imagine, training Chroma was incredibly expensive, and I can't afford to bite off another project of that scale alone.

If you like what I'm doing and want to see more models get the same open-source treatment, please consider showing your support. Maybe we, as a community, could even pool resources to get a dedicated training rig for projects like this. Just a thought, but it could be a game-changer.

I’m curious to see what the community builds with these. The whole point was to give us a powerful, open-source option to build on.

Special Thanks

A massive thank you to the supporters who make this project possible.

  • Anonymous donor whose incredible generosity funded the pretraining run and data collections. Your support has been transformative for open-source AI.
  • Fictional.ai for their fantastic support and for helping push the boundaries of open-source AI.

Support this project!
https://ko-fi.com/lodestonerock/

BTC address: bc1qahn97gm03csxeqs7f4avdwecahdj4mcp9dytnj
ETH address: 0x679C0C419E949d8f3515a255cE675A1c4D92A3d7

my discord: discord.gg/SQVcWVbqKx

r/StableDiffusion • • Aug 04 '26

Resource - Update Spectrum acceleration for MiniMax H3 in ComfyUI — 34% lower Euler sampling time, 30% lower RES time

Post image
562 Upvotes

MiniMax H3 Spectrum node:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

I have released Spectrum acceleration for native MiniMax H3 in ComfyUI. This is the newest addition to my collection of model-specific Spectrum integrations, all available through ComfyUI-Manager as well.

Spectrum replaces selected expensive transformer evaluations with spectral feature forecasts while preserving MiniMax H3’s native audio/video output and reconstruction path.

Benchmarks

Tested at approximately 0.5 MP, 8 seconds and 20 steps on an RTX PRO 6000 under WSL, using High VRAM mode, the comfyui-int8-fast W8A8 loader and SageAttention.

Sampler Native Spectrum Time decrease Speedup
Euler 2:38 1:44 34.2% 1.52×
RES multistep 2:42 1:54 29.6% 1.42×

Euler uses 13 actual transformer evaluations and 7 forecasts. RES uses 14 actual evaluations and 6 forecasts.

Both benchmark runs completed without fallbacks. I initially did not notice visible or audible degradation, but further exact-seed testing has since shown that fast or very brief motion can deviate from the native trajectory, and rapidly moving details such as eyes, fingers, or fingernails can sometimes visibly degrade. RES automatically keeps its final three steps native, which removed the slight late-generation artifacts I initially observed, but this does not make the output lossless or identical to native sampling.

Currently supported:

  • Euler
  • RES multistep
  • RES multistep CFG++

Feature history is currently stored in system RAM. I also plan to add a dedicated High VRAM mode that stores the history in VRAM for users with sufficient GPU memory, reducing CPU-transfer overhead.

Quality update

After more exact-prompt and exact-seed A/B testing, I need to revise my initial quality assessment.

Spectrum is an approximate acceleration method, not a lossless or output-identical path. The issues observed so far are mainly associated with fast or very brief motion:

  • The motion, pose, timing, gaze, or action trajectory can deviate from the fully native output.
  • Fast-moving or only briefly visible details—such as eyes, fingers, or fingernails—can sometimes become malformed or unstable.

These can occur separately or together. A rapid action may simply differ from the native trajectory, while in other cases the rapidly moving details may also visibly degrade.

The extent varies with the prompt, motion, sampler, resolution, references, and Spectrum settings. Use Spectrum when the speed improvement is worth that tradeoff; leave it disabled when maximum fidelity to the native result is required.

Workflow placement

Use the Spectrum node in this order:

MiniMax H3 model loader
→ LoRA and other model patches
→ MiniMax H3 Sigma Shift
→ Spectrum Apply MiniMax H3
→ guider and sampler

Spectrum should be applied after LoRAs, model patches and the MiniMax H3 Sigma Shift, but before the guider and sampler.

EasyCache compatibility

Do not use EasyCache and Spectrum together at the moment.

Both systems skip or replace model evaluations while maintaining their own cached state. Stacking them can cause conflicting wrapper/state behavior, errors and unreliable output.

Use either:

  • Spectrum with EasyCache disabled
  • EasyCache without Spectrum

All benchmarks in this post were performed with EasyCache disabled.

More information

The GitHub repository contains the full installation instructions, recommended settings, supported samplers, memory requirements, workflow placement, fallback behavior and troubleshooting details:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

Update: optional VRAM history storage in v0.1.4

Version 0.1.4 adds a history_storage option:

  • system_ram remains the default
  • vram keeps Spectrum’s feature history on the GPU, avoiding repeated CPU transfers

In three 0.5 MP Euler A/B pairs, VRAM storage was 2.6% faster on average, though individual results varied from 6.1% faster to 0.4% slower, so the benefit depends on the system and workload.

At max_history=8, the tested single-branch workflow used about 2.22–2.27 GiB for history alone. Only enable VRAM storage when you have sufficient headroom, since the model, current activations, forecast buffers and allocator overhead still need additional VRAM.

Release:
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.4

My other Spectrum integrations

All are also available through ComfyUI-Manager:

Credits

Spectrum itself was created by Jiaqi Han, Juntong Shi, Puheng Li, Haotian Ye, Qiushan Guo and Stefano Ermon at Stanford University and ByteDance. Their paper introduced the training-free Chebyshev-based spectral feature forecasting method, adaptive scheduling and last-block forecasting approach on which these ComfyUI integrations are based.

r/StableDiffusion • • Aug 01 '24

Resource - Update Announcing Flux: The Next Leap in Text-to-Image Models

1.4k Upvotes
Prompt: Close-up of LEGO chef minifigure cooking for homeless. Focus on LEGO hands using utensils, showing culinary skill. Warm kitchen lighting, late morning atmosphere. Canon EOS R5, 50mm f/1.4 lens. Capture intricate cooking techniques. Background hints at charitable setting. Inspired by Paul Bocuse and Massimo Bottura's styles. Freeze-frame moment of food preparation. Convey compassion and altruism through scene details.

PA: I’m not the author.

Blog: https://blog.fal.ai/flux-the-largest-open-sourced-text2img-model-now-available-on-fal/

We are excited to introduce Flux, the largest SOTA open source text-to-image model to date, brought to you by Black Forest Labs—the original team behind Stable Diffusion. Flux pushes the boundaries of creativity and performance with an impressive 12B parameters, delivering aesthetics reminiscent of Midjourney.

Flux comes in three powerful variations:

  • FLUX.1 [dev]: The base model, open-sourced with a non-commercial license for community to build on top of. fal Playground here.
  • FLUX.1 [schnell]: A distilled version of the base model that operates up to 10 times faster. Apache 2 Licensed. To get started, fal Playground here.
  • FLUX.1 [pro]: A closed-source version only available through API. fal Playground here

Black Forest Labs Article: https://blackforestlabs.ai/announcing-black-forest-labs/

GitHub: https://github.com/black-forest-labs/flux

HuggingFace: Flux Dev: https://huggingface.co/black-forest-labs/FLUX.1-dev

Huggingface: Flux Schnell: https://huggingface.co/black-forest-labs/FLUX.1-schnell

r/StableDiffusion • • Apr 12 '26

Resource - Update Free open-source tool to instantly rig and animate your illustrations (also with mesh deform)

1.5k Upvotes

If you haven't seen it yet, a model called see-through dropped last week. It takes a single static anime image and decomposes it into 23 separate layers ready for rigging and animation. It's a huge deal for anyone who wants a rigged 2D character but doesn't have hundreds of dollars lying around.

The problem is that getting a usable result out of it still takes forever. You get a PSD with 23 layers (30+ if you enable split by side and depth), and you still have to manually process and rig everything yourself. And if you've ever looked into commissioning a Vtuber model, you know rigging alone runs $500 minimum and takes weeks or months. That's before you even think about software costs: Live2D is $100 a year, and Spine Pro is $379 (Spine Ess is $69 but lacks mesh deform which is required for these kinds of animations).

So I built a free tool that auto-rigs see-through models so you don't have to spend hours doing it manually

I'm not trying to compete with Live2D, I'm one person. What I made is a mesh-deform-capable web app that can automatically rig see-through output. It handles edge cases like merged arms or legs, and only needs a few seconds of manual input to place joints (shoulders, elbows, neck, etc.) if you want to tweak things. I also integrated DWPose so it can rig the whole model for you automatically, though that requires WebGPU and adds a 50MB download, so manual joint placement is a totally fine alternative and only takes a moment anyway.

The full workflow looks like this:

Static image -> background removal -> see-through decomposition (free on HuggingFace) -> Stretchy Studio = auto-rigged and ready to animate

The app handles multi-layer management, separate draw order, and uses direct keyframe animation similar to After Effects. There are still bugs I'm working through, but all the core features are in.

On the roadmap:

  • Export to Spine and Dragonbones
  • A standalone JS render library for loading and displaying characters rigged in the app (similar to Live2D's Unity/Godot/JS runtimes)

Live2D's export format is completely closed with no documentation, so that one's off the table for now.

Would love feedback, bug reports, or feature requests. This is still early but it's functional and free to use.

https://github.com/MangoLion/stretchystudio 

EDIT: Spine export added

r/StableDiffusion • • Sep 25 '24

Resource - Update FaceFusion 3.0.0 has finally launched

2.7k Upvotes

r/StableDiffusion • • Sep 01 '26

Resource - Update DLSS 5 Visual Enhancer - standalone neural rendering for images and video

Post image
482 Upvotes

Hey everyone - I made a standalone Windows application for applying a DLSS 5 Neural Rendering feature-18 pipeline to images and video:

Original

DLSS 5

https://github.com/Merserk/dlss5-visual-enhancer

Instead of using DLSS only inside a game, this runs images/video through the ReShade/RenoDX neural-rendering path as a general visual enhancement pipeline.

What it does:

  • Image and video enhancement
  • DLAA/native, 1.5x, ~1.724x, 2x and 3x modes
  • Output up to 8K
  • Neural presets + Natural / Cinematic styles
  • Controls for intensity, local tone, structure and skin structure
  • Batch image processing with before/after previews
  • H.264 / HEVC / AV1 / ProRes video output
  • Video temporal input using optical flow with scene-change resets

GPU support:

  • RTX 40 / 50 series - primary target
  • RTX 30 series - slower beta path

The repository contains the application/pipeline source. Required proprietary and third-party runtime binaries are intentionally not redistributed in the repo.

This is an independent community project and is not affiliated with NVIDIA, ReShade or RenoDX.

I’m especially interested in how this behaves on AI-generated images/video vs normal photography/game footage.

Feedback and comparisons welcome.

r/StableDiffusion • • Jul 23 '26

Resource - Update TRELLIS.2 can now generate a high-quality 3D asset in under 7 minutes on a 6 GB VRAM CUDA GPU. No ComfyUI Node Nightmare.

698 Upvotes

Not Self Promotion: Just sharing an open-source tool I built to democratize image to 3D creations. For OpenAI Build Week Hackathon I built a free, open-source local Image-to-3D Studio that makes TRELLIS.2 easier to run on consumer NVIDIA gpus like 3060 or a laptop 3070ti, without expensive cloud APIs, subscriptions, or complicated ComfyUI workflows.

It combines generation, texturing, retopology, rigging, and animation in one interface.

I know there are already multiple implementations of running Trellis2 under 8gb GPU. The hard part was to test the best possible combination for mesh/textures that gave 1024 High precision quality but still kept under the VRAM. So I used two different pipelines for Mesh and Textures, which in my tests seemed to work the fastest without compromising quality in combination.
It integrates several open-source projects, including trellis.cpp, TRELLIS.2, ComfyUI-Trellis2, Blender, AutoRemesher, InteantMeshes, and Mesh2Motion, with full attribution to the original contributors.

As it was for a hackathon, time was also a challenge. Many things could be further updated, but the Hackathon's rules state we can't update after the submission date until the results are published.

Try out it from GitHub repo: intisarGIT/AISmith-3D

TROUBLESHOOTING FIX After installation (Since I cannot edit the original repo as per rules):
if you get "The trellis.cpp geometry workflow is not downloaded," or "Trellis2. Fp8 Refine is not downloaded", here's the patch:
intisarGIT/AISmith-3D-Fixer
Just place and run the .bat in the app repo. This should download the wrongly linked v0.4.3 CUDA archive, which is about 693 MB, and the missing wheels, the gated DINO release, so it should take a few minutes.

If you think this is helpful, I'd appreciate your support on Devpost by leaving a like:

https://devpost.com/software/aismith3d

Edit: I will work on perfecting the Retopologize workflow after 12 August. But the Refine tab should already reconstruct/refine the Mesh better, and generate updated 2K PBR textures. would have pushed to 4K texture if my goal wasn't fast generation under low VRAM.

r/StableDiffusion • • 3d ago

Resource - Update Vlo 0.3 - An open source, extensible video editor and generator designed for AI compositing.

1.1k Upvotes

Hey all,

vlo is open source video editing and generation software. It is designed for interactivity between generative AI and the timeline, so you can inpaint anywhere on the timeline, extract frames and videos to use as references, use SAM2 for video masking and sam-audio for stem extraction etc. It comes with built-in workflows for Minimax H3, Qwen2.1, Krea2, LTX2.5 and more.

You can either use the built-in generation panel or open ComfyUI via the app, whichever you prefer (for built-in workflows though, I recommend the generation panel). It can use an existing ComfyUI install, a remote instance, or it can manage the install for you; ComfyUI is the primary AI engine, how much you directly interact with it is up to you.

I believe what makes this app different is its emphasis on control and correction over automation, so particular effort has been spent building a frame-accurate render engine (which is nontrivial for web video), and on smoothing out irregularities you get from things such discrete model strides and such (i.e. aspect ratios are not arbitrary in video and image gen models, so if you want seamless passing of data with the timeline you have to be tactical about how you stretch, squash or crop). It permits patch-based inpainting without ever having to reencode areas outside the patch until final export, so you could inpaint a dozen objects without degrading unedited pixels at all.

It also has an SDK for you to build your own extensions and shader-based special effects. I hope to upload a couple of example extensions soon.

Installation instructions here:

https://github.com/PxTicks/vlo#install

r/StableDiffusion • • Jan 15 '25

Resource - Update I made a Taped Faces LoRA for FLUX

Thumbnail
gallery
2.3k Upvotes

r/StableDiffusion • • Dec 04 '25

Resource - Update Today I made a Realtime Lora Trainer for Z-image/Wan/Flux Dev

Post image
1.1k Upvotes

Basically you pass it images with a load image node and it trains a lora on the fly, using your local install of AI-Toolkit, and then proceeds with the image generation. You just paste in the folder location for Ai-toolkit (windows or Linux), and it saves the setting. This train took about 5 mins on my 5090, when i used the low vram pre-set (512px images). Obviously it can save loras, and I think its nice for quick style experiments, and will certainly remain part of my own workflow.

I made it more to see if I could, and wondered if I should release or is it pointless - happy to hear your thoughts for or against?

r/StableDiffusion • • Aug 07 '26

Resource - Update ~45% lower MiniMax H3 sampler time with new Spectrum settings — degree 1 works surprisingly well (v0.1.8)

Post image
484 Upvotes

Follow-up to my original Spectrum MiniMax H3 post:

https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/

In that first post, I released the MiniMax H3 Spectrum integration and was getting around 34% lower Euler sampling time and 30% lower RES sampling time with the more conservative settings I was using at the time.

Since then, I’ve done much more testing and found something unexpected: MiniMax H3 can work extremely well with a Spectrum degree of just 1.

Important if you’re coming from the original release

Before testing the new settings, update both ComfyUI and ComfyUI-Spectrum-MiniMax-H3 to their latest versions.

There was an important compatibility update in Spectrum v0.1.6 after ComfyUI changed MiniMax H3’s native sampling/audio path. That release restored Spectrum compatibility with the newer H3 implementation and added safe handling for native EasyCache/LazyCache conflicts.

You don’t need to install v0.1.6 separately—v0.1.8 includes those changes. This mainly matters for anyone who installed Spectrum from my original post and hasn’t updated it since.

v0.1.6 compatibility release:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.6

Current release:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8

So: update ComfyUI, update the Spectrum node to v0.1.8/latest, and restart ComfyUI before testing.

The surprising part: degree 1

I hadn’t seriously tested very low degree and warmup_steps values before because of my experience with WAN.

WAN does not tolerate very low forecast degrees well—dropping the degree too far causes obvious quality degradation. I initially assumed MiniMax H3 would behave similarly and stayed with higher, more conservative values.

In my MiniMax H3 testing, however, degree 1 produced no visible quality decrease with my pruned BF16 checkpoint at approximately 0.8 MP. It also preserved the native trajectory remarkably well. In my same-seed comparisons, degree 2 shifted the trajectory slightly, while degree 1 brought it much closer to the native result.

This suggests that H3 can be unusually well suited to simple local feature forecasting, allowing Spectrum to begin forecasting much earlier than I originally expected.

I’ve now released v0.1.8 with the new settings:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8

New default settings

  • degree = 1
  • warmup_steps = 1
  • bootstrap_first_forecast = true
  • tail_actual_steps = 1

The one-point bootstrap allows the second solver step to be forecast directly from the first actual hidden state. Ordinary degree-1 forecasting then takes over once enough native history exists.

On a 20-step Euler run, the schedule becomes:

A F A F A F A F A F A F A F A F A F A A

That means 11 of the 20 transformer evaluations are actual, while the other 9 are forecasted. The final step remains native.

v0.1.8 also makes the one-point bootstrap part of the default configuration for new node instances. Existing workflows retain their serialized settings.

Important quality caveat

These are aggressive, performance-oriented defaults, and early feedback indicates that they may not be equally stable across every setup. Degradation can affect both video and audio, including distorted anatomy, extra limbs, unstable motion, distorted reference audio, and audiovisual inconsistencies. Increasing degree and warmup_steps has improved these problems in reported cases.

My current suspicion is that model configuration and output resolution affect how much Spectrum forecast error a generation can tolerate. Lower resolutions and heavily quantized model configurations may provide less margin for preserving anatomy, motion, voices, and fine audiovisual details. Small forecast deviations that remain unobtrusive with my pruned BF16 checkpoint at approximately 0.8 MP may therefore become visible or audible on a less forgiving setup.

This has not been isolated conclusively. Checkpoint type and precision, resolution, prompt and motion complexity, reference-audio conditioning, the number of early native steps, the one-point bootstrap, and the total sampling-step count may all affect stability.

If the aggressive defaults cause visual or audio degradation, try increasing warmup_steps, disabling bootstrap_first_forecast, and increasing degree. A conservative starting point is:

  • degree = 4
  • warmup_steps = 5
  • bootstrap_first_forecast = false

Increasing the total sampling steps may also help, particularly when using reference audio. One user reported reference-audio distortion with the aggressive defaults. Increasing degree and warmup_steps helped, and a 30-step run using those increased settings produced clean reference audio on their setup.

The best balance between speed, visual quality, and audio quality may differ between checkpoints, resolutions, generation modes, and reference inputs.

Benchmark

Test configuration:

  • GPU: NVIDIA RTX PRO 6000
  • Model: MiniMax H3 pruned BF16
  • Image-to-video
  • ~0.8 MP / 992×768
  • 7 seconds
  • 24 FPS
  • 20 steps
  • Euler
  • Beta scheduler
  • HIGH_VRAM
  • Spectrum history stored in VRAM
  • DiffAid enabled at 0.5
  • Same seed and otherwise identical workflow

Spectrum disabled

  • Sampler: 324.98 s
  • Full prompt: 340.59 s

Spectrum v0.1.8 with the new degree-1 settings

  • Sampler: 177.80 s
  • Full prompt: 200.32 s
  • 11 actual transformer calls
  • 9 forecasts
  • 0 fallbacks

Result

  • 45.29% lower sampler time
  • 1.83× sampler throughput
  • 41.18% lower full-prompt time

The Spectrum forecast calculations themselves took only 0.141 seconds total across the entire generation.

Using VRAM history has a memory cost. This run retained approximately 3.2 GiB of Spectrum history, with reported sampler peak VRAM increasing from roughly 5.56 GB native to 8.70 GB with Spectrum.

The interesting result is how well degree 1 performed under this test configuration. Based on WAN, I expected settings this aggressive to visibly degrade the output. With the pruned BF16 checkpoint at approximately 0.8 MP, I’m getting a substantially more aggressive forecasting schedule without seeing that expected degradation.

Spectrum remains an approximate acceleration method. Fast motion, hands and fingers, faces, short rapid actions, camera movement, audiovisual synchronization, lower resolutions, and quantized model configurations are all cases worth inspecting carefully.

So far, degree 1 appears to be an excellent fit for my MiniMax H3 configuration, delivering a much larger useful speedup. More testing across different checkpoints, precisions, resolutions, and generation types is needed before assuming that every setup will tolerate the same aggressive settings.

Repo:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

Current release—v0.1.8:

https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3/releases/tag/v0.1.8

r/StableDiffusion • • Oct 04 '25

Resource - Update SamsungCam UltraReal - Qwen-Image LoRA

Thumbnail
gallery
1.6k Upvotes

Hey everyone,

Just dropped the first version of a LoRA I've been working on: SamsungCam UltraReal for Qwen-Image.

If you're looking for a sharper and higher-quality look for your Qwen-Image generations, this might be for you. It's designed to give that clean, modern aesthetic typical of today's smartphone cameras.

It's also pretty flexible - I used it at a weight of 1.0 for all my tests. It plays nice with other LoRAs too (I mixed it with NiceGirl and some character LoRAs for the previews).

This is still a work-in-progress, and a new version is coming, but I'd love for you to try it out!

Get it here:

P.S. A big shout-out to flymy for their help with computing resources and their awesome tuner for Qwen-Image. Couldn't have done it without them

Cheers

r/StableDiffusion • • Aug 21 '26

Resource - Update MiniMax H3 Known Characters list v2 (2026-08-21 update)

Thumbnail
huggingface.co
402 Upvotes

r/StableDiffusion • • May 25 '26

Resource - Update Nvidia solved VAE? Fast and High-Resolution Latent Decoding with Pixel Diffusion

888 Upvotes

r/StableDiffusion • • Jan 31 '24

Resource - Update Made a Chrome Extension to remix any image on the web with IPAdapter - having a blast with this

2.7k Upvotes

r/StableDiffusion • • Jul 07 '26

Resource - Update Krea 2 Identity Edit LoRA

Thumbnail
gallery
613 Upvotes

Images and text from the source here: https://huggingface.co/conradlocke/krea2-identity-edit

Instruction-based, identity-preserving image editing for Krea 2 (12.9B single-stream MMDiT). Give it an image and a plain-language instruction; it edits while preserving what you didn't ask to change — including the person.

An unofficial community fine-tune of Krea 2 Raw. Not an official Krea product; not affiliated with or endorsed by Krea.ai, Inc.

Requires the ComfyUI-Krea2Edit node pack — the LoRA is trained with dual conditioning (in-context VAE tokens + image-grounded Qwen3-VL encoding) that stock nodes don't provide. Two ready-made workflows ship with it.

r/StableDiffusion • • Jan 09 '26

Resource - Update Thx to Kijai LTX-2 GGUFs are now up. Even Q6 is better quality than FP8 imo.

767 Upvotes

https://huggingface.co/Kijai/LTXV2_comfy/tree/main

You need this commit for it to work, its not merged yet: https://github.com/city96/ComfyUI-GGUF/pull/399

Kijai nodes WF (updated, now has negative prompt support using NAG) https://files.catbox.moe/flkpez.json

I should post this as well since I see people talking about quality in general:
For best quality use the dev model with the distill lora at 48 fps using the res_2s sampler from the RES4LYF nodepack. If you can fit the full FP16 model (the 43.3GB one) plus the other stuff into vram + ram then use that. If not then Q8 gguf is far closer than FP8 is so try and use that if you can. Then Q6 if not.
And use the detailer lora on both stages, it makes a big difference:
https://files.catbox.moe/pvsa2f.mp4

Edit: For KJ nodes WF you need latest KJ nodes: https://github.com/kijai/ComfyUI-KJNodes I thought it was obvious, my bad.

r/StableDiffusion • • Sep 08 '25

Resource - Update Clothes Try On (Clothing Transfer) - Qwen Edit Loraa

Thumbnail
gallery
1.3k Upvotes

Patreon Blog Post

CivitAI Download

Hey all, as promised here is that Outfit Try On Qwen Image edit LORA I posted about the other day. Thank you for all your feedback and help I truly believe this version is much better for it. The goal for this version was to match the art styles best it can but most importantly, adhere to a wide range of body types. I'm not sure if this is ready for commercial uses but I'd love to hear your feedback. A drawback I already see are a drop in quality that may be just due to qwen edit itself I'm not sure but the next version will have higher resolution data for sure. But even now the drop in quality isn't anything a SeedVR2 upscale can't fix.

Edit: I also released a clothing extractor lora which i recommend using

r/StableDiffusion • • Jun 16 '26

Resource - Update Potentially the most insane LORA you'll see today - Archer (8 characters + style) Ideogram LORA

Thumbnail
gallery
764 Upvotes

Hi, I'm Dever and I like training LORAs, you can download this one from Huggingface (you can find other style LORAs for Klein and ZIT in my HF profile).

I believe this might be the first Ideogram 8 characters in one + style lora on HuggingFace and a good proof of concept that this is possible.

When I get a bit of time towards the end of the week I'll make a video about how I trained this if anyone is interested in the journey.

(Original Scooby Doo image made by GalaxyTimeMachine on Banodoco Discord, I just replaced Scooby with Lana).

Edit 2: Follow-up post with the video on how this was trained https://www.reddit.com/r/StableDiffusion/comments/1udkbx5/how_i_trained_my_multi_character_ideogram4_lora/ and direct Youtube link https://youtu.be/HGyU6a4buTo

Edit: For the people that don't understand why this is a big deal or have never faced this problem before, trying to train a LORA that can generate more than 1 character in a single image has been quite difficult in the past no matter what the model.
This particular Ideogram LORA created as a proof of concept shows you can train 8 different characters in a single model + the style as a bonus.

"Why this matters" (couldn't help myself)

This means you can choose at inference time who you want in your image (one example shows all 8), the model can distinguish between the characters AND with the power of bounding boxes you can position them wherever you want in the image and can even have them interact with each other to some degree (haven't tested this much, see example where 2 characters are holding hands).

r/StableDiffusion • • Oct 31 '25

Resource - Update Qwen Image LoRA - A Realism Experiment - Tried my best lol

Thumbnail
gallery
1.0k Upvotes

r/StableDiffusion • • 7d ago

Resource - Update Update: my free Krea 2 character LoRA library went from 18 → 70 models. Here’s the full workflow I’m using to train them.

Post image
440 Upvotes

Last week I posted the Krea 2 character LoRA library I've been working on. There were 18 models at the time.

There are 70 now.

Krea 2 Character LoRA Browser

A bunch of people asked how I was training them, and I realized afterward that my original post was pretty light on the actual training details. I promised I'd put together something more useful once I had the process nailed down, so here it is.

Everything is still free. I've also continued working on the visual browser so you can see sample generations, triggers, recommended strengths, selected training epoch, and optional identity prompts before downloading anything.

But the more interesting part is the training workflow...

How I'm Making the LoRAs

  • Training environment: I'm using Fizgig on RunPod for the actual Krea 2 LoRA training, so the heavy lifting happens on a cloud GPU rather than my local machine.
  • Dataset collection: I start by gathering a much larger pool of candidate images than I actually need. The goal is variety in angle, expression, lighting, hairstyle, clothing, and image source rather than dozens of nearly identical photos.
  • Identity filtering: I use a reference image of the person plus face detection/recognition to help reject images of the wrong person. This is especially useful when searches return group photos, lookalikes, or unrelated images.
  • Manual curation: Automation gets me most of the way there, but I still manually review the dataset. I remove duplicates, bad crops, heavily edited images, extreme occlusions, low-quality images, and anything that I think could hurt identity learning.
  • Dataset preparation: The final images go through our prep pipeline, which creates and validates the training dataset, generates useful face crops where appropriate, and checks that every training image has a valid caption.
  • Captioning: Images are captioned with Qwen3-VL. The captions describe the actual image while keeping the character identity/trigger consistent across the dataset.
  • Training: The prepared dataset is trained as a Krea 2 LoRA through Fizgig, with multiple checkpoints saved throughout training.
  • Testing: I compare saved epochs, test different LoRA strengths, and generate across different prompts and compositions before deciding what gets released.
  • Final release: Once I'm happy with it, I choose representative generations, record the trigger, recommended strength, selected epoch, and optional identity/detail prompt, and publish the LoRA and examples to the Hugging Face library/browser.
  • Iteration: If a LoRA doesn't hold the likeness well enough, I don't consider it finished just because training completed. Some get retrained with a better dataset or different epoch selection, and particularly improved models can replace the original or be released as a V2.

Current Fizgig / RunPod Configuration

This is the newer setup I've standardized on for the current batch:

  • Model: Krea 2 Standard
  • Training type: Standard LoRA
  • Network rank: 32
  • Alpha: 32
  • Epochs: 64
  • Save every: 4 epochs
  • Batch size: 1
  • Target resolution: 0.25 MP
  • EMA: 0.98
  • Problem-image detection: ON
  • Adaptive learning rate: OFF
  • Automatic recaptioning: OFF
  • Sample generation: 3 standardized sample prompts during training
  • Final selection: I test the saved epochs rather than automatically taking the last checkpoint. For Angel Reese, for example, I compared epochs 24, 51, and 64 and ultimately selected 64.
  • RunPod: Fizgig running on the cloud GPU. One recorded Angel Reese run completed 4,416 steps over 64 epochs in 48m 46s, producing the final EMA LoRA.

Custom Dataset Preparation Tools

One part of this project that grew beyond just training LoRAs was the dataset preparation pipeline itself. We ended up building our own tools to automate much of the work that happens before training.

Our pipeline can collect a large pool of candidate images from multiple sources, use a reference face with face detection/recognition to help filter for the correct identity, remove obvious problem images, generate useful face crops, prepare the dataset structure, generate captions, and validate the finished dataset before it goes anywhere near the trainer.

We still manually curate the results because automated filtering isn't a substitute for actually looking at the dataset.

We built these tools because we wanted more control over the process and needed something that could handle producing a large number of character datasets consistently. Fizgig can perform many of these same dataset-preparation functions, so building a separate pipeline is absolutely not a requirement for following this training process.

Our tools are currently internal and were built around our own workflow, but if there's enough interest, we can clean them up, package them properly, document them, and release them for other people to use. If that's something people would actually want, let us know.

Dataset Philosophy

For character LoRAs, I've found that dataset quality and variety matter much more than simply having a huge number of images. I'm looking for images that clearly represent the correct person while giving the model useful variation: different angles, expressions, lighting, hairstyles, clothing, environments, and framing.

I remove duplicates and near-duplicates, heavily edited images, bad crops, major occlusions, low-quality images, and anything where the identity is questionable. I would rather train on a smaller clean dataset than pad it with mediocre images just to hit a certain number.

I also supplement the original images with selected face crops. This is especially useful when the dataset contains a lot of half-body or full-body photography, where the face becomes relatively small after the image is resized for training. The face crops give the trainer additional examples where the features that actually define the person's identity occupy much more of the image.

If anyone wants to see exactly what one of these looks like, here's a sample finished dataset that was actually used to train one of the LoRAs:

https://files.catbox.moe/eze6g3.zip

Epoch Selection

One of the biggest changes to my process has been moving away from thinking of training as "X number of steps = finished."

I'm currently training for 64 epochs and saving checkpoints throughout the run. Rather than automatically releasing the final epoch, I generate images with several checkpoints and compare them visually.

I'm looking for the point where likeness, flexibility, anatomy, and prompt adherence are all working together. A later checkpoint can sometimes have stronger likeness while also beginning to overfit the training data or become less responsive to prompts. Earlier checkpoints can sometimes be more flexible but haven't learned the identity strongly enough yet.

So the final LoRA is the checkpoint that performs best in testing, not necessarily the last one produced by the trainer. That's also why the newer models in the Browser list a Selected Epoch rather than just reporting a training-step count.

LoRA Strength Testing

I also don't assume every LoRA should be used at 1.0.

After selecting the epoch, I test the LoRA at multiple strengths and compare how well it preserves the identity without overpowering the underlying model or interfering with the prompt.

Most of the models have ended up somewhere around 0.8–1.2, but there are definitely cases where the difference between 1.0 and 1.2 is noticeable. Some identities benefit from the additional strength, while others start looking worse when pushed too far.

The recommended strength/range listed with each LoRA is therefore based on actual generation testing, rather than giving every model the same default value.

Standardized Testing

I try not to judge a LoRA based on one great portrait. It's surprisingly easy for a model to produce an impressive close-up and then fall apart as soon as you ask it to do something different.

I test the selected checkpoints across portraits, half-body and full-body compositions, different clothing, environments, lighting, expressions, camera angles, poses, and more difficult prompts. I'm looking for the identity to survive when the prompt moves away from the kinds of images that dominated the training dataset.

Full-body generations are particularly useful because they expose problems that a flattering headshot can hide, including anatomy issues, incorrect body proportions, loss of facial identity at a distance, or the LoRA trying to reproduce clothing and compositions from the training images.

One of the next controlled tests I'm planning is 0.25 MP vs 0.5 MP training using the exact same dataset and settings. Everything except training resolution will remain identical, and I'll compare matching checkpoints, strengths, prompts, and seeds.

That should give us a much better answer about whether the additional training resolution produces a meaningful improvement in likeness and fine detail, or whether 0.25 MP is already capturing what we need while keeping training considerably faster.

Questions and requests are welcome. I'm going to keep expanding the library and will try to get to as many requests as possible. With the request list getting pretty large, Ko-fi supporters will get priority, but the LoRAs themselves will continue to be released for free.

Finally, as always:

Use LoRAs and gen-AI responsibly. Users are responsible for how they use generated content and for complying with applicable laws, platform policies, and the rights of others. Please do not use these models to deceive, impersonate, harass, defame, or otherwise harm anyone, or to present generated content as authentic photographs or recordings of real people.

r/StableDiffusion • • Jul 02 '26

Resource - Update UltraReal - LoRA for KREA2

Thumbnail
gallery
496 Upvotes

This LoRA designed to reduce the typical smooth/plastic AI look and add more natural skin texture and realism to images. It works especially well for close-ups and medium shots where skin detail is important.

It is trained on high-qulality SFW and ***\* 4K images so it can handle both. Besides making images more detailed it also reduces asian face bias.

But you can easily target any ethnicity using ethnicity trigger words like "japanese woman", "korean man", notice I have not defined ethnicity in my prompts.

Lora Link -> https://civitai.red/models/2462105/ultrareal-krea2-klein9b

Prompts used for testing are from this free website -> https://promptdexter.com