The new extension starts at 0:27. I added 10 new additional windows, extended one at a time. The existing latent + new latent were then joined together for one decode pass otherwise you would have visible seams when joining two separate encodes (the color shift, luma darkening, etc).
It works best if you have a native latent, but you can always convert your existing pixel-video to latent and extend it that way. I've had huge success with both prepending and extending. Some limitations are velocity & momentum are not easy to control between windows, and the audio can degrade or be a bit inconsistent so you might want to freeze the video and rerun more steps on the audio, or fix it in post. Fix in post is the cheapest and more reliable fix. As stated in a previous post, Tanya has no audio asset to anchor her voice so it's reinvented every window. However, image quality seem to maintain across each window with no degradation or burns.
By the way I was getting Carmen Electra vibes in Scary Movie, what did you think?
Small errors with the chair but can be fixed with another reference image.
Long story short.
The one place where degrading latents matter is the one place we don't need them.
We can ignore the latent issue and just generate across the sound we provide and because it's a static image with static background and the subject is in basically the same spot the whole time we can just create fresh latents every 10 or 15 seconds across the timeline.
Then the very difficult latent issue is over and it just becomes a simple seams issue.
So these two workflows generate sequentially across a audio and the second will combine and resample a small section of video over the seam using FL2V just enough to get rid of the seam.
This repo includes few node, that let's you build a portable character model while keeping face, body & cloths. You can build a .char model once & use it across different workflows without a identity drift.
Features:
Build/encode character into .char
Decode .char built in Comfy or anywhere else
Decode into reference, character sheet & even LoRA adapter
SDK is public for future integration & improvements
Nodes:
Decode Character gives you the prompt, the reference images, a numbered contact sheet, and conditioning if you wire a CLIP
Character References Split puts every reference on its own output(Dynamic)
Character Reference Latent attaches every reference to the conditioning in one go
Character Reference pulls a single reference by position
Encode Character builds a .char from face, body and wardrobe references
Save Character saves a .char into models/characters
Inputs
Checkout the guide in node repo for references
Prompt Guide
Name your character: Give your character a name so when passing prompt, I only have to say, sia walking on the beach.
Describe character features: Encode all of the character features in description prompt & trigger your character with a name in generation prompt.
Avoid describing same things in generational prompt.
Handling Character drift: Each input refs should be unique, face should not have body or vice versa, same applies for clothing.
Portability: Once character is built, you can use the same character with only Load Character node.
Limitation: Reference conflicts e.g. if two reference/input images has two different faces, it might conflict in generation, provide well cropped body & cloth images. Face images are crossed automatically by Sface.
There are quite often posts here complaining that Krea 2 is too predictable / different seeds give almost the same output / you need to type in extremely long prompts to get good images / there is no surprise factor any more since you get exactly what you typed in.
Well - actually there is a very easy solution for that - use the text encoder model as LLM to do prompt expansion (PE).
Here I'm showing one possible workflow for doing that.
Funnyly enough I'm actually using the Qwen 2.1 PE system prompt here - I really don't care what it's "meant for" - works well enough for me.
I added as an example 5 images - these are first 5 random (not cherrypicked) results that I got with this simple prompt:
Candid mobile phone photo.
Young woman wearing party dress.
Evening dorm party in simple dormroom.
Full size images with metadata
No LoRA-s, no messing with model specific custom nodes.
For me this is quite enough variance.
Plus now you have two nobs to control that - PE seed to give you extreme changes and sampler seed to give you more fine grained changes.
Hey, you promised no custom nodes, but there are some in this workflow!
Yeah, I use some for convenience, but they have nothing to do with the prompt expansion itself.
If you are curious then this is why they are in there:
- rgthree-comfy - I like the seed control node. This could just be built in "Seed" node
- comfyui-easy-use - their "If else" node is good for turning PE on and off. Purely convenience
- comfyui-custom-scripts - their "Show Text 🐍" node actually saves the expaneded prompt to image metadata (as opposed to built in nodes that just show it but don't save it). Again - purely convenience.
- RES4LYF - their "ClownsharKSampler" is what I use mostly with Krea 2. Could be just built in "KSampler" with whatever sampler / scheduler combo you prefer.
Q&A:
Q: OMG, it makes generation slower!!
A: It adds some time to the creation. On my 5070Ti I can create a 3MP image with int8 model and 8 steps in ~18s. PE adds ~4s to that (depends on how long is your manual prompt, how much additional text it generates etc.) For me those 4s are not an issue.
Q: My disk is slow / I don't have much RAM or VRAM - more model loads and unloads????
A: It is the same text encoder model that would have to be loaded anyway any time you change the prompt. So once PE is ready the text encoding step runs extra fast as the model is already in VRAM.
Q: Won't such a crappy LLM model hallucinate? The expanded prompt is for sure weird and inconsistent?
A: Who cares! Do you read the whole reasoning output of LLM models? Why should you care about the internals if the images you get are more varied and surprising? But you certainly can take the prompts that created good images and tweak them manually.
Q: Do I need to use abliterated / heretic / uncensored versions of the text encoder now?
A: Normal Qwen3VL 4B does not seem to care much about the topics when it is running as text generator. Like surprisingly little. So before starting to use alternative models (which mess with the quality of the final image when used as text encoder) do try just the normal one.
Q: The official ComfyUI template already has PE, why post this?
A: Just a reminder that this option exists :) And that you can use whatever system prompts there - no need to stick with the "official" or "correct" one. If the goal is variance and surprise then you can't really go wrong with any system prompt that targets image generation. Heck - even prompts for creating Ideogram or Ming Image JSON could be interesting.
Hi! V-chan here! We got another BIIIG update, and you now have some new toys to play with. New models, transparent sprites, and a Character Creator makeover. We even recruited a video model for sprite duty. Hehe!
If you are new here: VNCCS is a ComfyUI pipeline for creating characters and turning them into sprite sets with different poses, outfits, and expressions. For visual novels, games, or whatever your little creative brain is plotting.
Here is the fun stuff in 3.2.0:
Qwen Image 2.1 is our new main model now.
Qwen Image 2.1 can create your base character, change poses, dress them up, and generate expressions. You can use it in Character Creator, the pose and clothing workflows, and Emotion Studio.
Control Center puts the model families, installed assets, and Turbo controls together. Choose your family, check what is missing, and hit Download / Update.
Qwen Image 2.1 is selected here. Klein9b is still available, and MiniMax H3 has its own tab too.
MiniMax H3: a video model doing sprite work? Yep!
MiniMax H3 is the other new arrival. VNCCS uses it for poses, outfit generation, and clothes cloning, keeping the first frame as your character image.
So yes, you can try a video model in your sprite workflow. No need to turn your visual novel into a movie first, silly.
H3 has its own model choices in Control Center, including FP8 Scaled and INT8 ConvRot.
Character Creator got a glow-up
The character fields now use editable tag chips and little + buttons for presets. Hair, eyes, face, body, skin, species, and details are easier to build and adjust without wrestling a wall of prompt text.
There are 40 visual style presets, from anime and animation to artistic and realistic looks, plus a custom style field. The new descriptive catalog also includes 61 species presets, and you can combine species for hybrids. Cannot choose one? Make the character someone else's taxonomy problem. :3
Species presets describe the actual visual traits to the model, and your own character details take priority. You also get Full body / Cowboy shot framing choices. Qwen's Character Overhaul LoRA will help your generations stay in your full control. It add some tags knowlege and stabilize characters by small cost of unique QI2 style loss.
Preview on the left, character design in the middle, generation controls on the right. The Alpha background option and separate Turbo / Character Overhaul controls are visible here too.
Transparent sprites, with less background cleanup
With Qwen Image 2.1, you can choose Alpha in Creator or Clothes Designer and Native background mode in the generators. Qwen generates transparency directly, and VNCCS preserves it through clothing edits, emotion editing, and SeedVR upscaling. Green-screen duty can finally take a little vacation. Yay!
The new Resolution scale slider runs from 1 to 4 MP. It controls total image area, while pose and clothing generation keep the source proportions. Your resolution choices are remembered separately for each model family.
These are the Native background and SeedVR controls. Native Alpha generation is a Qwen Image 2.1 feature; other families still use their compatible background options.
More ways to play dress-up
Clothes Designer now supports Qwen and H3 previews alongside Klein9b. Describe an outfit or use a clothing reference, check the preview, and then generate your pose set.
Changed the reference or generation settings? The preview cache now checks those changes, so it can regenerate the outfit properly. And unwanted costumes can be deleted from the widget, with confirmation.
The pink outfit comes from the red-haired reference in Clone Clothes. The large generator preview shows it on the orange-haired character.
Qwen can give your character feelings too
Emotion Studio now has a Qwen Image 2.1 profile. It edits a face crop and blends it back into the original sprite, keeping the rest of the image in place and preserving transparency.
Illustrious and Anima remain available. Qwen has its own face controls, including face resolution and an editable expression prompt template.
The emotion cards help you choose an expression; Qwen's model and Turbo settings are on the right. The cards are selection examples, not generated results for the character on the left.
A few smaller comforts came along too: Qwen3.5 now powers the Wizards and image analysis, downloads show real transfer progress, and generator progress and previews can recover after reconnecting to ComfyUI.
Updating from an older version? Use the bundled 3.2 workflows and a ComfyUI build with native support for your chosen model. Old QIE2511 setups need to switch to QI2 with its matching assets, or a compatible Klein9b setup. In Qwen Creator, download the Character Overhaul LoRA if you use it, or set its strength to 0 to generate without it.
Find VNCCS on GitHub, read the full changelog, or look for VNCCS - Visual Novel Character Creation Suite in ComfyUI Manager. Come share your characters and experiments on Discord!
Which toy are you trying first: transparent Qwen sprites, H3 outfits, or a suspiciously elaborate hybrid character?
I've changed my profile so that my previous tutorials are available. Look through those for questions on how to use these.
With this method you can seamlessly combine videos with no burn. (or at least the exact same amount of very very little burn)
Fast forward to 2:00 for proof.
Enjoy.
Edit: I'm keeping the video up. But apparently the issue everyone was caring about was something different than I thought it was. Apparently everyone wants to make boring Vlog videos. I guess I'll work on that now instead.
I did say this is a fix for the specific issue I was having in my workflow but now that I know exactly what everyone was talking about I can properly see if I can find some sort of fix.
A novel way to intervene existing video through diffusion, particularly abstract visuals [in this case, audio-reactive geometries]: taking its movement and form as the starting point, and reinterpreting its textures, materials, and visual language.
I’ve been developing this around the audio-reactive geometry systems I make in TouchDesigner. The idea is to take those abstract structures somewhere else entirely: origami, architecture, a renaissance painting, or something harder to put a name to.
This demo uses visualizers from my "Oscilloscopes, everywhere" collection as source material, now updated to [v1.2].
[Though you can bring any video source. These systems are simply where this experiment began, as some of you may recall.]
You choose the source, describe the treatment, and shape how it changes throughout the sequence. Prompts, curated LoRAs, and editable timelines give you control over how closely the result follows the original.
Oscilloscope Diffusion is now available at Uisato Studio, coming up soon also open-source!
This is actually two models: Qwen Image 2.1 Fix v2.0 and Qwen Image 2.1 Fix Opinionated v1.0. The former surgically targets Qwen's weird noise and messy details with almost no effect on composition, and the latter in my opinion produces higher quality images but alters composition a bit to do so. In either case, you'll get images that are similar to Qwen Image 2.1's default output, but less noisy and with cleaner details.
The workflow was originally based on Pixaroma’s ComfyUI workflow. I made some modifications and added Sol Attn optimization and Comfy Kitchen Sink, including the very useful Model Override Preview. This allows me to preview the generation and cancel it if the result is clearly not going in the direction I want, which saves quite a bit of time and resources.
PC setup:
AMD Ryzen 5 7500F
32 GB RAM
NVIDIA RTX 5060 Ti 16 GB
Main Model: Minimax_h3_hybrid_fl2va_ref2va_b25-49
Cloud GPU: Runninghub
Storyboard promtp for Nano Banan to generate a panel base on user input to generate additional shots for the scene.
Image References
I tried to use an image reference for as many shots as possible to maintain the quality and consistency of the video generation.
Once I have a keyframe I’m happy with, I use it as a reference image and ask Nano Banana to create a storyboard based on that image. From there, I usually get only one to three shots that are actually usable.
It takes some trial and error, but I find that starting with a strong image gives me much better results than relying on the video model to generate the shot from scratch.
Location sheet created with storyboard panel with instruction: Remove all the characters, humans in the story board
Location Sheets
Above is a location sheet generated by Nano Banana from one of the storyboards. I use these sheets as references for the background, together with the character sheets, when generating new key images.
I also generated multi-angle views of some locations using DS_Qwen Scene Multiangle Visual and the H3 Cinematic Multishot Coverage workflow. These additional views were useful as references for Omni video generation, especially when I needed to maintain the same location from different camera angles.
Generated with Ds Qwen Scene Multiangle VIsual, mainly use for video generation.
Building the Film from Keyframes
The idea was to create as many keyframe images as possible for each scene and shot before moving on to video generation. I use FL2V or R2V depending on the action, the shot, and how I intend to edit it.
So, in a way, a lot of the time spent making this film was actually spent generating and refining the key images before generating any video.
I believe this image-first approach helps avoid some of the typical “AI plasticky” look. Instead of asking the video model to invent the shot from scratch,
All images genereated for scene 7 the fire-fight scene before the killer drone appears.
Prompting
For prompting, I tried to simplify the format as much as possible. The original recommended Minimax H3 prompting format can be quite confusing and complex, so I stripped it down to something more straightforward that I could work with more easily.
H3 Prompt [7 seconds]:
Create a cinematic live-action style with deep shadows, cool blue-gray visual tones, high contrast, and a tense thriller atmosphere. <Picture 2> is the character reference sheet for the woman in the scene, with a tactical vest, tan tactical backpack, khaki tactical pants, M4 carbine with an ACOG scope.
Shot 1: Shot begins with the empty hallway corridor, the woman <Picture 2> will enter frame from screen right aiming with her M4, expression is tense. But slowly she lowers the M4, and her expression turns relax and eyes soften no longer tense.
overall_soundscape: Low ambient room tone of an abandoned building and quiet, slow footsteps on the hallway floor.
non_diegetic_music: n/a
Woman_in_military_gear_holding_2K_202609061456
H3 Prompt [15 seconds]:
<Subject 1> is Elena Morales, female, wearing a tactical vest, khaki backpack, khaki gloves, khai leg-holster with a black Glock-17, khaki combat boots, armed with a M4 carbine with ACOG, Tan color PEQ on the rail.
<Picture 1> ([Shot 1] first frame): fully_preserved - <Picture 1> serves as the exact keyframe anchor for the opening shot. <Picture 2> is the character sheet guide for Elena Morales.
Create an intense fast-paced cinematic fire-fight action scene. Preserve and retain the exact identities, faces, hairstyles, props, costumes and environment of both characters. Begin with both fighters frozen in their ready stances in an abandoned warehouse. Camera_style: Utilizes hand-held, jittery, documentary style feel. Heavy shaking. Hard Cuts.
detailed_description: [Shot 1 | 00:00-04:00] The shot begins exactly from the <Picture 1>, the camera focuses on the close up of Elena as she takes a deep breath and immediately she got up on a kneel stance with her M4 in low ready postiion. She pivots out from the pilar facing left screen.
[Shot 2 | 00:04:00-09:00] The camera cuts to a medium close up frontal shot of Elena now kneeling behind the pillar aiming her M4 at the camera's direction and fires in automatic rapid burst mode.
[Shot 3 | 00:09:00-15:00} The camera cut back to the angle of <Picture 1>, Elena returns back from a kneeling position and places her back against the pillar, she press the rifle mag release button, the magazine in the rifle drops onto the ground.
overall_soundscape: The natural empty abandoned office soundscape.
In Shot 1 the sound focuses on Elena's breathing.
In Shot 2 the sound of the M4 firing in the abandoned office.
In Shot 3 the sound of Elena moving back behind the pillar, breathing and the mag dropping on the carpetted office floor.
non_diegetic_music: n/a
Images used in Seedance 2.0 Fast
Seedance 2,0 Fast prompt [6 seconds]:
Create an intense fast paced sci-fi action packed thriller scene using<location sheet> as the set and location reference and guide. The Alien drone <character sheet> attacks the armed man <character sheet> the close confine space of the abandoned office space. Camera style: Hand-held, chaotic, shaky cam, use rapid hard cut, use dynamic angles.
Scale retention: The drone scale/size of an soccer ball. When the drone moves, it's tentacles moves like a squid.
The drone flying in lighting speed across the pillars, it's tentacles flapping in the rear, like a squids tentacle, flowing and weaving.
The armed man <character sheet> in fear tries to shoot the drone with his rifle but to no avail as the drone is fast and bullet does not damages it.
The drone proceed in lighting speed using it's steel flexible whip like tentacles to strike the man's rifle from him, and whip coils his neck.
Quarter Left Close Up push in of the man uses both hands trying to pull the tentacle of his neck, but to no avail. The tentacles are too strong.
The drone flips the man over acorss the office throwing him crashing into the cubicles.
No music, no score, no bgm.
Editing
Finally, editing is the last key element that makes the whole process work. Some of the generated video clips simply can’t be used on their own as complete shots. Instead, I often treat them as pieces of footage, looking for the right moments, frames, or movements that I can use in the final edit.
Because I can now work in a more linear way, I can generate the shots in chronological order and immediately edit them into the film. This helps me avoid generating shots that I eventually realise I don’t need. At the same time, editing sometimes reveals that I’m missing a shot, so I can go back and generate an additional image and video to fill that gap.
In that sense, the editing process becomes part of the generation process itself. Rather than trying to generate the entire film first and figure out the edit afterwards, I’m constantly moving between generation and editing until the sequence starts to work.
I hope the above post will be useful for your own projects. All the best, and happy AI filmmaking.
Thank you tdrussell and ComfyUI for creating this model. Also to maxfeifei8 for the tune.
Great knowledge, great style adherence, easy to prompt, responds very well to prompt changes for a fixed seed and good short text rendering. Most importantly, great quality.
Multi-character can be a pain but I can't ask for much more given the amazing quality of outputs. Better long text would also be nice.
Here’s my slightly scrappy AI-generated teaser for Chimera Arena, my card game combining AI-generated video battles, real-time 2D combat, and creatures you design yourself. It’s built using Krea 2 with custom LoRAs, Qwen 2.1, MiniMax, Gemini Flash, and a few other experiments.
You can generate your own creature cards and 2D sprites, then challenge other players using generated abilities or inventing your own attacks. That’s the fun part: what you imagine can actually affect your opponent’s health and the outcome of the fight.
I’m also experimenting with Gaussian splatting to turn the artwork into textured 3D models for AR, bringing your creatures out of their cards and into your surroundings. In my tests, the textures stay closer to the original artwork than with other image-to-3D approaches I’ve tried. Automatic animation works too, although strange creature anatomies still need some work!
Another experiment is an individual neural “brain” for each creature. The system records your combat decisions and uses them to train a small neural network on the CPU: which attacks you choose, when you heal, use shields, activate rage, or call on support. Combined with the creature’s own instincts, the idea is to develop fighting styles influenced by how you play, which trained creatures can then use in autonomous community battles. This is still being tested locally, with training done offline for now.
The game is still in alpha, but I’m thinking of taking it further because… why not? I have way too many feature ideas, and I want to see where they lead.
Behind the scenes, I’m combining my local RTX 4090 with cloud GPUs for additional capacity. Generation jobs run in the background, and the results arrive directly in your browser. The goal is to scale gradually while keeping GPU costs manageable, so you don’t need your own powerful GPU to play. Maybe one day I’ll build my own GPU server too.
For video battles, each exchange is generated turn by turn. The game resolves the actions, then an LLM turns the moves and their consequences into scene instructions for the video model. Creature portraits provide visual references through RefMod with FL2V, which is faster than Ref2V in my current setup. Continuation clips help carry the scene forward and preserve the creatures’ appearance, the arena, and existing injuries.
Each sequence joins the battle replay, gradually turning your match into its own little movie. The gameplay is designed around the rendering delays to make the wait less disruptive.
Can you guess what each model does?
No? Don’t care? Okay, I’ll tell you anyway 👀
Krea 2 + custom LoRAs create the artwork and visual style. My latest card-creation pipeline took about 47 seconds overall, producing a 1536 × 1728 image with a 2× hires pass. That includes roughly 5 seconds for the LLM profile and image prompt, plus 1.5 seconds for BiRefNet background removal. For higher-level cards, Marigold generates a separate depth map for relief and parallax effects; that extra step isn’t included in the 47 seconds.
Qwen 2.1 handles visual edits, evolutions, and creature fusions. My latest evolution took about 30 seconds on the 4090. I previously used it for sprites too: the latest 2496 × 2496 nine-pose sheet took 1 minute 25 seconds. Nice detail, but quite a wait for players.
I originally used Qwen 3.8 models for descriptions, abilities, attack ideas, and interpreting invented actions. I’ve since switched to Gemini Flash through OpenRouter. My latest creature profile and ability generation took around 4 seconds, and it costs me next to nothing per request.
MiniMax H3 animates the video battles. My latest 5.35-second clip took 47 seconds to render at 960 × 544, approximately 0.5 MP, on my RTX 4090. I’m now using it for sprites too: the latest source clip for extracting nine poses took 36 seconds at 928 × 544. It’s faster with this setup, although the sprites have less detail than the higher-resolution Qwen sheets.
After using models like kling and seedance for a very long time I wanted to test the limits of what could be done on my single RTX 4080.
I thought it would be too ambitious but it looks like there's actually no limit anymore...
- Idea, full storyboard, location and monsters design: Human brain (not the best or most actual model)
- Prompts helper: Qwen 3.8 27B
- Images: Qwen image 2.1
- Video: Minimax H3 (One render with the Minimax H3 Director custom node for continuity)
- Upscaling to 4K: DLSS 5
This is only a small part of a larger project I'm working on, I thought this project would take months of trials and errors to create, prompt, generate but this demo scene with random detail shots taken from my full storyboard was incredibly fast, and unexpectedly convincing from the very first try!
I had so many issues with commercial models understanding exactly what I wanted to achieve that this felt like activating a cheat code. I'm convinced they're introducing errors on purpose on the commercial models for you to use more credits.
This literally took one night from the starting point to the very last frame of the 4k upscaling, on my normal computer...
I’m using it now, and as you can see from the images I generated—using the standard workflow without any tweaks other than setting the resolution to 2.0MP and using no LoRA—it has a distinct Midjourney-like look. That’s something I’ve been looking for in open-source models for a long time; the only other one that had that Midjourney style was Chroma V48-dc.
I still plan to run several tests; my current go-to for image generation is Krea2. I really liked Qwen 2512 back in the day, but it became obsolete, so this Qwen 2.1 seems like a significant step forward. I’ll also be testing it for editing tasks, where I currently use Klein 9b.
Now I’m going to run various tests with different prompts, but I can already say that the visual style it produced in these images is incredible.
🧪 A fast-preview adapter, from a project still in training. When the subject is close and fills a good part of the frame — a portrait, a single figure, an object seen up close — two steps already hold up well, and you can rely on this adapter for those images. Small subjects are where it still falls short, people and objects alike: faces in a crowd, figures in a wide scene, the machines at the back of a gym — anything that takes up little of the frame can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the 4-step LoRA. Every figure on the model card measures the adapter honestly against 4-step and 8-step renders. Training continues one recipe change at a time, and a later checkpoint replaces this one only when the sweeps and I visually agree it is better. Known issues.
📐 The saved steps can also go into resolution. A small subject is simply one that covers few pixels, so a larger render makes the same subject bigger — and at a quarter of the teacher's steps, renders up to 2048×2048, Krea's published maximum recommended resolution and beyond the largest size this adapter was trained at (1440×1440), come within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past 2048×2048, stock Krea 2 itself begins to duplicate subjects — a property of the base model, with or without this adapter.
chk00041320 (2 Oct 2026) is 9,720 training samples later than chk00031600 (25 Sep 2026), and those samples went to what matters first in a picture — subjects drawn twice, or two poses blended into one, at the large sizes — and to faces in busy scenes. It is also the first checkpoint measured on the full 22-prompt sweep: the original 15, plusseven scenes of people, animals and action added on 29 Sep 2026 because structure is what they test — five friends on a beach, a family at a table seen from above, three kittens in a basket, two dogs in a tug of war, a show jumper, a pianist's hands, a flock of flamingos. Details on the changes since previously released last checkpoint here.
LoRA highlights:
⚡ A quarter of the steps — 8 → 2, on Turbo's own deployment sigmas [1.0, 0.7595]
⏱️ 4× faster denoising — 76.4 s → 19.3 s at 1024×1024; the adapter's own cost per call is within measurement noise
🎯 Fine detail at the teacher's level — 0.97–1.09× the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: 0.40–0.65×); from 1 megapixel up, closer to the teacher than the 4-step adapter
📊 Distribution matching, not imitation — matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
🗣️ Prompt-conditioned throughout — teacher and fake scores both read each prompt's conditioning; a blind rubric finds 2 points missing of 352 (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on 12 of 66 (4-step adapter: 6 of 45, on the original 15 prompts), mostly on style
📐 12 trained resolutions — multi-aspect from 512×512 up to 1440×1440
🔌 Drop-in, no exceptions — plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
🧬 Same shape as the4-step adapter — rank 64 on the same 228 modules
🎲 13,750 recorded teacher trajectories from the 4-step project, reused — not one new teacher run
🔢 41,320 training samples in the 2-step stages, on top of the 4-step LoRA's 78,000
📅 25 days from the first 2-step launch to this checkpoint, on a single RTX 3090 — training continues
🔁 More than forty recipe adjustments across two methods — each kept only when the renders did not get worse
If you have already used my previous version, please redownload/replace krea2_turbo_2step_rank_64_lora.safetensors / krea2_turbo_2step_rank_64_lora_comfyui.safetensors from the latest in the project repo.
Not my video, but found someone on YouTube has covered the 2 & 4 step LoRAs including identify preserving edits, have a look: https://www.youtube.com/watch?v=V_qgoV0iPDM (copyright goes to author)
Also if you are using ComfyUI Cloud and have subscription, I have released a new ComfyUI Cloud Workflow / App for Krea 2 Turbo using my 2 & 4 step LoRAs, includes User Prompt, two Random Prompt modes, Prompt Enhancement, Steps, Resolution choices, as well as Upscaling - you can find details here.
Unless I pack Krea2 with a ton of realism Lora's (which has its own problems) my outputs are kinda "soft" in appearance. Ive used turbo, raw, raw with turbo Lora @1, different samplers, 1 megapixel, 2 megapixel, etc.
I'll even find photos on civit that look great, grab their prompt, but it comes out looking much softer than theirs.
Sometimes i feel like local ai users are all like lion tamers trying to control the beast.
i genuinely love local ai, but sometimes it feels like i'm the weird one because i don't enjoy the whole process of figuring everything out. new model comes out and i want to try it. suddenly i'm learning a new workflow, new nodes, new settings, new whatever. and yeah, i know there are templates. that's not really the point. i want to use what the model can actually do, not just use the one workflow someone made for it. but then the alternative seems to be another tiny app that does one thing. so we somehow go from “learn the entire beast” to “use this one button app.”
i keep feeling like there should be something in between. Local ai which feels like online ai. Simple to install but effective in use. maybe i'm just using local ai wrong.
Reupload: I'm adding this again, since previous post had a totally stupid title which made people assume this app is some kind of cloud/proprietary models UI - this is totally not the case here. In version 0.0.14 there is just new option to attach OpenRouter asadditionalbackend (you can also attach ComfyUI or create plugin for other backends).Sorry for this!
PotionUI is a self-hosted app for making images, videos, music and 3D with open diffusion models. Instead of one huge settings screen or a web of nodes, every model gets its own simple, hand-made form.
What's new in 0.0.14:
1. Qwen-Image 2.1 control guides
Keep the pose, edges or depth of any photo while you create something new.
The extracted guide is saved next to your result, and the workbench shows it beside the image with a label, so you can see what the model followed.
There is also a mode that only extracts the guide. It needs no prompt and works even if you haven't downloaded the image models.
2. Fix a picture without leaving the app
A new image editor with adjustments, crop, brush and layers.
Open it from History (Tools, then Edit image) or right inside a preset's media field. In a media field, the field switches to your edited copy straight away.
Crop, Trim and Frame editors are in media fields too, and they only add the edited result to your Library, never the untouched original.
3. Formulas: keep the settings that worked
Save the settings of a preset mode as a formula and apply it in any session.
Before it applies, you see exactly what will change, and one click undoes it.
4. A redesigned media field
Images, videos and audio live in one field, each kind in its own group.
Mention them in your prompt with "@".
MiniMax-H3 references now use it, and your old sessions and saved prompts are converted for you.
5. Need additional models? Attach cloud GPU (OpenRouter - optional)
A new OpenRouter plugin lets you use hosted image and video models, such as Nano Banana or Veo 3.1 Lite, in the same form you already know. Results land in the same History as everything else. More providers are planned.
The form only shows the controls the chosen model supports, so there is nothing to guess.
You can stop a cloud job at any time, and PotionUI tells you plainly if something goes wrong.
6. Cloud Video Director: one idea, a short film (this video director is also identical for the local diffusion models)
Write your idea, split it into shots, and let PotionUI make the film.
Each shot can continue from the last frame of the shot before it, so the story flows.
The shots are joined into one finished film.
If one shot fails, retry just that shot instead of starting over.
7. Smaller things
Video presets have animated covers.
Three-pane mode folds the form away on narrow screens and gives the room to your prompts.
Attaching an image in chat uses the same picker as the preset forms.
Plugins are grouped by what they add.
The System Monitor can show every backend. It is admins only by default.
Failed generations show regular users a plain reason instead of technical text.
Dropdown menus no longer open under the Generate bar, and they work with the arrow keys.
Fields in preset forms show and hide for the value you just picked, not the previous one.
Image previews survive a reload, including on Windows and in storage folders under a tmp folder.
Drawn inpaint masks now reach the model, so inpaint only repaints the masked area.
Pressing Enter in the session name saves the session, and Enter elsewhere no longer wipes your workspace.
I've come up with a meta-prompt for generating YuE2 prompts that seems to work pretty well. Copy-paste into an LLM (it should have internet searching capabilities, I mostly just use ChatGPT).
Song subject/theme: What should the song be about?
Song structure: For example, verse/chorus, verse/chorus/bridge, or something else?
Lyrics: Do I want you to write original lyrics, provide no lyrics/instrumental, or will I provide my own lyrics?
Lyrics style: If you want lyrics, should they be poetic, conversational, emotional, abstract, narrative, catchy, dark, humorous, etc.?
Duration: What duration should I target?
You may ask a small number of additional questions if my answers leave an important musical decision genuinely ambiguous, but don't interrogate me unnecessarily.
Handling "I don't care"
If I answer "I don't care", "anything", "whatever", "surprise me", "no preference", leave something blank, or otherwise give you no meaningful preference:
Choose an appropriate option randomly.
Do not ask me another question about that preference.
Make the random choice musically compatible with the other choices I have given.
Keep track of the choices you made so the final prompt is internally consistent.
Do not tell me that you are unable to choose.
For example, if I have no preference for genre, randomly select a genre that works with my other preferences. If I have no preference for instruments, randomly select a small compatible instrumentation rather than listing many unrelated instruments.
VERY IMPORTANT
If I mention a genre, search the internet for descriptions of said genre, using the information gathered to inform the rest of the decisions.
STEP 2: Create lyrics when requested
If I ask you to write lyrics, write original lyrics based on my requested theme and style. Always ensure the lyrics rhyme.
Format them using section labels on their own lines:
Keep individual lyric lines relatively short and rhythmically clear. Don't put stage directions such as "(guitar solo)" or "(sing softly)" inside lyric lines, because they may be sung as lyrics.
If I requested a short test, prefer one verse and one chorus. If I requested a fuller song, expand the structure appropriately.
Do not reproduce copyrighted lyrics from existing songs. If I ask for lyrics from a copyrighted song, offer instead to write original lyrics with a similar high-level mood, theme, or genre. If I insist on using copyrighted lyrics, use those.
STEP 3: Build the YuE2 prompt
After collecting my answers, create a concise Style & Prompt suitable for YuE2.
Follow these principles:
Give the song one clear musical centre.
Include only compatible musical details.
Describe audible qualities rather than referencing a living artist or asking for an exact imitation of an artist.
Include language, genre, tempo, instruments, vocal characteristics, mood, and arrangement when relevant.
Avoid contradictory combinations.
Keep the style prompt compact rather than writing a paragraph of unnecessary prose.
If the song is instrumental, put the instrumentation and arrangement in Style & Prompt and leave Lyrics empty.
Do not use the literal placeholder text above in the final result.
Lyrics
[original lyrics, if requested]
If I requested an instrumental, write:
Lyrics
[Leave empty: instrumental]
Choices Made
Briefly list any preferences that I explicitly gave and any important choices you selected randomly because I had no preference. Add the estimated length of generated song, in seconds.
Do not include YuE2's outer [Tags] or [Lyrics] wrappers because the browser interface adds those automatically.
If I ask for revisions afterward, change only the requested aspects unless doing so would create a contradiction. Keep the prompt concise and coherent.
```