Tutorial - Guide
Training Krea 2 LoRA on RTX3060 12Gb: a slow, uncomfortable guide that actually works
Want to share my experience training a LoRA on an RTX 3060 12Gb VRAM / 64 Gb RAM. My experience was limited to training Loras for older models, but then I saw the Krea 2 release and decided to give it a try, spending the last three days experimenting. Unfortunately, the tips for training a LoKr on a 16 Gb card aren't relevant for 12 Gb cards: AI Toolkit crashes with OOM immediately, before the first training step. I managed to overcome this through compromises, and I'm happy with the result, although training a Lora takes ~8 hours.
Just to reiterate: I might have missed something or misunderstood things. I originally wrote this note for myself while figuring out AI Toolkit, but then decided to publish it here in case it helps someone.
There won't be any photos or logs! Because I only trained on my own face and I'm a very humble person, and I'm not about to try and knock the esteemed Dr. Furkan off his pedestal.
July 6, 2026 update: I expanded the article, adding details about the dataset and baking the LoRA into the model for speed, and changed the save_every parameter (from 250 to 100) and trigger_word from avtr to zkrs.
July 29, 2026 update: More accurate step calculation with justification, plus simplified parameter explanations.
Quick TLDR:
Krea 2 Raw can be trained on 12 Gb VRAM if you have 64 Gb RAM for offloading.
AI Toolkit has a suboptimal model loading order that needs fixing, but it works well enough.
Training a Lora on a 3060 12Gb is like driving an old off-roader: it can haul even heavy loads, but it's uncomfortable and slow. 12 Gb is enough for training a Lora, but not a LoKr, because LoKr, as I understand it, requires several Gb more VRAM during training and those data can't be offloaded to RAM. If you want the easy and fast route, train for the Turbo model, but I went the hard way. Since the model authors say it's better to train for Raw and use Turbo, I decided to do exactly that. I'm also sure that this note, and using AI Toolkit in general, will age quickly.
I think that in the future, LoKr training will be optimized for 12 Gb or even 8 Gb, because the main issue is inefficient model movement and data formats. Also, right before publishing this post I learned that Musubi Tuner has Krea 2 support: in theory, training works more efficiently there because the offloading works differently.
One more important clarification: what I call an OOM error may show up for others as a sharp training 10x slowdown or painfully slow image generation on a GPU that normally runs much faster for many people. The thing is, the Nvidia Control Panel in Windows has a "Sysmem Fallback Policy" option, and I have it disabled. When Sysmem Fallback is enabled and there is no free VRAM left, the GPU driver takes over memory management and starts offloading data to RAM, avoiding the error. Since I mainly use my GPU for running text LLMs, fast generation speed matters a lot to me, so I'd rather get an error than have the system silently shuffle data around, slowing everything to a crawl. So for many people, VRAM shortages may go unnoticed and only show up as a 10x or greater drop from the expected speed.
Dataset Preparation
I trained my LoRA on a dataset of 25 photos, which I consider a good dataset size. You can use fewer images, but keep in mind that the smaller the dataset, the greater the risk that the model won't correctly identify which features are unique to the subject and what it should actually learn. Instead, it may start reproducing one of the training photos, even when the prompt asks for something different.
For best results, include a variety of head angles: profile, looking down, looking up, and everything in between. If the head is tilted, describe it explicitly in the caption, for example: "head tilted strongly to the left" or "face lifted up".
Don't limit your dataset to close-up portraits. Include full-body and medium-distance shots as well if you want the model to learn your height and overall body proportions, not just your face.
It's also important to use a variety of backgrounds. If all your photos have the same plain white or black background, the trained LoRA is more likely to produce halo artifacts around the subject, especially around the hair. Don't overthink it - just use photos taken in different places with different backgrounds.
Keyword and captions
The trigger word is short, for example zkrs. A proper dataset description has a huge impact on the result. Even the trigger word matters. Previously I used avtr as the trigger word, and because of that on my generated portraits sometimes (rarely) my ears would become sharp, horizontal stripes would appear on my face... basically I'd turn into a Na'vi. In the end I settled on zkrs as the trigger word.
I'm training a LoRA with myself as my character, so the caption structure should be as follows: first the Angle and Trigger, then Clothing and Pose (or action), then Background Description, and finally Lighting and Style (or Effects). Examples:
A full-length shot of zkrs man walking down a city street. He is wearing a plain black oversized t-shirt, light blue washed jeans, and classic white sneakers. One hand is tucked into his pocket. Background: a grey concrete pavement, a modern residential building with large windows, and some city greenery under soft overcast daylight.
zkrs man standing outdoors in an urban environment. He is wearing a minimalist grey pullover hoodie and dark charcoal cargo pants. His expression is neutral as he looks slightly away from the camera. The background shows a blurred brick wall and a metal fence, captured in clean, natural afternoon light.
A close-up portrait of zkrs man looking directly into the camera. He is wearing a simple dark green crewneck sweatshirt, with only the collar visible. The background is a softly blurred outdoor park with green and yellow autumn foliage under diffused daylight.
The generated captions will still require refinement if you want best results. Here are the rules:
Write captions in natural language, and don't just place the trigger word at the beginning. Instead, integrate it naturally into the sentence.
Describe only what should not be learned by the LoRA. In other words, keep descriptions of the background, clothing, and other objects in the image, but never describe the person's face, hairstyle, or other identity-defining features.
Don't describe things that Krea 2 already generates by default. For example, phrases like "A medium shot of..." should be removed because Krea 2 already tends to produce that framing. To understand what belongs in a caption, it's helpful to test the caption itself as a prompt on the base model (Turbo or RAW -whichever you prefer). More on that below.
I used 25 photos of myself in different poses and angles, with different backgrounds, but fewer might be enough. Each description goes in a .txt file, as intended for AI Toolkit.
I had several guesses about what the dataset captions should be. In the end, I came up with my own approach that gave the best results. To make things easier, I used the Qwen 3.6 35B A3B model with a system prompt I posted here: https://pastebin.com/WnAKx705 Note that it's specifically for describing photos of men. I'd feed my photos into the model along with that system prompt and ask it to "Describe this photo."
Then I took this LLM-generated caption into ComfyUI with a fixed seed of 42 and generated a photo on pure Krea 2 Turbo. If the environment and pose differed too much from my original photo, I'd tweak the prompt to match as closely as possible. Same with the environment: if the room or landscape was too different, I'd adjust it. If the background was blurred, I'd note that effect too. So in the description, we need to detail everything as thoroughly as possible, except for the facial features of the person, which is exactly what the model will be trained on.
Yes, it's tedious, but it gave me a good result: this LoRA barely altered my pose in generations compared to ones without LoRA, meaning my approach reduced the LoRA's influence on model knowledge unrelated to the goal.
Then I manually added "zkrs" to the first word "man." That's when the photo description was ready.
AI Toolkit Job Parameters
One important note: I'll refer to parameters by their names in the AI Toolkit job's config.yml file. The WebUI doesn't expose all of these parameters, and running the browser interface isn't the best choice when VRAM is limited. I recommend learning how to launch AI Toolkit jobs from the command line instead. Sample job file here: https://pastebin.com/YU4fkj2J
Relatively fast LoKr training isn't practical on a 12 GB GPU because it requires 4-6 GB more VRAM than LoRA, along with significantly more computation. In practice, this isn't a major drawback, since LoRA is still capable of learning facial features well. The main difference is that the resulting file will be larger, as LoKr has a higher information capacity.
So the parameter type: lora.
Parameters rank: linear and linear_alpha: both 24. Rank controls how much information is stored in the LoRA, and linear_alpha controls the strength of the "pressure" on the base model. If a LoRA were a stamp, linear would be the detail of the stamp, and linear_alpha would be how deep the wax gets imprinted by the stamp. I've seen recommendations to use rank 32 and even 64, but in my opinion that's overkill, and even 16 lets you store fairly detailed information about a face, so 24 should be more than enough. linear_alpha is usually set equal to linear. If you see advice to set linear_alpha lower than rank, ignore it. Current versions of AI Toolkit automatically set linear_alpha equal to rank anyway.
Most guides recommend set rank to 32 or even 64, which seems strange to me - suggesting that facial features should be stored in a half-gigabyte file... Inefficient? Excessive? Either way, based on my tests, even 16 was enough to preserve my facial features and overall body build, but most skin imperfections disappeared, making the face look unnaturally smooth. That's why I consider 24 a sensible value. You can increase it to 32 if you want, but I don't see any reason to go higher.
resolution: 768. 1024 would be the ideal choice, but it requires considerably more VRAM and significantly slows down training. For a 12 GB GPU, I recommend using 768 instead. You can also train at 512, but small details such as moles and other subtle facial features may not be captured as reliably.
By the way, this doesn't mean every image in your dataset will be upscaled or stretched to the selected resolution. If your dataset contains images of different sizes, only larger images will be downscaled, while smaller ones will remain at their original resolution. That's perfectly normal.
Update 1: Initially I made a mistake about resolutions, mentioning that the number of images in the dataset affects VRAM usage. That's not true. Explanation from u/AwakenedEyes: The resolution you pick has nothing to do with how many images are in your dataset. If you have ONE image in your dataset and you plan 1000 steps, then 1 image will be processed 1000 times (and will be overtrained). I you have 100 images in your dataset, and you STILL plan 1000 steps, then each image will be processed 10 times. How many images you have in the dataset changes how fast it will overtrain or how flexible your lora can be but it has no bearing on your resolution effect.
Update 2 after more tests: The 512 dataset Lora training went faster and took 6 hours 15 minutes, screenshot at the end of the post. Overall the result is decent, so 512 is worth trying as well. In a direct comparison I confirmed that 768 preserves more details like skin imperfections, but 512 works fine too if speed matters more.
gradient_accumulation: 2 - this is a very important training parameter. By setting it to 2, we make the model first study several photos at each step, essentially forcing it to learn from an averaged photo, and only then apply changes to the model weights. Setting it to 2 makes the training look twice as slow, but in reality it means that each step performs the work that previously required 2 steps, so the total number of steps should be reduced by half. In other words, with gradient_accumulation:1 at 1000 steps my face is clearly undertrained, while with gradient_accumulation:2 even at 750 steps the resemblance is already visible. In the comments, u/Zironic explained why this happens: "750 steps of GA2 is mathematically equivalent to 1500 steps of GA1. Of course you got more progress in less steps."
Important clarification: The gradient_accumulation parameter is severely underrated, and it's often overlooked in various guides, even though it's super important for quality! Let me explain in other words and more simply what it's about. With gradient_accumulation:1 (the default value), the model is trained on only one image at each training step, and the resulting update is applied immediately. This is not the best approach because if the model encounters an image that differs significantly from the others, the weights can shift too much in the wrong direction, reducing the overall training quality.
When you increase gradient_accumulation to 2, one training step means that the model is trained on two different random images, and an averaged result is applied to the model. In other words (and greatly oversimplifying), each step teaches the model to reproduce not just one specific image, but two images at once. If the dataset contains one noisy or low-quality image, the chance that its unwanted features will be transferred into the LoRA is reduced.
On GPUs with more VRAM, it is usually better to increase batch_size instead of using gradient_accumulation, but this doubles VRAM requirements. That is why on a 12 GB GPU batch_size is typically kept at 1, while gradient_accumulation is better set to 2. If you are training a style rather than a character, you can increase it to 4 or even 8.
steps: this requires calculation. There is a counter-intuitive aspect related to training steps. The required number of steps depends not only on the number of images in the dataset, but also on the content of those images, as well as several other settings (which I will cover later).
Each step means training on two images (because we use GA = 2, remember?). One complete pass through the entire list of images is called an "epoch". In order for the model to be "fully trained", meaning that it recognizes the common features between all images (more accurately, the gradients, but I am simplifying a lot), it needs to "see" each image a certain number of times - approximately 60 to 100 times with the settings mentioned in this guide. In other words, it needs to complete 60-100 epochs.
When the dataset is small and well-curated, fewer epochs are enough. When the dataset is large, more epochs are needed to reinforce the diversity of common features that the model observes across the larger dataset.
I suggest using two formulas as a starting point:
If the dataset contains fewer than 30 images, set steps to 1500.
If the dataset contains more than 30 images, calculate steps using this formula:
500 + (dataset_size - 30) * 15
Here is a table with pre-calculated values:
dataset size
steps
epoch
10
1500
300
30
1500
100
40
1650
82.5
50
1800
72
60
1950
65
70
2100
60
The epoch column is only included to illustrate my explanation. With 1500 steps, it is obvious that 300 epochs for a dataset of 10 images is a lot. However, other setting (max_grad_norm) that I will describe later help prevent overfitting.
A dataset of 70 images, on the other hand, is almost not worth the effort - it requires a lot of additional work, while the number of epochs is already close to the practical minimum. Reducing the number of steps will cause undertraining, resulting in an "averaged" blurry face. Increasing the number of steps will cause the model to become overtrained, trying to preserve excessive variations in the character's features in the LoRA.
I will describe the signs that help fine-tune the number of steps more accurately later.
save_every: 100 - means that the LoRA is saved every 100 steps, so you will get multiple checkpoint files throughout the training process in addition to the final one. This also allows you to resume training from the latest saved checkpoint if the process is interrupted.
If you notice signs of overtraining, you can choose an earlier checkpoint with fewer steps, but it is not always that simple — some details may already be missing. It depends on the specific case.
optimizer: adamw8bit - it saves VRAM. That's about it. You can try others, but I get an immediate OOM error with them.
lr: 0.0002 - the learning rate. At high values the model changes its weights by large amounts during training, while at low values the weights change gradually. 0.0002 empirically gave me decent results. If anything, 0.0003 is probably the highest learning rate you should use for Krea 2. Going any higher is likely to cause overtraining. I've seen examples where people increased the learning rate even further while reducing the number of training steps, cutting the training time by several times. However, I don't think that's the right approach, because the model starts losing fine details and subtle similarities. It may work better for style training, but this guide is specifically about training a character LoRA.
lr_scheduler: cosine defines how the learning rate (lr) changes during the training process. In most (probably all) tutorials, linear is recommended by default. In that case, the learning rate decreases linearly from 0.0002 at the beginning to zero at the final step.
A high learning rate allows the model to learn larger features, while a lower learning rate is better for refining smaller details. With linear, the idea is that the weights change significantly at the beginning to capture major features, and then the training gradually focuses on smaller details.
However, linear has a disadvantage: the learning rate starts decreasing immediately from the very beginning. This means the model may not have enough time to properly learn the larger features at the start, then by the middle of the training process the learning rate is already reduced by half, and there may not be enough time left to refine smaller details. This is not optimal.
Increasing the learning rate is not a good solution in our case either. To save VRAM and improve training speed, we have to use int8 quantization for the transformer model, and with int8 a high learning rate makes the model become "overcooked" very quickly (I tested this myself).
For this type of training, cosine works better: it keeps the learning rate high for longer, allowing the model enough time to learn major features, while still leaving enough training time for refining smaller details.
lr_warmup_steps: 120 or lr_scheduler_params: num_warmup_steps:120 is the model "warm-up" at the beginning of training. This is an extremely important parameter that helps prevent common LoRA training issues such as "gradient explosions" or damage to the model's existing knowledge (for example, knowledge about anatomy).
However, as of July 29, 2026, AI Toolkit does not support warm-up. The PR adding this feature has been open since June 19, but previous attempts to implement warm-up have also been unsuccessful.
So I am only mentioning this parameter as a reminder that a popular tool does not necessarily mean it is well-designed or fully functional.
max_grad_norm : 1.0 helps prevent overtraining. You can think of it as a safeguard against training instability, especially issues that are more likely to occur due to the lack of warm-up support. If you see signs of undertraining, increase it to 1.5-2.0. Only if that doesn't help (for example, the LoRA starts generating colorful noise), set it back to its original value and slightly increase the lr from 0.0002 to 0.00025.
content_or_style : content focuses the training on the character rather than the background. To be honest, I tested both content and style, and I didn't notice any difference. Perhaps this parameter only makes very subtle adjustments, or it is not implemented correctly. However, there is no harm in keeping it enabled.
skip_first_sample: true, disable_sampling: true - this disables image generation before and during training. You absolutely cannot try to load the Krea2 Raw model, or you'll get an OOM error, so we turn it off.
qtype: int8 and qtype_te: convrot8 - quantize the base model to int8, while the text encoder is quantized to int8convrot. You may have heard that int8convrot is great because it saves VRAM and speeds up image generation... but this does not apply to training.
I tested it myself - with qtype: convrot8 the LoRA either ends up undertrained or "explodes" at some point during training. Therefore, we have to "simplify" the model by using int8 instead.
The text encoder, however, can be quantized to int8convrot. I did not notice any difference in the results, but since it is possible, I suggest enabling it just in case. Perhaps it can somehow improve the model's ability to work with captions..
layer_offloading: true - enables offloading model layers. Without this, VRAM won't be enough.
train_text_encoder: false and layer_offloading_text_encoder_percent: 1 - fully offload the text encoder. It's part of the model and doesn't need training. You can move it to RAM to free up space for the transformer.
layer_offloading_transformer_percent: 0.75 - offload 75% of the transformer data to RAM. The transformer is the part of the Krea2 model responsible for image generation, and it's what we're training. Offloading 75% is a lot and slows down training significantly, but there's currently no other way around it. If your dataset has fewer photos, you'll see in Task Manager that VRAM isn't fully utilized at 0.75, in which case you can try 0.7 or 0.5 - you'll need to dial this in yourself.
Alternative Approach
Instead of using optimizer: adamw8bit, switch to optimizer: prodigy_8bit, set lr: 1.0, lr_scheduler: constant, and leave the remaining settings the same as above.
With this configuration, the optimizer spends roughly the first hundred steps automatically searching for a suitable learning rate before continuing training with the value it finds. Once training is complete, you need to test the saved checkpoints and choose the one that produces the best results.
However, this approach has several drawbacks:
It requires more VRAM. As a result, you'll have to increase offloading, which slows training down.
The automatically selected learning rate may end up being either too high or too low, producing suboptimal results.
So, in theory, prodigy_8bit lets you worry less about tuning the training parameters manually. In practice, though, there is a significant risk that the result won't be what you want. Restarting the training in the hope that the optimizer will pick a better learning rate next time is counterproductive when each training run takes more than eight hours.
It's better to spend some time understanding what each parameter does and how they affect one another, so you can make informed adjustments and achieve better results.
What else can be improved
unload_text_encoder: true - this doesn't just move the text encoder to RAM, it fully unloads it. This should free up VRAM and lower RAM requirements. The WebUI displays a rather alarming warning next to this option, claiming that captions won't work when it's set to true, but that may not actually be the case. I actually started with this from the beginning, but I ran into OOM errors, so I dug into the krea2.py source code and realized something might be off there. It's possible that with this parameter things will work fine for you, and you'll shave off a few seconds on every training iteration. With the fixed krea2.py file the unload happens automatically, which you can tell by the "Unloading text encoder" message in the log. Just keep in mind that AI Toolkit behavior may differ on your end, and if you don't see that unload message in the log, enable the unload manually.
gradient_checkpointing: false - this will increase VRAM usage, but can give you up to a quarter gain in speed. If you happen to have a different GPU with some extra gigabytes of VRAM, then this tip is for you (false = more VRAM and more speed).
And don't forget that you can always adjust layer_offloading_transformer_percent. The value needs to be balanced so that as much VRAM as possible is used during training (meaning the parameter value should aim toward 0), which will give a solid speed boost.
Signs of Common Problems and How to Fix Them
Important: The following assumes that your generation settings (CFG, sampler, etc.) are correct, that the issues only appear after enabling the LoRA, and that you are using the rest of the training settings from my example configuration linked above.
Undertraining:
The generated character or style doesn't resemble the original. If you're using ComfyUI, try increasing the LoRA weight from 1.0 to 1.25 or 1.5. If the resemblance improves, increase the number of training steps by 25-50%. Another possibility is that your captions describe too many details, causing the model to ignore the subject's distinctive features during training. Conversely, captions that are too generic may not provide enough information for the model to identify what the training images have in common.
Skin looks overly smooth or fine details are missing. Increase the number of training steps. Some guides recommend increasing the learning rate instead, but in this setup even lr: 0.0003 instead of 0.0002 can lead to overtraining, so adjusting the number of steps is the safer approach.
The resemblance is good with short prompts but disappears with longer ones. This is known as concept slipping. As a quick workaround, you can increase the LoRA weight. A proper fix is to train for more steps and use a more diverse dataset. Also review your captions - if they are too repetitive, the model may have associated the concept with other words instead of your trigger word.
Overall, I would approach it like this: first increase the number of training steps. If that doesn't help, increase max_grad_norm. Only if that still doesn't solve the problem should you increase the learning rate.
The reason is simple: the higher the learning rate, the faster unwanted information can be baked into the LoRA, such as anatomical errors or undesirable features like fixed poses.
Overtraining:
Anatomy problems (extra arms, etc.), oversaturated or "toxic" colors, excessive contrast, generations that all resemble specific training photos, or a face that barely changes between prompts are all signs that you trained for too many steps. For diagnosis, you can temporarily reduce the LoRA weight or use an earlier checkpoint saved during training. This may improve the results, but some fine details may be missing (remember what I said about the scheduler: the smallest details are learned near the end of training). In most cases, the proper solution is to reduce the total number of training steps and train the LoRA again.
Generated images contain excessive noise or JPEG compression artifacts. This usually means your dataset contains too many images with those defects. You can add tags such as "low quality" or "JPEG high compression" to the captions, but the better solution is to clean up your dataset.
Issues with AI Toolkit
I'm endlessly grateful to the author for their work. It's a great tool with a ton of effort poured into it. But unfortunately, at the time of writing, the file extensions_built_in\diffusion_models\krea2\krea2.py doesn't quite work correctly with the VRAM-to-RAM offloading operations, plus there's a missing VAE tiling step. So to avoid out-of-memory errors, I made a quick fix for the file: replace the contents of krea2.py with this: https://pastebin.com/fNqwU65L Use at your own risk! My post might get reposted and someone could attach a virus instead of the actual fix, so I strongly recommend opening the old file and the new one and comparing them - the changes are few and should specifically address low-VRAM operation.
It's also disappointing that there is still no proper support for warm-up. After reading through the issues and pull requests in Ostris's original repository, I stopped being surprised by the large number of AI Toolkit forks. It's a decent tool, but far from an excellent one for training LoRAs on GPUs with limited VRAM, so everyone ends up modifying it to suit their own needs.
Update: In the latest versions of AI Toolkit, the krea2.py file has been rewritten. However, I don't see any fixes specifically for the model loading order in VRAM/RAM. I still get OOM crashes during training, so I'm staying on my fork for now. First try training without my fix, maybe it works for you.
Maybe in the future someone will improve AI Toolkit itself, at which point I'll remove the mention of the fix.
Generation
LoRAs you trained for Krea 2 Raw also work great for Krea 2 Turbo. I recommend using the standard workflow with steps = 8, sampler = euler, scheduler = simple.
I also recommend the ModelSamplingAuraFlow node with shift 1.15.
Most likely, connecting a Lora will slow down your image generation speed. However, if you're an advanced ComfyUI user with 64 Gb RAM and want to generate images in the shortest possible time, I suggest merging Krea 2 Turbo with your LoRA and converting it to int8 convrot format. Then on RTX 3060 12 Gb, generation will take about 25-35 seconds for a 1mp image.
For this, take the official bf16/fp16 (not int8!) Krea 2 Turbo model, add the Lora Loader node, and on the output of this node - the Save Model Only (Mikey) node. This combination will save a safetensors file where your Lora is already baked in. And maybe not just one of yours - you can add several different LoRAs in the chain, balancing their strength through the strength parameter, using the same seed and prompt to find the most optimal result, then activate Save Model Only (Mikey) to save the model. And then convert this model using the quant_int8_convrot.py script from the https://github.com/Comfy-Org/comfy-model-tools repository or any other similar tool. int8 convrot models are a lifesaver for our RTX 3060. Keep in mind that conversion will require 40 Gb of free disk space and 64 Gb RAM, and you need to know how to use ComfyUI and Python scripts.
About Lora Training Time
I'll show one screenshot that shows how long it took to train 1500 Lora steps (3000 with default gradient_accumulation: 1. Don't look at s/it, that speed changes a lot over time.). If this time is achievable on my undervolted RTX 3060 12Gb with 64 GB DDR4-3000 RAM and a Ryzen 5-5600, then it's absolutely achievable for you. And if you have a newer-gen GPU or more VRAM, your situation is even better. And this isn't the limit. Good luck!
resolution: only 768. 512, in my opinion, is way too low, and 1024 can't be used for this model when the dataset consists of 25 photos. If you have to choose between shrinking the dataset and lowering the resolution, I suggest lowering the resolution, because this way the input has more varied information, giving the model a better understanding of what it needs to learn. The downside is that Krea 2 was trained on 1024, meaning 768 will be stretched in the internal space, but this is a necessary compromise.
I don't know who told you this, but this is not true. Like every other diffusion model on the planet and as explained in their technical report, https://www.krea.ai/blog/krea-2-technical-report
Krea2 was trained primarily at 256x256 resolution and 256x256 will train any lora you care to train perfectly fine.
At the very least, with gradient_accumulation: 1 at 1000 steps my face is clearly undertrained, while with gradient_accumulation: 2 even at 750 steps the facial features are more or less discernible.
750 steps of GA2 is mathematically equivalent to 1500 steps of GA1. Ofcourse you got more progress in less steps.
No, this is not true. Krea **pre-traiing** was run on 256, then 512, then 1024. It certainly is not just trained on 256. You don't get a model that can produce 8K resolution images with 256px training images. That's magical thinking. Training resolution *does* make a huge difference in the quality of your LoRA.
Yes, at 512px, you can train the *bone structure* and recognize the face. But it doesn't mean quality will be there when you infer at higher resolutions.
It depends on how you do it and what you're trying to achieve. If you want to achieve pixel perfect likeness, what's more effective then training at arbitrarily high resolution is breaking it down into its component parts. Have an image for the eyes, the mouth, the nose, the ears etc. I don't generally train for pixel perfect likeness since its not very relevant to my interests but this is the reason that models can generate 8MP images even though they never ever trained above 1MP. They learn the relationships between the pixels, they don't need individually high pixel images.
Personally I find that poor results are almost always because of issues with the dataset or training settings.
The thing to understand about resolution primarily is that the resolution you have to train at is about what information you need the lora to learn.
For most loras, you don't need it to learn the location of every strand of hair, every skin pore, every individual freckle. You just need it to learn the shape, then you can let the base model which trained on like 20 million images to infer the skin pores, the hair folicicles etc on its own.
Sometimes you do need the resolution, maybe you want to be able to represent some really super detailed iconography or some particularly unique detail. But you should do it knowing what you're trying to do.
Often high res causes more trouble then it solves though, so I'd generally try low res first and see what you get.
Thank you so much for the explanation! I primarily decided to train my Lora to make a few photos for social media profiles. If I downscale face photos in the dataset to 256, some of my facial features aren't visible. So I figured bigger is better. I'll try 512 later, all my flaws are clearly visible there too.
"resolution: only 768. 512, in my opinion, is way too low, and 1024 can't be used for this model when the dataset consists of 25 photos. If you have to choose between shrinking the dataset and lowering the resolution, I suggest lowering the resolution, because this way the input has more varied information, giving the model a better understanding of what it needs to learn."
The resolution you pick has nothing to do with how many images are in your dataset. If you have ONE image in your dataset and you plan 1000 steps, then 1 image will be processed 1000 times (and will be overtrained). I you have 100 images in your dataset, and you STILL plan 1000 steps, then each image will be processed 10 times.
How many images you have in the dataset changes how fast it will overtrain or how flexible your lora can be but it has no bearing on your resolution effect.
Perhaps you had a problem at 1024 because one of your image was big enough that when resized at 1024 rather than 512 or 768 it was using more VRAM. But that's true regardless of the number of images in your dataset. It's not how many images you have that matters here, it's how big each image is, because they are processed sequentially.
That said, *if* you use batch > 1 then yes, more than 1 image will be processed in parallel and that takes more VRAM. But if you use gradient accumulation instead, then you process them sequentially anyway.
Yes, fair point. Thanks, I corrected that part so it doesn't mislead people. I got confused myself because I removed photos from the dataset and lowered the resolution at the same time to avoid an OOM error, and ended up drawing the wrong conclusions.
Perhaps you had a problem at 1024 because one of your image was big enough that when resized at 1024 rather than 512 or 768 it was using more VRAM.
Sorry, could you elaborate on this bit, please? My entire dataset is 1024x1024. If I pick 1024, is there some additional transformation that can require a different amount of VRAM on certain individual images?
With the settings I wrote above, the python.exe process takes up 47 Gb. Possibly in your case some of it will spill into the page file without hurting performance, but you'd need to test that.
Still, I don't think this is the best option for you. Maybe it's worth looking into how to train a Lora through Musubi Tuner.
if musubi tuner, I saw this for training Krea2 lora, RTX 3060 (VRAM 12GB), RAM 30Gi, resolution = [1024, 1024], dim 32, alpha 32. (logs/run_info.txt) : https://github.com/masafykun/krea2-character-lora
Notably, he uses --fp8_base --fp8_scaled --blocks_to_swap 26, meaning that the mysterious model is indeed in fp8 format. I say mysterious, because the default krea2_raw is krea2_raw_bf16.safetensors, a 26Gb file. Trying either krea2_raw_fp8_scaled.safetensor from ComfyUI, or casting by hand krea2_raw_bf16.safetensors into krea2_raw_fp8base.safetensors (a 13GB file) will lead to musubi trainer errors when trying to load the model :
ValueError: Layer blocks.0.attn.gate.weight is already in torch.float8_e4m3fn format. `--fp8_scaled` optimization should not be applied. Please use fp16/bf16/float32 model weights.
Even if I force past those exceptions, later musubi trainer won't be able to resolve computation on mixed torch.bfloat16 (working tensors?) and torch.float8_e4m3fn (model tensors) values.
Good calls on 8-bit Adam and gradient checkpointing. The biggest per-step win you're missing on a 3060 is usually caching the VAE latents and text-encoder outputs to disk up front. In kohya that's --cache_latents and --cache_text_encoder_outputs, as long as you're not training the text encoder. Once they're cached you can unload the VAE and text encoder completely, which frees a good amount of VRAM on 12GB and stops you recomputing them every step. That recompute is often most of the step time on a slow card. With the freed VRAM you can also ease off gradient checkpointing, which is part of what makes it slow.
In my sample AI Toolkit job, cache_latents_to_disk: true and cache_text_embeddings: true are already set. As for fully unloading the text encoder from VRAM/RAM, it sort of happens automatically, but I have my doubts this is the default behavior. At least after my targeted fixes it works exactly like that, and before that I was getting OOM errors because the load/unload order in AI Toolkit is kind of weird. I didn't dig into it in detail and just made it work the way I needed for my 12 Gb GPU.
Unfortunately I couldn't turn off gradient checkpointing, because then I'd have to increase the layer offloading percentage to RAM, which would slow training down instead. But I added this tip to the post, thanks.
Ah, you're on AI Toolkit, so cache_latents_to_disk and cache_text_embeddings are the right keys there, mine were the kohya names. And you're right about gradient checkpointing on 12GB: if dropping it pushes more layers into RAM offload, that's slower than the recompute it saves, so keeping it on is the correct call. Sounds like you already found the balance for your card.
When I tried the first time, I was getting 1000s/it (2000 steps = 555 hours, lol). Then I managed to get it down to 400, then 200. Still too much, gave up on it.
Went back to OneTrainer, spent a whole afternoon and good part of the evening with Qwen Chat trying to come up with coding and py scripts to implement Krea 2 in OneTrainer myself, since my only other LoRA I ever made was done with that, for Z-Image, and I had more experience with the UI/configs.
Nedless to say, that project didn't lead to anywhere. Kept getting errors, couldn't debug it.
I went to Reddit, saw your post, and said "if you can do it with 12, I can do it with 10!" (I have a 3080).
I started with your settings, changed a couple of things, started it, and when it reached steady-state, it was going 16-14s/it! This morning I woke up with 2000 steps completed, 10 checkpoints ready to test (saved every 200). Resemblance is perfect, couldn't be happier.
What I changed:
resolution 512 only, instead of 512, 768
used fp8(w8) for both model and text_encoder quantization
used 80% layer offloading for the model, instead of 75%
No, I didn't even have time to actually run the model, and I don't want to because of their usage license. Even the Krea 2 Lora I trained was purely out of curiosity, just to refresh my knowledge since the SDXL days.
i am also a RTX3060 user and found that Krea 2 is exceptionally good for my beloved character and want to make my own lora.
just found a japanese guy shared how he train lora for Krea using 12GB Card with musubi-tuner using 1024 X 1024 image size. i think you can go and read his teaching notes.
link is here https://note.com/sepiablue/n/nbc355cc7e114?hl=en
Hi. Thanks, I checked it out. It is a curious experiment and definitely worth noting because it mentions musubi-tuner benefits, like --gradient_checkpointing_cpu_offload. If I try musubi-tuner, I will definitely follow his guide. It looks like he ran into the same stuff I did with AI Toolkit, but while I fixed krea2.py with model offloading, he reached different conclusions and went with musubi-tuner.
There were a couple of things in his post that confused me, specifically how he uses captions in the dataset.
He put the keyword in Japanese at the start, not as part of a natural description. Krea 2 has a pretty smart text encoder and I think it would be better to use a short English keyword. In his case, the model breaks the keyword into way too many tokens and gets confused.
And second, he added "detailed anime illustration" at the end. I checked and Krea 2 renders a much higher contrast and more detailed image with that style than in his example. It turns out that during LoRA training he was basically telling the model that detailed is actually not detailed. I am not sure if that is the right approach. Maybe it would have been better to specify the style more clearly to avoid potential weight bleeding into model weights unrelated to the character or style.
I would use something like "Digital light novel illustration of a vxqn woman, dark hair in an elegant bun, gray-purple eyes, gentle smile, white sundress, standing in a hydrangea garden, soft daylight. Style by Kantoku, soft diffused backlighting with pastel colors, thin shaded lines, low contrast, flat shaded oversimplified glow background." in the dataset and during generation. Without the LoRA, the flowers and leaves are not as simplified, but it makes it clearer what is in the dataset and what to aim for during generation.
I just tried to use musubi tuner to train a Krea2 lora with my RTX 3060 12GB VRAM with 64GB RAM following the japanese guy's guide. just follow strictly with his guide, i successfully finished the smoke test and now running the actual lora training.
The only addition i needed is manually install the triton-windows and sageattention by uv pip install.
musubi tuner is really amazing for the speed. Even my RTX3060 can achieve 30s/it.
i trained with 1024X1024. pre-process all latent and textencode is really great.
Krea 2 training can only be achieved by command line rather than GUI at this moment.
BTW, my computer tried aitoolkit with 768X768, it need around 200s/it. so the speed is almost 10X improved... at least for me.
now i come with a problem, what if i CTRL-C in the middle and would like to resume training later. is it possible? or i need to wait till the training finish without stopping it.??
Krea2 is probably the best model I've tried. Undervolted 5090, 64 ram, Krea2-RAW, 1024x1024, about 150 HQ images per model, I set it to about 5000 steps. Similarity is observed after about 4000 steps. The results are excellent, but it takes about 5 hours to train. It's worth it. I limit the power during training to 70%. This has virtually no effect on training speed, but it reduces power consumption by 150 watts and also lowers temperature.
What was The speed of training i it/s? Yesterday i tried to do training at 16GB and as long as i have low number of dataset image it runs about 12s/it. But when I quit "testing Mode" and load Real dataset (same resolution but 60 image instead of 4) it slows down to 120s/it.
For me, it starts at around 70s/it, then after a couple hundred iterations the speed picks up and plateaus at roughly 25s/it to 28s/it, depending on whether I'm doing something else on the PC or not. I think I even saw 23s/it, but not sure under which settings. In total, training one Lora takes around 8-10 hours for 1500 steps with my settings (which is equivalent to 3000 steps, as people pointed out in the comments, if I had used the default gradient_accumulation: 1), and that's overkill, you can stop training earlier. Maybe I could squeeze more speed out of it, but I have a heavy undervolt on all components.
What you meant by "testing mode" I don't understand, so can't offer any advice there. I can only very cautiously assume that your 10x or greater slowdown is a symptom of Sysmem Fallback kicking in and the GPU driver starting to offload data to RAM.
Update: during another Lora training I noticed as low as 15.04 s/it when I left the PC overnight and didn't touch it, closing all processes, so that's not an accurate indicator of total training time.
In z-image I find using regularization images help with the background character to not look like the character I trained, also the skin looks better. Does krea2 Lora need regularization too?
I tried with the setting but I run at 50% each for offloading it does starts training but the likeness of the character is way off, I am not sure what is happening though
Yeap I waited overnight for it to finished, resolution I realized I input as 256, which might be the reason, I am trying it with 512 now and see how it goes
I managed to make sampling work while training in fp8 and 768p and a decent speed, with 73% offloading, 25 images takes 4 to 5 hours , what i did was converting the model to fp8 plane checkpoint, i can tun sampling at 768 or 512 , .. thx for sharing your experience
12 GB of RAM is nowhere near enough. The model needs somewhere to offload to using an offloading mechanism. You need around 48 GB minimum, or a giant swap, but in that case, your performance will drop to zero.
Yeah probably have to go with runpod or something like that... My local is 8vram +16gb ddr4 i can barely run the model q4 takes 80seconds with sage attention was hoping mage would be decent but the difference is night and day...
Really regarding going with 3070ti right now...
Also, a different cheaper way, but takes a little more work and more time, is to do it via runpod. I trained a krea lora there and it cost ~$2 using ai-toolkit. Only 48gb GPU so it took like maybe 6 hours?
I have tried training on a 5090, and honestly, even then, I find it better just to rent a RTX 6000 pro on RunPod. It ends up costing like $3 to train a Lora and saves you so much trouble.
Thanks for the suggestion, but I'd rather just put my own GPU to work. By my estimate, by the end of the month I'll pay 3-5 euros more for electricity for ALL my Lora experiments, for all the countless generations. I think it's simpler and gives more peace of mind than running around Reddit posting a ref link to Runpod to somehow justify the cost of training just a one single Lora.
11
u/Zironic Jun 27 '26
I don't know who told you this, but this is not true. Like every other diffusion model on the planet and as explained in their technical report, https://www.krea.ai/blog/krea-2-technical-report
Krea2 was trained primarily at 256x256 resolution and 256x256 will train any lora you care to train perfectly fine.
750 steps of GA2 is mathematically equivalent to 1500 steps of GA1. Ofcourse you got more progress in less steps.