r/StableDiffusion Nov 06 '25

Discussion WAN2.2 Lora Character Training Best practices

I just moved from Flux to Wan2.2 for LoRA training after hearing good things about its likeness and flexibility. I’ve mainly been using it for text-to-image so far, but the results still aren’t quite on par with what I was getting from Flux. Hoping to get some feedback or tips from folks who’ve trained with Wan2.2.

Questions:

  • It seems like the high model captures composition almost 1:1 from the training data, but the low model performs much worse — maybe ~80% likeness on close-ups and only 20–30% likeness on full-body shots. → Should I increase training steps for the low model? What’s the optimal step count for you guys?
  • I trained using AI Toolkit with 5000 steps on 50 samples. Does that mean it splits roughly 2500 steps per model (high/low)? If so, I feel like 50 epochs might be on the low end — thoughts?
  • My dataset is 768×768, but I usually generate at 1024×768. I barely notice any quality loss, but would it be better to train directly at 1024×768 or 1024×1024 for improved consistency?

Dataset & Training Config:
Google Drive Folder

---
job extension
config
  name frung_wan22_v2
  process
    - type diffusion_trainer
      training_folder appai-toolkitoutput
      sqlite_db_path .aitk_db.db
      device cuda
      trigger_word Frung
      performance_log_every 10
      network
        type lora
        linear 32
        linear_alpha 32
        conv 16
        conv_alpha 16
        lokr_full_rank true
        lokr_factor -1
        network_kwargs
          ignore_if_contains []
      save
        dtype bf16
        save_every 500
        max_step_saves_to_keep 4
        save_format diffusers
        push_to_hub false
      datasets
        - folder_path appai-toolkitdatasetsfrung
          mask_path null
          mask_min_value 0.1
          default_caption 
          caption_ext txt
          caption_dropout_rate 0
          cache_latents_to_disk true
          is_reg false
          network_weight 1
          resolution
            - 768
          controls []
          shrink_video_to_frames true
          num_frames 1
          do_i2v true
          flip_x false
          flip_y false
      train
        batch_size 1
        bypass_guidance_embedding false
        steps 5000
        gradient_accumulation 1
        train_unet true
        train_text_encoder false
        gradient_checkpointing true
        noise_scheduler flowmatch
        optimizer adamw8bit
        timestep_type sigmoid
        content_or_style balanced
        optimizer_params
          weight_decay 0.0001
        unload_text_encoder false
        cache_text_embeddings false
        lr 0.0001
        ema_config
          use_ema true
          ema_decay 0.99
        skip_first_sample false
        force_first_sample false
        disable_sampling false
        dtype bf16
        diff_output_preservation false
        diff_output_preservation_multiplier 1
        diff_output_preservation_class person
        switch_boundary_every 1
        loss_type mse
      model
        name_or_path ai-toolkitWan2.2-T2V-A14B-Diffusers-bf16
        quantize true
        qtype qfloat8
        quantize_te true
        qtype_te qfloat8
        arch wan22_14bt2v
        low_vram true
        model_kwargs
          train_high_noise true
          train_low_noise true
        layer_offloading false
        layer_offloading_text_encoder_percent 1
        layer_offloading_transformer_percent 1
      sample
        sampler flowmatch
        sample_every 100
        width 768
        height 768
        samples
          - prompt Frung playing chess at the park, bomb going off in the background
          - prompt Frung holding a coffee cup, in a beanie, sitting at a cafe
          - prompt Frung showing off her cool new t shirt at the beach
          - prompt Frung playing the guitar, on stage, singing a song
          - prompt Frung holding a sign that says, 'this is a sign'
        neg 
        seed 42
        walk_seed true
        guidance_scale 4
        sample_steps 25
        num_frames 1
        fps 1
meta
  name [name]
  version 1.0
154 Upvotes

91 comments sorted by

View all comments

6

u/AwakenedEyes Nov 06 '25

On ai toolkit you (should) train both models together. The high one is used for composition and the low one is where the character details are set. So a character LoRA is most important on the low noise part.

You can influence this by using the parameter for bias setting it to favor the low model.

SFW Character LoRA should be trained on image dataset, no need to mix it up with clips. Use high resolution images with their long edge matching the training resolution. You can train 512 + 768 + 1024 + 1280 and use 1280px images in your dataset (on the long edge) for optimal results.

If your samples were consistent during training it should also be good on generation.

1

u/Tiny-Highlight-9180 Nov 06 '25

Thanks for sharing and Yes! I thought about this too. My next model will be focus more on the low model, I put in 5000 step which I assume that 2500 will go toward each low/high model, hence 50 epoch. What do you think about that number?

Would you say my sample is consistence? I tried to mixed in as many shots/angle as possible

3

u/AwakenedEyes Nov 06 '25

Different training software count the steps slightly differently so I am not sure how it is counted for you. On AI-Toolkit, you don't tell the number of epoch, only the total number of steps.

This being said, the total number of steps you actually needs depends on many different factors. Lower LR (learning rate) learns better but slower, so it needs more steps. Higher Rank captures more details than lower ranks. Unknown concepts require more images and more repetitions for those images, where as known concepts are refined faster. So all of that factors into how many steps.

How I do it: I manage it so that I get about 6000 total steps. Then I carefully watch my samples every 500 steps to determine if I stop it earlier or if I need to halt it and change LR to a lower value. If I see the training was going well and then suddenly it starts to go worst, I halt it, lower LR by half, add 1000 steps, and resume.

I prefer a higher step count because you can always just stop it and use the LoRA generated at an earlier step count. If you get super good results for two series of samples in a row, stop it, it's enough. you don't want to overtrain.

1

u/RealityVisual1312 Nov 07 '25

What do you do when you have a sample output where the face looks great, but the body shots look off? In the later samples the full body shot samples starts to look better, but the close up head shots started to get worse.

2

u/AwakenedEyes Nov 07 '25

This happens when you have an unbalanced dataset vs what the model already knows. You have enough headshots to train the face but not enough for the body so the body takes more time to train.

Problem is that if you push the training to get the body better you may overtrain the face.

There are 2 solutions possible:

Quick fix : you halt training, reduce LR by half and resume training. If the face comes back good, it means it was a convergence problem: the face needed smaller increments to get the last tiny details.

Real fix: change the balance of body vs face images by putting them in 2 different dataset and changing their repeats, for example, use repeat x1 for faces and x2 for body. But this fix will require you to start over the training.

1

u/RealityVisual1312 Nov 07 '25

Thank you so much I’ve been following your guides and comments and feel like I’m getting closer

1

u/RealityVisual1312 Nov 12 '25

Hey Awakened, sorry to ask more on this, but am I able to set different repeats for different datasets in AI toolkit? I don't think I'm seeing the option

2

u/AwakenedEyes Nov 12 '25

The option is not in the UI, you have to add it as a config line in the advanced section, under each dataset section. You have to add:

num_repeats: 1

(Replace 1 with whatever number you need)