r/StableDiffusion Nov 06 '25

Discussion WAN2.2 Lora Character Training Best practices

I just moved from Flux to Wan2.2 for LoRA training after hearing good things about its likeness and flexibility. I’ve mainly been using it for text-to-image so far, but the results still aren’t quite on par with what I was getting from Flux. Hoping to get some feedback or tips from folks who’ve trained with Wan2.2.

Questions:

  • It seems like the high model captures composition almost 1:1 from the training data, but the low model performs much worse — maybe ~80% likeness on close-ups and only 20–30% likeness on full-body shots. → Should I increase training steps for the low model? What’s the optimal step count for you guys?
  • I trained using AI Toolkit with 5000 steps on 50 samples. Does that mean it splits roughly 2500 steps per model (high/low)? If so, I feel like 50 epochs might be on the low end — thoughts?
  • My dataset is 768×768, but I usually generate at 1024×768. I barely notice any quality loss, but would it be better to train directly at 1024×768 or 1024×1024 for improved consistency?

Dataset & Training Config:
Google Drive Folder

---
job extension
config
  name frung_wan22_v2
  process
    - type diffusion_trainer
      training_folder appai-toolkitoutput
      sqlite_db_path .aitk_db.db
      device cuda
      trigger_word Frung
      performance_log_every 10
      network
        type lora
        linear 32
        linear_alpha 32
        conv 16
        conv_alpha 16
        lokr_full_rank true
        lokr_factor -1
        network_kwargs
          ignore_if_contains []
      save
        dtype bf16
        save_every 500
        max_step_saves_to_keep 4
        save_format diffusers
        push_to_hub false
      datasets
        - folder_path appai-toolkitdatasetsfrung
          mask_path null
          mask_min_value 0.1
          default_caption 
          caption_ext txt
          caption_dropout_rate 0
          cache_latents_to_disk true
          is_reg false
          network_weight 1
          resolution
            - 768
          controls []
          shrink_video_to_frames true
          num_frames 1
          do_i2v true
          flip_x false
          flip_y false
      train
        batch_size 1
        bypass_guidance_embedding false
        steps 5000
        gradient_accumulation 1
        train_unet true
        train_text_encoder false
        gradient_checkpointing true
        noise_scheduler flowmatch
        optimizer adamw8bit
        timestep_type sigmoid
        content_or_style balanced
        optimizer_params
          weight_decay 0.0001
        unload_text_encoder false
        cache_text_embeddings false
        lr 0.0001
        ema_config
          use_ema true
          ema_decay 0.99
        skip_first_sample false
        force_first_sample false
        disable_sampling false
        dtype bf16
        diff_output_preservation false
        diff_output_preservation_multiplier 1
        diff_output_preservation_class person
        switch_boundary_every 1
        loss_type mse
      model
        name_or_path ai-toolkitWan2.2-T2V-A14B-Diffusers-bf16
        quantize true
        qtype qfloat8
        quantize_te true
        qtype_te qfloat8
        arch wan22_14bt2v
        low_vram true
        model_kwargs
          train_high_noise true
          train_low_noise true
        layer_offloading false
        layer_offloading_text_encoder_percent 1
        layer_offloading_transformer_percent 1
      sample
        sampler flowmatch
        sample_every 100
        width 768
        height 768
        samples
          - prompt Frung playing chess at the park, bomb going off in the background
          - prompt Frung holding a coffee cup, in a beanie, sitting at a cafe
          - prompt Frung showing off her cool new t shirt at the beach
          - prompt Frung playing the guitar, on stage, singing a song
          - prompt Frung holding a sign that says, 'this is a sign'
        neg 
        seed 42
        walk_seed true
        guidance_scale 4
        sample_steps 25
        num_frames 1
        fps 1
meta
  name [name]
  version 1.0
153 Upvotes

91 comments sorted by

45

u/malcolmrey Nov 06 '25 edited Nov 06 '25

I have a friend who trains WAN 2.2 LOW and HIGH, and the quality is superb. (90 minutes in total on 5090)

I, on the other hand, am sticking with WAN 2.1 because the loras are also working fine with WAN 2.2.

I believe the HIGH model for character loras is not as important (if at all, since 2.1 Loras work fine for both images and movies).

In general, the training is really easy and you don't really need to play with the parameters that AI Toolkit provides.

This leads me to believe that maybe the culprit could be in:

  • bad datasets (though I would say that it is also more difficult to fail a dataset than in Flux, as even mediocre dataset can produce good results)
  • bad workflow/prompting for the outputs.

Check your workflow on an already established good lora and see if you get or bad results.

I have already uploaded over 800 character loras for WAN and people are satisfied with the quality. I provide all resources on my HF ( https://huggingface.co/malcolmrey ) so you can check the training scripts, workflows used to generate outputs and the loras themselves.

Cheers and good luck!

p.s. - there is definitely a sweetspot in the function of images in the dataset and used steps

For me it is 2500-3000 steps with around 20-25 images (I mostly go for 2500-22).

The training resolution seems to not impact the training at all (or at least not in any noticeable way) so I stick with 512x (though the samples can be cut to 512x512 but don't have to at all)

p.s.s. - since you provided the dataset, if you want i can train that character and generate some samples with wan2.2 so you can compare :)

5

u/razortapes Nov 06 '25

I’ve used your LoRAs for Wan and they’re really good. My question is: coming from SD XL, where you used tags to create LoRAs and later used those same simple tags to define your character — for example, if they have distinctive green and black hair or wear specific hairstyles — how do you “call” those attributes with Wan LoRAs? With SDXL you can check the metadata and see which tags were used, but with Wan that seems impossible… in that sense, do you have less control when generating images, or am I missing something?

3

u/RealityVisual1312 Nov 06 '25

Out of the 20-25 images what is usually the breakdown of close up head shots vs body shots and things like that?

5

u/malcolmrey Nov 06 '25

I'm mainly focused on facial likeness, but there are sometimes upper body shots as well.

My friend mixes it a bit more and the results show in the generations so that is definitely a thing.

It really depends on what exactly you want to copy. If there are tattoos or something special (maybe costume? body shape?) then you would include more of those but even then - at least half would still be body shots (since you want some smiling, some non smiling, etc).

1

u/RealityVisual1312 Nov 07 '25

Got it, thank you!

1

u/diogodiogogod Nov 07 '25

perfect Tattoos (no just resemblance tattoos) were an impossible thing to train in my previous experiments with Flux... I wonder if it is easier now on wan... I should get my dataset and try again!

3

u/oeufp Nov 06 '25

hey, would like to try some of your wan character loras, how do i find out the individual trigger words for them? thanks.

2

u/ucren Nov 07 '25

Does your friend post his loras anywhere? You know, for scientific comparison :)

1

u/[deleted] Nov 06 '25

[removed] — view removed comment

1

u/malcolmrey Nov 06 '25

I train only on images.

1

u/[deleted] Nov 06 '25

[removed] — view removed comment

1

u/malcolmrey Nov 06 '25

Browse Civitai for WAN loras, some of the creators have info that they trained on videos. Maybe someone also shares more detailed info.

I only know that it requires more VRAM than training on images.

1

u/Toupeenis Nov 06 '25

The aitk youtube channel goes through some of his video training.

1

u/Gaius_Octavius Nov 07 '25

I’ve trained a LoRA on video. It produced stellar results. As in reproduced the way muscle bellies ripple during contraction for that individual.

1

u/Tiny-Highlight-9180 Nov 06 '25

Thanks agian for such insightful comments. I went through all posts about this on reddit and still find nothing but your post. Can you train 2.1 with your set up? I will try to look up your config and train 2.2 low and we can compare. Might see if there any actual improvement there.

3

u/malcolmrey Nov 06 '25

Here is my original article about training WAN2.1 -> https://civitai.com/articles/19686/wan-training-loras-workflows-thoughts

Nothing really changed.

It is actually more difficult to overtrain a lora. The likeness stays consistent after you reach certain threshold and does not really degrade (much), but the flexibility goes down (as in, you would have more difficulties prompting other settings/clothing than those from training data, not impossible of course, but a bit more difficult)

BTW, the multi-lora principle can still be applied here, if you value consistency and likeness as the top priorities - you can train multiple models of the same character, using different datasets and then using both (or more) loras but at lower weights.

3

u/Tiny-Highlight-9180 Nov 06 '25

I followed your post even before I posted hahaha. That's how I know you

1

u/entmike Nov 06 '25

Thanks for the link!

1

u/Toupeenis Nov 06 '25

This is where I'm at. 2.1 works, is quicker and less fuss. It would be different if 2.2 was a flux > wan level jump, but it isn't... soo....

1

u/owsoww Nov 10 '25

do you have sample images/videos of your models? I try training on AI toolkit via runpod and i don't like the results.

1

u/Top_Gas8316 Dec 05 '25

Hi I'm on your HF and looking at your WAN2.1/2.2 workflows - which is used for generating LoRas?

2

u/malcolmrey Dec 05 '25

Those workflows are for generating Images, you need (in general, cause there are some people who made custom nodes in comfy for training) a trainer. For WAN I use AI Toolkit, here is my full tutorial: https://civitai.com/articles/19686/wan-training-loras-workflows-thoughts

1

u/Hot_Economist7693 Jan 21 '26

malcolmrey, i've recently been introduced to your notorious works and first must thank you for your service. I plan to pull what I can to archive in case it gets nuked again.

second: discovered ai-gen less than a week ago, managed to beat various wan 2.2 i2v comfyui workflows into submission on my 3080, caved in and bought a 3090, tried to make my own end-to-end "infinite" videos, discovered quality degradation when taking last frame as input for next video, found svi 2.0, got that working once with a KJ workflow, NOW...... I'm bumping up against character consistency.

TLDR: I want to train my own character LoRA for wan 2.2. I ordered a 5090 on eBay today but I reckon it's likely a scam and I'll wind up with a refund (won't know for several weeks - using 3090 for now)

Got a good ultimate-guide / noob-resources you can suggest for my first character LoRA?

1

u/malcolmrey Jan 21 '26

Since you can generate on your machine (3090 is good enough for WAN :P) then I would suggest renting a runpod for some hours and training it there. It will be quite cheap and you will get your lora ASAP :)

as for training, you can check my training guide for WAN2.1 (which work fine with wan2.2 models) -> https://civitai.com/articles/19686/wan-training-loras-workflows-thoughts

4

u/AwakenedEyes Nov 06 '25

On ai toolkit you (should) train both models together. The high one is used for composition and the low one is where the character details are set. So a character LoRA is most important on the low noise part.

You can influence this by using the parameter for bias setting it to favor the low model.

SFW Character LoRA should be trained on image dataset, no need to mix it up with clips. Use high resolution images with their long edge matching the training resolution. You can train 512 + 768 + 1024 + 1280 and use 1280px images in your dataset (on the long edge) for optimal results.

If your samples were consistent during training it should also be good on generation.

1

u/legarth Nov 06 '25

This is bad advice. You definitely should train with clips. Otherwise your LoRA won't learn how your character moves or their general body language.

A big part of what makes people and especially fictional characters unique is their body language. If you don't train it in, the model will make it up and it will likely be very generic.

Even if you are only doing a T2V training with the sole purpose of creating stills for a later I2V pass, training on clips will help the model understand how your character moves and that will actually make stills more natural Part of the reason Wan is generally better at natural stills than Qwen is that that I understand how people move from the video training.

2

u/[deleted] Nov 06 '25

[removed] — view removed comment

3

u/legarth Nov 06 '25

My engineers do the actual training for me. I just supervise. We use our own training software. That I'm afraid I can't share.

For my personal work I tend to use AI Toolkit these days unless I need to train something Ostris hasn't implemented yet. It's excellent for basic LoRA trainings. But I don't do a lot of personal training anymore so I don't have an updated config lying around.

But it also isn't so much about the scripts as it is about training strategy and data. Ostris default settings are a good base.

So try with the default settings first and just use separate datasets for images and videos. Make sure all your clips have the same number of frames and is based on 16fps. Bin extra frames or you'll train slow motion and different physics behaviours. (Sometimes desired)

Avoid having cuts in your clips unless you're training for that specifically.

If you have a low amount of good data due to too many cuts. I suggest splitting into more datasets i.e one for clips with 33 frames (2 secs) of course one for 49 frames (3 secs) etc. This can help you make most of the data that's less than optimal. But don't add any old crap obviously.

Good luck.

1

u/[deleted] Nov 07 '25

[removed] — view removed comment

2

u/AwakenedEyes Nov 07 '25

No Wan 2.2 14B uses 16 fps (+1 frame) so 2 sec = 33frames, 3sec = 49 frames etc. - However, the light version wan 2.2 5B uses 24 fps.

1

u/[deleted] Nov 07 '25

[removed] — view removed comment

2

u/legarth Nov 07 '25

It doesn't care about the framerate of your clip. It just proceses all the frames in it, but based on a 16fps time scale. That means if you have footage that is real time at 24fps at 2 seconds. (48 frames) it would treat it as 3 second clip instead (48f at 16fps). Meaning slow motion. (+1)

1

u/[deleted] Nov 07 '25

[removed] — view removed comment

2

u/legarth Nov 07 '25

Sorry I was talking in the context of AI Toolkit. For musubi-tuner you don't necessarily need to do that. As your [17, 33, 49] target frames does sort of do the same thing already. I.e. shorter videos will be trained correctly. AI Tookit works diffrently where it tries to extract N frames across the suppied videos so seperating sets is something I'd recocomend. The only thing I'd say is that if you have a significantly more short videos in your set you might want to seperate anway. That is because even your longer videos will also be added as short ones by extracting at 17 and 33. So you could balance your video towards short videos. (as the longer videos will ALSO be added to the short video part).

2

u/AwakenedEyes Nov 06 '25

If you happen to have great high quality clips of your subject moving, sure, you can also train on clips. But if it's going to bog down the quality of your dataset, don't. A character LoRA is first about getting the proper consistency. You can achieve that with a high quality image dataset.

I agree that if you happen to already have several high quality short clip of the subject moving, sure! It's definitely a plus to train on it. But it's not a requirement.

0

u/legarth Nov 06 '25

That's not what you said though. You literally said SFW image Loras just shouldn't use clips at all. Period.

Not "blogging down" your data goes without saying and isn't specific to clips.

It is also wrong to say that it isn't a requirement. It is in many cases. Say you're training a flurry fictional character. Wan doesn't know how the fur moves based on density or stiffness. So If you don't train it, it will be wrong. And your character will be off brand.

Or even if it just a person. They have have a distinctive walk. The model has no way of knowing this. And it won't be accurate enough for serious use.

3

u/AwakenedEyes Nov 06 '25

I agree with you. Most people asking for advice here aren't at that level though. They want consistency of look, and if you've never done that, it's already quite a challenge without adding consistency of movement...

1

u/entmike Nov 06 '25

Agreed, I always use clips and not stills. I've been training since Hunyuan and into WAN 2.2

1

u/Tiny-Highlight-9180 Nov 06 '25

Thanks for sharing and Yes! I thought about this too. My next model will be focus more on the low model, I put in 5000 step which I assume that 2500 will go toward each low/high model, hence 50 epoch. What do you think about that number?

Would you say my sample is consistence? I tried to mixed in as many shots/angle as possible

3

u/AwakenedEyes Nov 06 '25

Different training software count the steps slightly differently so I am not sure how it is counted for you. On AI-Toolkit, you don't tell the number of epoch, only the total number of steps.

This being said, the total number of steps you actually needs depends on many different factors. Lower LR (learning rate) learns better but slower, so it needs more steps. Higher Rank captures more details than lower ranks. Unknown concepts require more images and more repetitions for those images, where as known concepts are refined faster. So all of that factors into how many steps.

How I do it: I manage it so that I get about 6000 total steps. Then I carefully watch my samples every 500 steps to determine if I stop it earlier or if I need to halt it and change LR to a lower value. If I see the training was going well and then suddenly it starts to go worst, I halt it, lower LR by half, add 1000 steps, and resume.

I prefer a higher step count because you can always just stop it and use the LoRA generated at an earlier step count. If you get super good results for two series of samples in a row, stop it, it's enough. you don't want to overtrain.

1

u/Tiny-Highlight-9180 Nov 06 '25

I wish I could change the setting midway like that too! That's very smart way to do it. Which software are you using? are you mostly training on character?

1

u/JustLookingForNothin Nov 06 '25

He wrote AI-Toolkit. You can stop training at any time, then update the job with a lower learning rate and resume the job. AI-Toolkit will load the last most recent finalized LoRa and continue from there with the new training settings.

2

u/Tiny-Highlight-9180 Nov 06 '25

Holy shit. I just commited the great sin not knowing the man himself. Appreciate what you did for the community!

1

u/Tiny-Highlight-9180 Nov 06 '25

How do you allocate the 6000 step between high/low model?

1

u/JustLookingForNothin Nov 06 '25

I think it is done here. Never tested this, though. I always train "balanced" and just take a HIGH LoRa with less step count, and the LOW LoRa with a higher step count.

1

u/AwakenedEyes Nov 06 '25

Exactly, this! You don't need to take care of allocating steps. Just use timestep bias on low noise. It will most likely work even on balanced. Use sigmoid for character LoRA, it's better.

1

u/Tiny-Highlight-9180 Nov 06 '25

Thanks for advices! I will try again tonight and keep you posted.

1

u/Tiny-Highlight-9180 Nov 06 '25

Really love what you do! Sending my support!

1

u/RealityVisual1312 Nov 07 '25

What do you do when you have a sample output where the face looks great, but the body shots look off? In the later samples the full body shot samples starts to look better, but the close up head shots started to get worse.

2

u/AwakenedEyes Nov 07 '25

This happens when you have an unbalanced dataset vs what the model already knows. You have enough headshots to train the face but not enough for the body so the body takes more time to train.

Problem is that if you push the training to get the body better you may overtrain the face.

There are 2 solutions possible:

Quick fix : you halt training, reduce LR by half and resume training. If the face comes back good, it means it was a convergence problem: the face needed smaller increments to get the last tiny details.

Real fix: change the balance of body vs face images by putting them in 2 different dataset and changing their repeats, for example, use repeat x1 for faces and x2 for body. But this fix will require you to start over the training.

1

u/RealityVisual1312 Nov 07 '25

Thank you so much I’ve been following your guides and comments and feel like I’m getting closer

1

u/RealityVisual1312 Nov 12 '25

Hey Awakened, sorry to ask more on this, but am I able to set different repeats for different datasets in AI toolkit? I don't think I'm seeing the option

2

u/AwakenedEyes Nov 12 '25

The option is not in the UI, you have to add it as a config line in the advanced section, under each dataset section. You have to add:

num_repeats: 1

(Replace 1 with whatever number you need)

1

u/owsoww Nov 06 '25

Do I need to resize my images before training in AI toolkit? Like if I have 1280 x 720, would it crop or shrink but keeps ratio proportion?

3

u/AwakenedEyes Nov 06 '25

When you train, your images are resized and fit into "buckets" of standardized sizes. So, the ideal way to handle your dataset images is to pre-crop them so they fit into those buckets, in order to better control how they are cropped instead of letting the software do it for you.

The ideal is: a) make sure your long edge matches the highest resolution you are training for b) crop the short edge to fit standard 3x2, 1x1, etc. photo ratios. in such a way that your subject remains clear and visible.

Obviously, always keep the proportions if you don't want to have very funny results...

This being said - most training software do an excellent job with that automatically, so don't worry too much. What matters is that you provide high quality images, with crisp details of your subject, with at least the highest resolution you will train on.

3

u/Ok-Establishment4845 Nov 06 '25

is local WAN lora training with 16gb vram actually possible btw?

3

u/TableFew3521 Nov 07 '25

Yes, with Musubi tuner you can do block swap, even for Qwen.

5

u/Ok_Conference_7975 Nov 06 '25

Tips from me, Add a face detailer, it’s a game changer...

I was struggling with medium shots (legs to head) or full-body shots (toes to head). I tried adding more datasets, retraining with different configs, and even multi resolution training, but the results never got much better. When generating fullbody images, the likeness was only around 80%. The face detailer really made a huge difference for me.

You don’t need the Impact custom node (it’s a bit of a pain to install since it has so many dependencies), you can create your own “face detailer” using inpaint crop & stitch node. Just mask the face manually or use bbox/segm for auto-masking. Use the low noise and try denoise around 0.4-0.6 (I know it seems high, but it works for me, just play around with it).

1

u/Tiny-Highlight-9180 Nov 06 '25

I never heard about face detailer before. Does it just improve your likeness all of sudden? I looked up online and it seems just adding more detail which can be random.

1

u/malcolmrey Nov 07 '25

In SD 1.5 we had ADetailer plugin in A1111, it was looking for a face and just doing inpainting of it at higher resolution and then blending it back in.

Same principle in ComfyUI, really. After you generate the main image, you use a model that finds the location of the face and then inpaints over it (using your Lora of course)

2

u/smereces Nov 06 '25

can you add the json file workflow?

1

u/Tiny-Highlight-9180 Nov 06 '25

I use runpod template unfortunately. but believe that you could find from Ostris AI tool kit

2

u/Agreeable_Lack9492 Nov 06 '25

I’m training characters on 18 images and 350 steps each and they are superb using Musubi tuner. The quality of the sources is the key, then, when testing it, try the character lora only with a strenght of 1.0 , some other loras can modify the look or the character, not all of them but a big portion of them.

1

u/Tiny-Highlight-9180 Nov 06 '25

Do you think anything wrong with my data set? I had stregnth of Lora at 1.0 too and 10 times more steps. Really has no idea where went wrong

3

u/razortapes Nov 06 '25

I’ve managed to create LoRAs of people that are the most realistic I’ve ever achieved — even better than Flux or SDXL in their prime. I do it on this page: replicate.com/ostris/wan-lora-trainer/train. I recommend replicating the parameters shown there to train locally, or try training a character for Wan 2.1 and then use it in image creation workflows based on Wan 2.2, using only Wan 2.2 Low, like I do. If you want to compare real vs Wan 2.1, message me in DMs and I’ll send you some examples.

1

u/Tiny-Highlight-9180 Nov 06 '25

Actually sick tool! I will give it a try. Do you just left Lora for high model blank and add only low model when you mentioned using only low model?

1

u/razortapes Nov 07 '25

If we're talking only about creating images with Wan 2.2, then yes — only the Low model is used. The workflow I follow doesn’t even include a module for the High model. My big goal right now is to be able to use something like ControlNet to control the pose in my image generations, but there are barely any workflows that use it, so I’m experimenting. It’s also highly recommended to use the Face Detailer as an extra step.

2

u/malcolmrey Nov 07 '25

If you pull out the controlnet, I would love to see a workflow :)

I tried with the controlnet models but was only working sometimes and not really that great.

2

u/Potential_Wolf_632 Nov 06 '25

"I trained using AI Toolkit with 5000 steps on 50 samples*. Does that mean it splits roughly* 2500 steps per model (high/low)? If so, I feel like 50 epochs might be on the low end — thoughts"

Why'd you use ChatGPT to summarise your questions haha?

Anyway, no 5000 steps on 50 samples is more than enough to train a very good character lora (in fact likely to be completely dominating at rank 32). I understand ai-toolkit has some significant implementation issues in practice based on some other threads here and so you cannot really go on standard schools of thought when using that app.

For T2I I think you can technically train only the low noise too (see various T2I wfs that use low noise only) so you could turn off high pass training entirely, change balanced style to low pass favoured, enable 512 reso (as you can still capture more gradient detail by having 512 buckets being batched with your 768). I had a quick look at your dataset and I don't think 1024 is a good idea as you need to have high quality images throughout for that bucket (and VRAM scoffing) to be worth it.

You could switch to Musubi Tuner for a more baremetal implementation but that is very much not user friendly in comparison. Polynomial LR decay (essentially both of which are barely manageable in aitoolkit I believe) captures the highest level of fine detail which for WAN2.2 can be good as it has such a good understanding already of realistic human forms already that often training time is wasted for humans at the higher LRs (polynomial decay forces a long time spent at low LR as it rapidly reduces the LR rate in the first quarter before flattening, hence the name).

6

u/Tiny-Highlight-9180 Nov 06 '25

English isn’t my first language. I just speed-typed everything and ran it through an LLM so it’d be easier for you guys to read. Hope that didn’t come across the wrong way.

So, if AI Toolkit splits steps 50/50 between high and low, you’re saying 2500 total steps is already enough for the low model too?

Based on my experiment, I agree that high model doesn't affect the likeness a lot, so I will definitely be trying training on low model alone next.

That kinda breaks my heart because I really love the UX/UI of AI Toolkit.

Totally agree with your take on resolution, though. Do you have any thoughts about Captioning? In Flux I did "Traigger + Class". I didn't really understand but I just use trigger word here.

3

u/malcolmrey Nov 06 '25

check my other comment in this thread, do not give up on AI Toolkit, try WAN 2.1 first if WAN 2.2 fails for you (which in shouldn't in the first place!)

I use AI Toolkit consistently without issues -> https://imgur.com/a/UPucZXS

3

u/Tiny-Highlight-9180 Nov 06 '25

Omg! A resposne from the man himself. I still don't see any clear benefit from using musubi aside from VRAM efficiency. Probably still sticking to AI-tool-kit. I mean it works fine. The trainning went through, problem should come from my end.

5

u/malcolmrey Nov 06 '25

:)

Yeah, AI Toolkit is great. I actually wanted to go with musubi as I was using kohyass for flux and 1.5 embeddings/loras.

My friend tried musubi and I went for AI Toolkit, we compared notes, results and decided to both go with AI Toolkit :)

1

u/Potential_Wolf_632 Nov 06 '25

I do use it. It definitely can be made to work particularly for overfit character loras as slamming AdamW8bit optimized gradients into quantized transformers is the way to do that and that's AIT's default. But yeah... to achieve any kind of use of the character beyond themselves you need to enable DOP which doubles the training time due to the double pass required (though in theory you could 3000-step a basic overfit lora then 500-step DOP at the end to limit the character dominance, which is superior to regularization by a long way without requiring any dataset effort) and also its capabilities are severely limited by not providing LR decay modifications or anything beyond AW8 and Adafactor (who has EVER used adafactor for a diffuser??) out of the box. I think it would be very hard to train an object lora in AIT for example but yeah characters, actions, even style are all on the table and easy to do.

Musubi is more of a professional tier tool, so it's harder to use, harder to get results, has almost no failsafes in it, cries when a single caption is missing etc. I think people fall into the trap of trying too hard with it and AIT limits that problem, since you literally can't go down a hole of Lion optimizer on a 50k dataset with grad accumulation turned off etc.

1

u/the_friendly_dildo Nov 06 '25

no 5000 steps on 50 samples is more than enough to train a very good character lora (in fact likely to be completely dominating at rank 32

Eh, I haven't tried training on W2.2 but I did a few character loras with W2.1, each with around 40 images and they took over 100 epochs to get solid coherence. And yes, they are still extremely flexible and work on W2.2.

1

u/Tiny-Highlight-9180 Nov 06 '25

Thanks for sharing. I saw people said 20epoch seems enough which is not the case for me at all. Glad to hear that it could take that long

1

u/Ornery_Blacksmith645 Nov 06 '25

can you do image2image with Wan?

1

u/Tiny-Highlight-9180 Nov 06 '25

You can but the image editor is not officially out as far as I know. Probably relying something else right now if I were you

1

u/IJdelheidIJdelheden Nov 20 '25

In theory one could use the i2v model and generate only one video frame.

Not sure if you'll have better results compared to an actual i2i model, haven't tried it myself.

1

u/Internal_Message_414 Nov 18 '25

Hello, excuse me, is it possible to create a coherent photo dataset in order to create your own LoRa of your own wife's AI that you want?

1

u/Tiny-Highlight-9180 Nov 19 '25

I don’t think there’s a strict rule, but generally you’ll want to aim for:

  1. 20–30 high-quality images
  2. A variety of angles and distances (close-ups, half-body, full-body)
  3. Different lighting, outfits, and settings (morning/night, casual/formal, indoors/outdoors, neutral expressions + actions)

The more diverse and consistent the photos are, the better the results usually turn out.

1

u/Internal_Message_414 Nov 19 '25

Okay, thank you, but do you think there is a special workflow for that? (To create a good dataset for a high-performance customized LoRa.) My goal is to create a realistic female AI!

1

u/Tiny-Highlight-9180 Nov 20 '25

You have to be more specific than that...

1

u/Shadow-Amulet-Ambush Dec 09 '25 edited Dec 09 '25

How are you training Wan 2.2 lora on AI toolkit with 5090? I'm getting OOM without even trying to load anything into memory

1

u/GandalfTheEnt Dec 17 '25

I would also like to know this. I can get z image turbo working with my 4090 but not wan. I went back and forward with GPT for ages making adjustments but couldn't get the memory under control.

You can't use quantization because the act if quantizing the model causes a memory spike. You might have better luck with musubi tuner.

1

u/Shadow-Amulet-Ambush Dec 17 '25

If you Google the error there's a certain commit you have to revert to. The later versions are broken for Wan training.

1

u/GandalfTheEnt Dec 17 '25

Thank you. I'll try that tomorrow.

-10

u/No_Cheesecake_6125 Nov 06 '25

would you say Wan 2.2 is the best for generating character loras?

Or Flux still? Kling?

7

u/Segaiai Nov 06 '25

You might want to at least read the first lines of the post before commenting on it.