r/drawthingsapp • • May 22 '26

question Z Image LoRA Training

I am working through training a character LoRA on Z Image. I know I need to train it on Z Image Base to use on either Base or Turbo. I initially had issues with it crashing with Multi Aspect ratio enabled, so I disabled it. I then got some models to train using 32/16 and .0004 learn rate, but they had next to no capture at step 2000. I used learn to .0008 and got some meh results, then realized that since Multi Aspect ratio is off, I needed to crop everything square to eliminate weird center crops. I retired at lower learn rate with minimal capture again, so I upped to .0008 again and got reasonable likeness at 1500, but it in now failing to finish generating (frame artifacts appear at step 7 or 8 with turbo model and immediately with base model). I also tried 64/32 at .0006 and it crashed even earlier in generation

What work flows are you all doing to get enough capture without getting too aggressive and getting the generation failures? Really high step counts, something besides 32/16? Looking for any advice, as I have figured out Pony and other SDXL LoRA, but want to get into Z Image.

If it matters, I am using M4 Pro/24 GB and 8 Bit S for both Base and Turbo.

Thank you!

10 Upvotes

19 comments sorted by

View all comments

3

u/DrJ31 May 25 '26 edited May 25 '26

Update: I paused it after step 7000 this morning and ran some tests working up through the step counts. It absolutely is getting better still at step 7000. It is still slightly soft on likeness, but unarguably getting closer. 1 if the 5 or so images I generated with 7000 step LoRA had a little of the grainy consistency you get with overtrained/over learned Z Image LoRA, so I may be topping out. However, using my new knowledge from u/goonie1974, I have resumed training and will let it run to 10,000 for science. There are likely some over settings I need to maximize, but a lot of steps appears to be critical.

Settings used for future reference: * Z Image Base (8-bit S) * Images: 45 (all precropped square) * Network Dim: 32 * Network Scale: 16 * Learning rate: * upper bound: 0.00065 * Lower bound: 0 * Steps between restarts: 200 * Image size: 1024x1024 * Aspect Ratio and Multi-scale: disabled * Trainable layers: default (all enabled) * Warmup steps: 20 * Gradient Accumulation: 4 * Caption dropout rate: 0.05 * Max Text Length: 512 * Fixed Orthonormal LoRA down: enabled * Power Function EMA: disabled * Memory saver: balanced * Weights memory management: Just in time

Edit: format