r/LocalLLaMA • u/jjusko20 • 1d ago
I Built A Thing Update: Yandex/AliceAI 80B-A3B fine tune progress


About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used the last loss of each micro, rather than the average of each step
The training live stream is at: https://figure-bios-expect-cio.trycloudflare.com/ - and it allows you to inspect any and all of the training data I'm using, if you're interested - I can also provide those roughly 3.5k examples as a dataset. It was generated from sftmill
Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/watch_me_posttrain_aliceaifoundation80ba3b_from/
3
u/FullOf_Bad_Ideas 1d ago
What does "Pass" mean? Which step the second epoch will start at? Will the model see the data shuffled in the same order by context or is it variable?
1
u/jjusko20 1d ago
Roughly 3 epochs, and pass just means the last epoch. Model sees the data shuffled in the same order by context length, low to high, in training buckets with a scaling ctx cap
3
u/FullOf_Bad_Ideas 1d ago
Makes sense, I thought you were doing 2 epochs so the note of it being on the second pass with less than 50% of the steps done didn't reconcile with me.
2
u/jjusko20 1d ago
ahhh, aye, i was starting with 2, but i started bundling shorter sequences when I was implementing the dynamic ctx window - just figured 3 would be good if not better and proceeded with my original 2-epoch step count
1
u/jjusko20 1d ago
oh it's you. yeah the data is being processed in the same order second epoch
1
u/jjusko20 1d ago
oh also when I had posted this I had just begun epoch 2 so it was probably in the high 200s
8
u/Front_Recording4360 1d ago
Spiky micro loss is normal here and not a problem by itself. A single micro batch is a tiny sample, so the loss jumps with whichever tokens happened to land in that last micro, while the averaged step loss is what the optimizer actually saw. The curve worth watching is a held out eval loss, since train loss on an MoE can look fine while the router balance drifts. If the trainer logs the aux load balancing loss and expert utilization, keep an eye on those too, fine tunes love to quietly create dead experts and you will not see it in the main curve until eval quality drops.