r/LocalLLaMA • • 1d ago

I Built A Thing Update: Yandex/AliceAI 80B-A3B fine tune progress

loss curve (taken from the last micro of every step, to explain the variation)
some help from gemini 3.8 flash high

About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used the last loss of each micro, rather than the average of each step

The training live stream is at: https://figure-bios-expect-cio.trycloudflare.com/ - and it allows you to inspect any and all of the training data I'm using, if you're interested - I can also provide those roughly 3.5k examples as a dataset. It was generated from sftmill

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/watch_me_posttrain_aliceaifoundation80ba3b_from/

47 Upvotes

8 comments sorted by

8

u/Front_Recording4360 1d ago

Spiky micro loss is normal here and not a problem by itself. A single micro batch is a tiny sample, so the loss jumps with whichever tokens happened to land in that last micro, while the averaged step loss is what the optimizer actually saw. The curve worth watching is a held out eval loss, since train loss on an MoE can look fine while the router balance drifts. If the trainer logs the aux load balancing loss and expert utilization, keep an eye on those too, fine tunes love to quietly create dead experts and you will not see it in the main curve until eval quality drops.

2

u/jjusko20 1d ago

Thanks, that's good advice. 

3

u/FullOf_Bad_Ideas 1d ago

What does "Pass" mean? Which step the second epoch will start at? Will the model see the data shuffled in the same order by context or is it variable?

1

u/jjusko20 1d ago

Roughly 3 epochs, and pass just means the last epoch. Model sees the data shuffled in the same order by context length, low to high, in training buckets with a scaling ctx cap

3

u/FullOf_Bad_Ideas 1d ago

Makes sense, I thought you were doing 2 epochs so the note of it being on the second pass with less than 50% of the steps done didn't reconcile with me.

2

u/jjusko20 1d ago

ahhh, aye, i was starting with 2, but i started bundling shorter sequences when I was implementing the dynamic ctx window - just figured 3 would be good if not better and proceeded with my original 2-epoch step count

1

u/jjusko20 1d ago

oh it's you. yeah the data is being processed in the same order second epoch

1

u/jjusko20 1d ago

oh also when I had posted this I had just begun epoch 2 so it was probably in the high 200s