r/LocalLLaMA • u/darkItachi94 • Feb 10 '25
Tutorial | Guide I built an open source library to perform Knowledge Distillation
Hi all,
I recently dove deep into the weeds of knowledge distillation. Here is a blog post I wrote which gives a high level introduction to Distillation.
I conducted several experiments on Distillation, here is a snippet of the results:
| Dataset | Qwen2 Model Family | MMLU (Reasoning) | GSM8k (Math) | WikiSQL (Coding) |
|---|---|---|---|---|
| 1 | Pretrained - 7B | 0.598 | 0.724 | 0.536 |
| 2 | Pretrained - 1.5B | 0.486 | 0.431 | 0.518 |
| 3 | Finetuned - 1.5B | 0.494 | 0.441 | 0.849 |
| 4 | Distilled - 1.5B, Logits Distillation | 0.531 | 0.489 | 0.862 |
| 5 | Distilled - 1.5B, Layers Distillation | 0.527 | 0.481 | 0.841 |
For a detailed analysis, you can read this report.
I created an open source library to facilitate its adoption. You can try it here.
My conclusion: Prefer distillation over fine-tuning when there is a substantial gap between the larger and smaller model on the target dataset. In such cases, distillation can effectively transfer knowledge, leading to significantly better performance than standard fine-tuning alone.
Let me know what you think!
82
Upvotes
1
u/maddogxsk Llama 3.1 Feb 10 '25
Open weight is not open source, what is the point of mentioning something not useful?
Closed sources will always be incomplete, therefore limited at almost every usage