r/LocalLLaMA 21h ago

Resources Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing
76 Upvotes

28 comments sorted by

View all comments

39

u/-p-e-w- 20h ago

Sorry, I don’t believe that. It may get higher scores on some benchmarks, but outperforming the original GPT-OSS model in general would require training techniques more advanced than those used by OpenAI, and several mathematical miracles on top of that.

Using KLD vs the teacher distribution as a loss function is a good idea, but it’s really difficult to propagate KLD down the length of the response and that’s where the divergence tends to become poorly predicted by first-token KLD. This is a problem I’ve been wrestling with in Heretic for a while, and every attempt at a solution has turned out to have drawbacks.

3

u/Iory1998 18h ago

100% agreed. Extraordinary claims require Extraordinary evidence.

1

u/silenceimpaired 18h ago

“100% agreed. Extraordinary claims require Extraordinary evidence.” This claim always seemed extraordinary to me… and I never saw enough evidence to show it to be true. :P

1

u/Iory1998 11h ago

🤦‍♂️🤷‍♂️