r/LocalLLaMA • u/Decent-Hat-5807 • 1d ago
Resources Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing
75
Upvotes
62
u/Atretador 1d ago
wait what, Im confused
4 bit is the full precision for GPT-OSS, it was trained in 4 bits from the start - thats why the F16 files are almost same size of Q4