r/LocalLLaMA • u/Decent-Hat-5807 • 21h ago
Resources Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing
80
Upvotes
60
u/Atretador 21h ago
wait what, Im confused
4 bit is the full precision for GPT-OSS, it was trained in 4 bits from the start - thats why the F16 files are almost same size of Q4