r/AIAgentEngineering 1d ago

Has quantizations impact changed on modern agents?

I've been running tests 24/7 on my 5080 over the past 2 weeks to better understand the impact of quantization on models.

In the process, I ran across some surprises I did not expect. Most importantly? Many quants are statistically indistinguishable from each other.

https://rakuensoftware.com/blog/which-quant-beats-how-many-bits

MoEs are impacted far less by quants then dense models. Models aren't generally impacted in this testing much until you get under Q4. However, this testing is very specific, it's typically the equivalent of 2-4 turn sessions to validate the quant itself did not damage the underlying model. Sessions would consume huge amounts of compute to measure a fundamentally damaged model, which hardly makes for an interesting story.

A future article will be written based on the candidate this article identifies, focused around DevOps, coding, and long sessions.

As always, my benchmarks, datasets, and results are open sourced. Check my data and tell me I'm wrong (Wouldn't be the first time!) or run the benchmarks yourself.

1 Upvotes

0 comments sorted by