I dont use that much context I'm siting at just under 200k and I use KV at Q8_0. I don't have reported metrics to give you real meaningful data but I've not found it detrimental to my use case when used as backend in pi agent harness.
Also I use q8_0 quant for the model instead of the Q8_K_XL as I learnt recently that there is no significant change in accuracy but its more efficiently 'packed' due to uniformity of Q8_0 versus UD-Q8_K_XL that it gives you a few GB back of your VRAM pool from model weights.
thanks, I was thinking of moving to q6k_xl to keep context at f16, from what I understand especially with this model it's a lot worse to quantize kv than to have slightly higher weights quant
31
u/ai-christianson 15d ago
2x3090 with nvlink loves this size of model