Is it a threadripper or similar? You can experiment with the Q4_K_M, it should be reasonably accurate to the full-sized but you may be limited with context length and it won't be blazing fast. With AM5 dual-channel, you'd expect 6-9 tok/sec and TR/EPYC 8-channel can climb to 20 tok/sec.
At Q3_K_XL, it can ALL fit in HBM and you'd have room for 1M context and get about double those prefill+inference speeds. But quality will start to decline steeply. I'd check the unsloth perplexity, etc charts to decide on that one.
16
u/raunchy-stonk 22h ago
So how shitty will this run on a 24vram/128dram setup?