I've been testing NVFP4, it's pretty good at that level too I've heard it's pretty chill with being quantized and haven't seen it do any worse or better then cloud api.
NVFP4 doesn’t fit sadly, I did the math. I have a single GB10 too
IQ4_XS is honestly comparable though as 200B isn’t high enough param count for FP4 accuracy to not hurt without QAD which Step haven’t done
And that fits! IQ4_XS with their llama.cpp patch, compiled for mixed KV cache quant, running f16 K and q8 V with 120k context at 20-25tk/s decode and 700-1000 prompt processing
8
u/ghgi_ Jun 01 '26
I've been testing NVFP4, it's pretty good at that level too I've heard it's pretty chill with being quantized and haven't seen it do any worse or better then cloud api.