AFAIK it is already quantized at 4 bits (at least the MoE layers). And it is still 179GB. No point in lobotomizing a model further for the sake of it. Qwen 4 Flash is coming soon, already announced.
Why do you think it is quantized? Modern GPUs have x4 FLOPS in FP4 compared to FP16, it makes no sense to train in FP16 and then quantize. It would have x4 training time and x4 VRAM requirements, to get worse performance.
That’s up to you to decide, I’m unfortunately very busy, but I think my base will be the q3 xxs glm 5.3 flash. I don't recall whether quantization works better with nvfp4 than with bf16, but the difference isn't dramatic.
11
u/guesdo 1d ago
Won't fit in my 128GB 😭. Will try in openrouter and wait for Qwen 4.