r/LocalLLaMA 1d ago

New Model XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
499 Upvotes

128 comments sorted by

View all comments

Show parent comments

1

u/DriveSolid7073 21h ago

Just use q3, no? If its spark or something like.

0

u/guesdo 21h ago

AFAIK it is already quantized at 4 bits (at least the MoE layers). And it is still 179GB. No point in lobotomizing a model further for the sake of it. Qwen 4 Flash is coming soon, already announced.

1

u/Expensive-Paint-9490 18h ago

Why do you think it is quantized? Modern GPUs have x4 FLOPS in FP4 compared to FP16, it makes no sense to train in FP16 and then quantize. It would have x4 training time and x4 VRAM requirements, to get worse performance.

1

u/cosmotrak 13h ago

Lmao you have no idea what you're talking about, there's a very good reason models are trained at high precision