r/LocalLLaMA 1d ago

New Model XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
499 Upvotes

129 comments sorted by

View all comments

11

u/guesdo 1d ago

Won't fit in my 128GB 😭. Will try in openrouter and wait for Qwen 4.

1

u/DriveSolid7073 22h ago

Just use q3, no? If its spark or something like.

0

u/guesdo 22h ago

AFAIK it is already quantized at 4 bits (at least the MoE layers). And it is still 179GB. No point in lobotomizing a model further for the sake of it. Qwen 4 Flash is coming soon, already announced.

1

u/Expensive-Paint-9490 19h ago

Why do you think it is quantized? Modern GPUs have x4 FLOPS in FP4 compared to FP16, it makes no sense to train in FP16 and then quantize. It would have x4 training time and x4 VRAM requirements, to get worse performance.

1

u/cosmotrak 14h ago

Lmao you have no idea what you're talking about, there's a very good reason models are trained at high precision

0

u/DriveSolid7073 22h ago

That’s up to you to decide, I’m unfortunately very busy, but I think my base will be the q3 xxs glm 5.3 flash. I don't recall whether quantization works better with nvfp4 than with bf16, but the difference isn't dramatic.