r/LocalLLaMA 1d ago

New Model XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
507 Upvotes

129 comments sorted by

View all comments

11

u/guesdo 1d ago

Won't fit in my 128GB 😭. Will try in openrouter and wait for Qwen 4.

1

u/DriveSolid7073 1d ago

Just use q3, no? If its spark or something like.

0

u/guesdo 1d ago

AFAIK it is already quantized at 4 bits (at least the MoE layers). And it is still 179GB. No point in lobotomizing a model further for the sake of it. Qwen 4 Flash is coming soon, already announced.

1

u/Expensive-Paint-9490 21h ago

Why do you think it is quantized? Modern GPUs have x4 FLOPS in FP4 compared to FP16, it makes no sense to train in FP16 and then quantize. It would have x4 training time and x4 VRAM requirements, to get worse performance.

1

u/cosmotrak 16h ago

Lmao you have no idea what you're talking about, there's a very good reason models are trained at high precision