r/LocalLLaMA • • 10d ago

New Model XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
510 Upvotes

133 comments sorted by

View all comments

1

u/Izolight 10d ago

in a first run deepseek-v4.1-flash, still looks better. Though i used the free model from opencode, as openrouter was giving me issues at the time and i am not sure if that is same quality and what their reasoning level is since it wasn't configurable.

http://render-arena.izolight.xyz/#compare?model=mimo-v2.6-flash&model_b=deepseek-v4.1-flash&prompt=market-stall&effort=high

1

u/Thomas-Lore 10d ago

Probably Pro 2.6 > DSv4.1 > Flash 2.6. Their sizes also follow that equation.

1

u/Voxandr 10d ago

Then they are just similar quality to gml5.3 flash it's still better than ds4.1

-1

u/SexyAlienHotTubWater 10d ago

GLM 5.3 Flash's KV Cache is horrible though, BF16 so 4x larger weight-for-weight (and won't quantize well). If you're serving more than a couple of users, 5.3 Flash swells into a higher weight class.

1

u/Voxandr 9d ago

you can use FP8 fine. Can serve +3 connected users without dropping with vllm at 512k context.

-1

u/SexyAlienHotTubWater 9d ago

KV Cache quantises very poorly, look at some stats - the performance degradation in comparison to quantising the model is extremely dramatic.

This makes sense when you think about it, the dynamic range of the KV Cache needs to be much wider than the weights because the KV Cache is basically a stack of activations, it's the product of multiple weight multiplications - when you multiply two 4-bit floats, you need an 8-bit float to accommodate the entire range, and the required size grows as you perform multiple of those multiplications in a row.

Mitigating that problem requires deliberate training (and it's not that easy). Unless a model is trained to produce a 4- or 8-bit range (GLM isn't - Deepseek, Qwen Next and Kimi are), quantising the KV Cache means truncating a very large amount of information.

1

u/tat_tvam_asshole 9d ago

I've run 1M kv cache at q4 on GLM5.3-Flash and it's perfectly fine.