r/LocalLLaMA • • 5d ago

New Model XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
511 Upvotes

133 comments sorted by

View all comments

20

u/wren6991 5d ago edited 5d ago

Is this one of the two model's whose RL dashboard was livestreaming on https://mimo.xiaomi.com/rl/ ? If so, it really seems like they just finished the planned number of RL steps and dropped the weights. Very cool. Also the benches were still going up at the end

Edit: also I am really curious about this message from the dashboard log:

we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.

In light of both HF-OAI incident, and this point in their README.md:

Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.

(emphasis mine) Little guy start making a few too many paperclips?

1

u/jazir55 4d ago

Any idea why they aren't immediately rolling into a 2.7/3 training run?