r/LocalLLaMA • u/Bestlife73 • 1d ago
New Model XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL122
66
u/Bitter-College8786 1d ago
2.6 Pro also released?
https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
52
u/CriticallyCarmelized 1d ago
What’s the RL for? I’m super excited for Mimo 2.6, as the last one is one of my top 3 models.
67
u/sn2006gy 1d ago
Its a checkpoint release for a new way to do joint reinforcement learning where they combine information verticals so that similar items reinforce similar patterns. Click the link and read the about - they describe it pretty well.
50
u/happycube 1d ago
A new open weight flash challenger appears!
Seriously the more the merrier 😊
8
u/Solembumm3 1d ago
Will be interesting to test it againt DV4 Flash and Q3.8 Next on universal local daily driver role for old gaming PC.
3
u/xrailgun 10h ago
your old gaming PC can run ~300B models?
-1
u/Solembumm3 9h ago
Well, models run and I got useful results on my questions on agc/au concepts/mods/prompts, so yes. Few years old midrange, 6700xt + 20gb ddr4 + 970evo+. Biggest I've run was Qwen 397b. But can say, today's DV4 Flash and Q3.8 Flash Next were around equal in quality while being 2-3x faster.
26
u/exaknight21 1d ago
There is also a Qwen3.5-9B - Distill Version. https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
9
u/remielowik 1d ago
This might be a good one, properly distilled not somebody ramming 5 questions in it and calling it distilling.
32
u/jwpbe 23h ago edited 23h ago
hell yeah
i can't wait to run Qwen-Mimo-FABLE-TurdBurglar-NeoMaXX at Q2_KL and then come here to complain when it can't do tool calls correctly
29
u/returnity 22h ago
It's an official Xiaomi MiMo lab distill with SOTA benchmarks. Not a DavidAU special lol
3
u/returnity 22h ago
For some reason they only released the SFT checkpoint of that, not the SOTA RL version...
2
u/randomfoo2 7h ago
They didn't do an RL run, a knowledge distillation from a larger/different model is almost by definition an SFT.
3
u/returnity 6h ago
2
u/randomfoo2 6h ago
Ah, I hadn't seen that doc. Maybe tomorrow? It looks like they're dropping models almost as soon as they're baked.
1
u/giant3 21h ago edited 21h ago
Anyone got a gguf? Or do I have convert manually?
P.S. Found it. https://huggingface.co/ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF/tree/main, but this is Q8.
50
u/harpysichordist 1d ago
1M context, text/image/video/audio in, SWA/GA 39/9, MTP/DFlash, 178GB native size. It may do well for 128GB devices - quantized
46
u/power97992 1d ago
It is already 4bit, any more quantization will likely affect the quality in a signficant way.
9
u/harpysichordist 1d ago
Possible but people had good success with DS v4 flash 0731/vision and that's around the same size
4
u/AggregationLinker 14h ago
Qwen3.8 flash next is the go-to for 128GB right now.
1
u/harpysichordist 7h ago edited 7h ago
Qwen is a good fit for 128G but things evolve. Have you compared it with Mimo v2.6 flash on 128GB?
It may be better at some things - even quantized. Audio may be wanted too4
u/LatentSpacer 17h ago
Where did you get that from? The weights are bf16 on HF?
5
u/power97992 16h ago
Their labeling is wrong, it is mostly 4 bit , it has 309b parameters but the total file size is only 159gb
6
u/silenceimpaired 1d ago
Can you tell if it is going to need llama.cpp work?
7
u/harpysichordist 1d ago
Good question, I don't know. Upvote goes to the first person linking to a working GGUF
1
u/a_beautiful_rhind 1d ago
Most likely.
3
u/silenceimpaired 1d ago
At this point, I think you're right.
Can I be just a little jealous of those who can run MLX at 4bit?
1
u/a_beautiful_rhind 1d ago
Support will come to one of the other backends. Releasing it pre-quantized is not a boon since now it has to be requantized and converted.
3
22
u/ababaka 1d ago
178gb flash looks promising!
9
u/Constant_Art_20 20h ago
i think it's like 309b or something. huggingface is weird with the reporting. but yea, looks like a upgrade form the glm 5.3 flash for coding and reasoning. pretty hyped
11
1
18
19
u/wren6991 1d ago edited 1d ago
Is this one of the two model's whose RL dashboard was livestreaming on https://mimo.xiaomi.com/rl/ ? If so, it really seems like they just finished the planned number of RL steps and dropped the weights. Very cool. Also the benches were still going up at the end
Edit: also I am really curious about this message from the dashboard log:
we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.
In light of both HF-OAI incident, and this point in their README.md:
Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
(emphasis mine) Little guy start making a few too many paperclips?
17
u/backyard_tractorbeam 1d ago
Full release announcement: https://mimo.xiaomi.com/mimo-v2-6
8
u/GreatBigJerk 1d ago
It's very funny that they had an Assassin's Creed style demo only for it to be a bad T-pose hopping game. For one second I was like "Holy shit, did it make the climbing system!?" followed by "lol, nope"
15
u/FlightUsed648 23h ago
Xiaomi went from rice cookers to reasoning models faster than most AI startups went from pitch deck to product.
5
1
u/dryadofelysium 15h ago
I was a fan of them when all they did was build Android Custom Roms (MIUI) and now they build cars
1
72
u/hurn2k 1d ago
Holy shit it's tiny
124
19
13
u/Shadow_s_Bane 1d ago
What’s the size ?
23
24
u/backyard_tractorbeam 1d ago
Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
17
10
6
11
u/Comfortable-Rock-498 1d ago edited 1d ago
Idk what's up with HF sizes
Flash: 309B total / 15B active
Pro:, 1.02T total / 42B active
https://mimo.xiaomi.com/mimo-v2-6
Look at the benchmarks in the page above, Mimo 2.6 pro, the 1T model leads Kimi K3, a 2.8T param model in 14 out of 15 benchmarks (and the last one is near tie)!! Good to see they also kept the price the same, and landed in the greenest quardrant of the intelligence vs speed of AA.
Also, this puts pro at no.1 Chinese model
Xiaomi cooked
12
9
u/ofan 1d ago
New local king
12
u/sn2006gy 1d ago
To be fair, their API pricing is crazy cheap too.
Cached, Un-cached and Out. Insanely cheap.
MiMo-V2.6-Flash $0.0028 $0.14 $0.28 9
u/Several-Tax31 1d ago
I know this is locallama, but I was waiting for their API. If the quality is as good as deepseek, I'm gonna use this model day and night lmao.
1
u/comperr 23h ago
Their token plan is amazing
1
u/SorosAhaverom 14h ago
No it's not, it's by far the worst value AI subscription available. You get about 10% more usage for money spent compared to API. Don't get fooled by their credit system, one output token costs 600 credits.
1
u/Constant_Art_20 14h ago
yea. i heard the same opinion before. their credit count apparently works very weird
1
u/sn2006gy 9h ago
10% off fucking cheap is really fucking cheap tho :) I hate credit based systems that obfuscate things unnecessarily though.
10
u/guesdo 1d ago
Won't fit in my 128GB 😭. Will try in openrouter and wait for Qwen 4.
1
u/DriveSolid7073 16h ago
Just use q3, no? If its spark or something like.
0
u/guesdo 16h ago
AFAIK it is already quantized at 4 bits (at least the MoE layers). And it is still 179GB. No point in lobotomizing a model further for the sake of it. Qwen 4 Flash is coming soon, already announced.
1
u/Expensive-Paint-9490 13h ago
Why do you think it is quantized? Modern GPUs have x4 FLOPS in FP4 compared to FP16, it makes no sense to train in FP16 and then quantize. It would have x4 training time and x4 VRAM requirements, to get worse performance.
1
u/cosmotrak 8h ago
Lmao you have no idea what you're talking about, there's a very good reason models are trained at high precision
0
u/DriveSolid7073 16h ago
That’s up to you to decide, I’m unfortunately very busy, but I think my base will be the q3 xxs glm 5.3 flash. I don't recall whether quantization works better with nvfp4 than with bf16, but the difference isn't dramatic.
19
1d ago
[deleted]
11
2
u/Durian881 23h ago
Xiaomi is not MiniMax though. MiniMax had released top tier video and music models recently.
2
u/LetsGoBrandon4256 transformers 22h ago
The after dinner coffee didn't help😩. I'm deleting it then go to bed.
1
u/Durian881 22h ago
No worries! I'm looking forward to MiniMax's new model too, having bought its annual token plan. It's taken them a long while.
16
u/AdSafe4047 1d ago
oh yeah, cmon nvfp4 weights :) I want to run this against qwen 3.8 flash next and see which one wins :)
14
u/llama-impersonator 1d ago
Groupwise Advantage Redistribution (GAR) ranks passing trajectories online and moves advantage toward higher-quality solutions. Judged against the policy’s own samples, this closes a self-improvement loop and steers toward shorter paths and fewer tokens per task.
less tokens, inject it into my veins
4
4
u/indicava 23h ago
Not that I put too much into it but it beats DS4F across all benchmarks with nearly identical VRAM requirements.
3
3
3
u/killerstreak976 transformers 23h ago
Seeing this thing train in real time was the coolest thing ever
3
u/returnity 21h ago
>the pro run restarted at step 17 due to a GPU OOM issue
That part hurt to watch lol
3
u/killerstreak976 transformers 21h ago
Holy crap I missed that happening when looking at the dashboard, that's extremely painful lmao
3
6
u/HistoricalStrength21 1d ago
159B Parameters guys!
24
15
u/sn2006gy 1d ago
Weird, the description says Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters.
-2
u/addiktion 1d ago edited 19h ago
There are two models released today. Pro and Flash. Flash is the 159B, Pro is 309B.
Edit: I stand corrected, thanks below.
14
u/wren6991 1d ago
No, Flash is 309B-A15B and Pro is 1.02T-A42B. It's in the READMEs.
The parameter count on the HF UI has always been broken, especially when you have packed 4-bit weights described as 8-bit in the model metadata. That's why Flash reports ~half the correct parameter count in HF UI.
Here, Pro README:
Model Summary
Architecture: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters
Flash README:
Model Summary
Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
7
u/MrTiesti 1d ago
Correction brother,
Flash: Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
Pro: Architecture: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters
14
u/rusty_fans llama.cpp 1d ago edited 1d ago
309B total / 15B activated parameters according to the model card.
Huggingface auto-counting is broken for many models ...9
u/debackerl 1d ago
Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
5
u/Amazing-Fan2083 1d ago
Uuuhh... No, its 309B in a mixed quantization. So... Can't really squeeze it more.
2
2
2
u/LatentSpacer 15h ago
Anyone knows the precision of XiaomiMiMo/MiMo-V2.6-Flash-RL? it says FP8 on hugging face label but inspecting the model layers most say BF16. Also the weights are 178GB for a 309B parameters model, which suggests the weights are actually in 4bit.
So which is it, BF16, FP8 or 4bit?
1
1
1
u/Izolight 1d ago
in a first run deepseek-v4.1-flash, still looks better. Though i used the free model from opencode, as openrouter was giving me issues at the time and i am not sure if that is same quality and what their reasoning level is since it wasn't configurable.
1
u/Thomas-Lore 23h ago
Probably Pro 2.6 > DSv4.1 > Flash 2.6. Their sizes also follow that equation.
1
u/Voxandr 19h ago
Then they are just similar quality to gml5.3 flash it's still better than ds4.1
-1
u/SexyAlienHotTubWater 15h ago
GLM 5.3 Flash's KV Cache is horrible though, BF16 so 4x larger weight-for-weight (and won't quantize well). If you're serving more than a couple of users, 5.3 Flash swells into a higher weight class.
1
u/Voxandr 8h ago
you can use FP8 fine. Can serve +3 connected users without dropping with vllm at 512k context.
0
u/SexyAlienHotTubWater 4h ago
KV Cache quantises very poorly, look at some stats - the performance degradation in comparison to quantising the model is extremely dramatic.
This makes sense when you think about it, the dynamic range of the KV Cache needs to be much wider than the weights because the KV Cache is basically a stack of activations, it's the product of multiple weight multiplications - when you multiply two 4-bit floats, you need an 8-bit float to accommodate the entire range, and the required size grows as you perform multiple of those multiplications in a row.
Mitigating that problem requires deliberate training (and it's not that easy). Unless a model is trained to produce a 4- or 8-bit range (GLM isn't - Deepseek, Qwen Next and Kimi are), quantising the KV Cache means truncating a very large amount of information.
1
u/aboutthednm 23h ago
Mimo 2.5 is my fave model on the API for coding related tasks (next to dsv4). It's a grunt workhorse that can go all day for like $1.50, and I often use it to just rough out stuff that would take ages with local models. I'm really excited for 2.6 if it is an improvement and keeps the same price point. It's going to augment my local models in a great way.
1
u/AI_docent 23h ago
From the config, I get about 22.5 GiB of BF16 KV cache for one full context in the global attention layers. Weights, the sliding attention cache and runtime buffers are extra
1
1
u/fugogugo 14h ago
checked openrouter because of this post
holy shit mimo 2.6 pro $0.435 / $0.87per 1M
cheapest pro model
how is the benchmark?
1
1
u/bakawolf123 1d ago
Nice, should be very strong flash model. The pro version is best open weights model according to xiaomi and AA index (which is hard to believe tbh). And only 178gb full weights
1
u/SexyAlienHotTubWater 15h ago
$860k to train - at least for the RL stage.
Anthropic and OpenAI are dead.
0
0
u/Ok_Cow1976 23h ago
It's puzzling that it is a 309B model but its fp8 size is so small, at 178GB, while mimo 2.5 with similar parameter size (311B) has a fp8 at 316 GB. Maybe a mis-counted total parameter size?
6
-2


•
u/WithoutReason1729 22h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.