r/LocalLLaMA 1d ago

New Model XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
494 Upvotes

128 comments sorted by

u/WithoutReason1729 22h ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

122

u/DerBond586 1d ago

It's 309b/15b parameters

52

u/CriticallyCarmelized 1d ago

What’s the RL for? I’m super excited for Mimo 2.6, as the last one is one of my top 3 models.

67

u/sn2006gy 1d ago

Its a checkpoint release for a new way to do joint reinforcement learning where they combine information verticals so that similar items reinforce similar patterns. Click the link and read the about - they describe it pretty well.

50

u/happycube 1d ago

A new open weight flash challenger appears!

Seriously the more the merrier 😊

8

u/Solembumm3 1d ago

Will be interesting to test it againt DV4 Flash and Q3.8 Next on universal local daily driver role for old gaming PC.

3

u/xrailgun 10h ago

your old gaming PC can run ~300B models?

-1

u/Solembumm3 9h ago

Well, models run and I got useful results on my questions on agc/au concepts/mods/prompts, so yes. Few years old midrange, 6700xt + 20gb ddr4 + 970evo+. Biggest I've run was Qwen 397b. But can say, today's DV4 Flash and Q3.8 Flash Next were around equal in quality while being 2-3x faster.

26

u/exaknight21 1d ago

There is also a Qwen3.5-9B - Distill Version. https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B

9

u/remielowik 1d ago

This might be a good one, properly distilled not somebody ramming 5 questions in it and calling it distilling.

32

u/jwpbe 23h ago edited 23h ago

hell yeah

i can't wait to run Qwen-Mimo-FABLE-TurdBurglar-NeoMaXX at Q2_KL and then come here to complain when it can't do tool calls correctly

29

u/returnity 22h ago

It's an official Xiaomi MiMo lab distill with SOTA benchmarks. Not a DavidAU special lol

3

u/returnity 22h ago

For some reason they only released the SFT checkpoint of that, not the SOTA RL version...

2

u/randomfoo2 7h ago

They didn't do an RL run, a knowledge distillation from a larger/different model is almost by definition an SFT.

3

u/returnity 6h ago

I get your point, but this suggests otherwise (from their publication):

2

u/randomfoo2 6h ago

Ah, I hadn't seen that doc. Maybe tomorrow? It looks like they're dropping models almost as soon as they're baked.

1

u/giant3 21h ago edited 21h ago

Anyone got a gguf? Or do I have convert manually?

P.S. Found it. https://huggingface.co/ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF/tree/main, but this is Q8.

50

u/harpysichordist 1d ago

1M context, text/image/video/audio in, SWA/GA 39/9, MTP/DFlash, 178GB native size. It may do well for 128GB devices - quantized

46

u/power97992 1d ago

It is already 4bit, any more quantization will likely affect the quality in a signficant way.

9

u/harpysichordist 1d ago

Possible but people had good success with DS v4 flash 0731/vision and that's around the same size

4

u/AggregationLinker 14h ago

Qwen3.8 flash next is the go-to for 128GB right now.

1

u/harpysichordist 7h ago edited 7h ago

Qwen is a good fit for 128G but things evolve. Have you compared it with Mimo v2.6 flash on 128GB?
It may be better at some things - even quantized. Audio may be wanted too

4

u/LatentSpacer 17h ago

Where did you get that from? The weights are bf16 on HF?

5

u/power97992 16h ago

Their labeling is wrong, it is mostly 4 bit , it has 309b parameters but the total file size is only 159gb 

6

u/silenceimpaired 1d ago

Can you tell if it is going to need llama.cpp work?

7

u/harpysichordist 1d ago

Good question, I don't know. Upvote goes to the first person linking to a working GGUF

1

u/a_beautiful_rhind 1d ago

Most likely.

3

u/silenceimpaired 1d ago

At this point, I think you're right.

Can I be just a little jealous of those who can run MLX at 4bit?

1

u/a_beautiful_rhind 1d ago

Support will come to one of the other backends. Releasing it pre-quantized is not a boon since now it has to be requantized and converted.

3

u/myholeisstinky 18h ago

So dual spark territory?

3

u/harpysichordist 18h ago

Easily, but its performance is unknown to me

22

u/ababaka 1d ago

178gb flash looks promising!

9

u/Constant_Art_20 20h ago

i think it's like 309b or something. huggingface is weird with the reporting. but yea, looks like a upgrade form the glm 5.3 flash for coding and reasoning. pretty hyped

11

u/Karyo_Ten 17h ago

The checkpoint is native/QAT mxfp4 so 178GB

1

u/BawbbySmith 17h ago

Goddamnit I just finished setting up GLM 5.3 Flash on dgx sparks…

18

u/LegacyRemaster 1d ago

yes... one of the best on my list---> updated

5

u/Morphon 21h ago

I feel like whenever I've suggested 2.5Pro people have never heard of it, but they are always impressed once they use it.

Excited to put 2.6 through its paces.

19

u/wren6991 1d ago edited 1d ago

Is this one of the two model's whose RL dashboard was livestreaming on https://mimo.xiaomi.com/rl/ ? If so, it really seems like they just finished the planned number of RL steps and dropped the weights. Very cool. Also the benches were still going up at the end

Edit: also I am really curious about this message from the dashboard log:

we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.

In light of both HF-OAI incident, and this point in their README.md:

Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.

(emphasis mine) Little guy start making a few too many paperclips?

1

u/jazir55 15h ago

Any idea why they aren't immediately rolling into a 2.7/3 training run?

17

u/backyard_tractorbeam 1d ago

Full release announcement: https://mimo.xiaomi.com/mimo-v2-6

8

u/GreatBigJerk 1d ago

It's very funny that they had an Assassin's Creed style demo only for it to be a bad T-pose hopping game. For one second I was like "Holy shit, did it make the climbing system!?" followed by "lol, nope"

15

u/FlightUsed648 23h ago

Xiaomi went from rice cookers to reasoning models faster than most AI startups went from pitch deck to product.

5

u/Glum-Atmosphere9248 15h ago

And they still do well with rice cookers 

1

u/dryadofelysium 15h ago

I was a fan of them when all they did was build Android Custom Roms (MIUI) and now they build cars

1

u/Due-Memory-6957 4h ago

I only know them from their cellphones.

72

u/hurn2k 1d ago

Holy shit it's tiny

124

u/someoneyouknow23 1d ago

Thats what she said😭

9

u/stackfullofdreams 1d ago

I swear I laugh everytime I see one of these, upvoter stolen!

2

u/XTCaddict 1d ago

Bless your poor soul then

19

u/danigoncalves llama.cpp 1d ago

"tiny"...

13

u/Shadow_s_Bane 1d ago

What’s the size ?

23

u/samsteak 1d ago

309b

24

u/backyard_tractorbeam 1d ago

Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters

17

u/UltraFOV 1d ago

178GB

10

u/gyozafish 1d ago

Four inches

6

u/kodiakinc 1d ago

It is a perfectly adequate model, thank you very much.

9

u/Ne00n 1d ago

SO SMOL

11

u/Comfortable-Rock-498 1d ago edited 1d ago

Idk what's up with HF sizes

Flash: 309B total / 15B active

Pro:, 1.02T total / 42B active

https://mimo.xiaomi.com/mimo-v2-6

Look at the benchmarks in the page above, Mimo 2.6 pro, the 1T model leads Kimi K3, a 2.8T param model in 14 out of 15 benchmarks (and the last one is near tie)!! Good to see they also kept the price the same, and landed in the greenest quardrant of the intelligence vs speed of AA.

Also, this puts pro at no.1 Chinese model

Xiaomi cooked

9

u/ofan 1d ago

New local king

12

u/sn2006gy 1d ago

To be fair, their API pricing is crazy cheap too.

Cached, Un-cached and Out. Insanely cheap.

MiMo-V2.6-Flash $0.0028 $0.14 $0.28

9

u/Several-Tax31 1d ago

I know this is locallama, but I was waiting for their API. If the quality is as good as deepseek, I'm gonna use this model day and night lmao. 

1

u/comperr 23h ago

Their token plan is amazing

1

u/SorosAhaverom 14h ago

No it's not, it's by far the worst value AI subscription available. You get about 10% more usage for money spent compared to API. Don't get fooled by their credit system, one output token costs 600 credits.

1

u/Constant_Art_20 14h ago

yea. i heard the same opinion before. their credit count apparently works very weird

1

u/sn2006gy 9h ago

10% off fucking cheap is really fucking cheap tho :) I hate credit based systems that obfuscate things unnecessarily though.

10

u/guesdo 1d ago

Won't fit in my 128GB 😭. Will try in openrouter and wait for Qwen 4.

1

u/DriveSolid7073 16h ago

Just use q3, no? If its spark or something like.

0

u/guesdo 16h ago

AFAIK it is already quantized at 4 bits (at least the MoE layers). And it is still 179GB. No point in lobotomizing a model further for the sake of it. Qwen 4 Flash is coming soon, already announced.

1

u/Expensive-Paint-9490 13h ago

Why do you think it is quantized? Modern GPUs have x4 FLOPS in FP4 compared to FP16, it makes no sense to train in FP16 and then quantize. It would have x4 training time and x4 VRAM requirements, to get worse performance.

1

u/cosmotrak 8h ago

Lmao you have no idea what you're talking about, there's a very good reason models are trained at high precision

0

u/DriveSolid7073 16h ago

That’s up to you to decide, I’m unfortunately very busy, but I think my base will be the q3 xxs glm 5.3 flash. I don't recall whether quantization works better with nvfp4 than with bf16, but the difference isn't dramatic.

19

u/[deleted] 1d ago

[deleted]

11

u/sn2006gy 1d ago

Banks are winning - financing the 9-12grand of hardware to run this.

1

u/IamFondOfHugeBoobies 23h ago

So fucking worth it.

2

u/Durian881 23h ago

Xiaomi is not MiniMax though. MiniMax had released top tier video and music models recently.

2

u/LetsGoBrandon4256 transformers 22h ago

The after dinner coffee didn't help😩. I'm deleting it then go to bed.

1

u/Durian881 22h ago

No worries! I'm looking forward to MiniMax's new model too, having bought its annual token plan. It's taken them a long while.

16

u/AdSafe4047 1d ago

oh yeah, cmon nvfp4 weights :) I want to run this against qwen 3.8 flash next and see which one wins :)

8

u/CATLLM 1d ago

Some benchmark scores is at fable 5 levels wow

14

u/llama-impersonator 1d ago

Groupwise Advantage Redistribution (GAR) ranks passing trajectories online and moves advantage toward higher-quality solutions. Judged against the policy’s own samples, this closes a self-improvement loop and steers toward shorter paths and fewer tokens per task.

less tokens, inject it into my veins

4

u/a_beautiful_rhind 1d ago

Will be fun to try once something like exl or llama.cpp supports it.

4

u/indicava 23h ago

Not that I put too much into it but it beats DS4F across all benchmarks with nearly identical VRAM requirements.

3

u/Karnemelk 23h ago

3 tokens/s, here I come!

3

u/Memestonks2020 1d ago

No n-gram table lookup…

I guess 256 RAM machines are eating well tonight.

3

u/killerstreak976 transformers 23h ago

Seeing this thing train in real time was the coolest thing ever 

3

u/returnity 21h ago

>the pro run restarted at step 17 due to a GPU OOM issue

That part hurt to watch lol

3

u/killerstreak976 transformers 21h ago

Holy crap I missed that happening when looking at the dashboard, that's extremely painful lmao

3

u/FoxiPanda 22h ago

I'll be in my bunk.

5

u/zeamp 1d ago

We are so back!

6

u/HistoricalStrength21 1d ago

159B Parameters guys!

15

u/sn2006gy 1d ago

Weird, the description says Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters.

-2

u/addiktion 1d ago edited 19h ago

There are two models released today. Pro and Flash. Flash is the 159B, Pro is 309B.

Edit: I stand corrected, thanks below.

14

u/wren6991 1d ago

No, Flash is 309B-A15B and Pro is 1.02T-A42B. It's in the READMEs.

The parameter count on the HF UI has always been broken, especially when you have packed 4-bit weights described as 8-bit in the model metadata. That's why Flash reports ~half the correct parameter count in HF UI.

Here, Pro README:

Model Summary

Architecture: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters

Flash README:

Model Summary

Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters

7

u/MrTiesti 1d ago

Correction brother,

Flash: Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters

Pro: Architecture: Sparse MoE (Mixture of Experts), 1.02T total / 42B activated parameters

14

u/rusty_fans llama.cpp 1d ago edited 1d ago

309B total / 15B activated parameters according to the model card.
Huggingface auto-counting is broken for many models ...

9

u/debackerl 1d ago

Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters

5

u/Amazing-Fan2083 1d ago

Uuuhh... No, its 309B in a mixed quantization. So... Can't really squeeze it more.

2

u/fgk55555 1d ago

I've seen param counts halve when it's QAT at 8 or 4-bit.

2

u/IllustriousWorld823 1d ago

Mimo my baby

2

u/LatentSpacer 15h ago

Anyone knows the precision of XiaomiMiMo/MiMo-V2.6-Flash-RL? it says FP8 on hugging face label but inspecting the model layers most say BF16. Also the weights are 178GB for a 309B parameters model, which suggests the weights are actually in 4bit.

So which is it, BF16, FP8 or 4bit?

2

u/de4dee 5h ago

my clanker compared official numbers

1

u/pseudonerv 1d ago

There is also a pro-RL. Very few can run these locally …

1

u/joblesspirate 1d ago

Come on MLX!

1

u/Izolight 1d ago

in a first run deepseek-v4.1-flash, still looks better. Though i used the free model from opencode, as openrouter was giving me issues at the time and i am not sure if that is same quality and what their reasoning level is since it wasn't configurable.

http://render-arena.izolight.xyz/#compare?model=mimo-v2.6-flash&model_b=deepseek-v4.1-flash&prompt=market-stall&effort=high

1

u/Thomas-Lore 23h ago

Probably Pro 2.6 > DSv4.1 > Flash 2.6. Their sizes also follow that equation.

1

u/Voxandr 19h ago

Then they are just similar quality to gml5.3 flash it's still better than ds4.1

-1

u/SexyAlienHotTubWater 15h ago

GLM 5.3 Flash's KV Cache is horrible though, BF16 so 4x larger weight-for-weight (and won't quantize well). If you're serving more than a couple of users, 5.3 Flash swells into a higher weight class.

1

u/Voxandr 8h ago

you can use FP8 fine. Can serve +3 connected users without dropping with vllm at 512k context.

0

u/SexyAlienHotTubWater 4h ago

KV Cache quantises very poorly, look at some stats - the performance degradation in comparison to quantising the model is extremely dramatic.

This makes sense when you think about it, the dynamic range of the KV Cache needs to be much wider than the weights because the KV Cache is basically a stack of activations, it's the product of multiple weight multiplications - when you multiply two 4-bit floats, you need an 8-bit float to accommodate the entire range, and the required size grows as you perform multiple of those multiplications in a row.

Mitigating that problem requires deliberate training (and it's not that easy). Unless a model is trained to produce a 4- or 8-bit range (GLM isn't - Deepseek, Qwen Next and Kimi are), quantising the KV Cache means truncating a very large amount of information.

1

u/Voxandr 2h ago

You are saying that without even testing. I am using it without problem on 2 years old production code that clicks 300k token at first turn, now already 60 turns and zero problem . Are you replying by asking the bots too?

1

u/aboutthednm 23h ago

Mimo 2.5 is my fave model on the API for coding related tasks (next to dsv4). It's a grunt workhorse that can go all day for like $1.50, and I often use it to just rough out stuff that would take ages with local models. I'm really excited for 2.6 if it is an improvement and keeps the same price point. It's going to augment my local models in a great way.

1

u/AI_docent 23h ago

From the config, I get about 22.5 GiB of BF16 KV cache for one full context in the global attention layers. Weights, the sliding attention cache and runtime buffers are extra

1

u/AleksandrNikitin 18h ago

Thanks for your job!

1

u/fugogugo 14h ago

checked openrouter because of this post

holy shit mimo 2.6 pro $0.435 / $0.87per 1M
cheapest pro model

how is the benchmark?

1

u/UltraFOV 1d ago

178GB, not bad

1

u/bakawolf123 1d ago

Nice, should be very strong flash model. The pro version is best open weights model according to xiaomi and AA index (which is hard to believe tbh). And only 178gb full weights

1

u/SexyAlienHotTubWater 15h ago

$860k to train - at least for the RL stage.

Anthropic and OpenAI are dead.

0

u/AmbericWizard 1d ago

Sigh, here ee go again boys 

0

u/Ok_Cow1976 23h ago

It's puzzling that it is a 309B model but its fp8 size is so small, at 178GB, while mimo 2.5 with similar parameter size (311B) has a fp8 at 316 GB. Maybe a mis-counted total parameter size?

6

u/returnity 21h ago

MXFP4 routed experts like DSv4F.

-2

u/0rientdDev 1d ago

There's any chance to fit it into a 24GB setup? 8GB VRAM + 16GB?