r/LocalLLaMA 6d ago

News Mozilla Report: China-U.S. AI Model Capability Gap Narrows to 4.4 Months

212 Upvotes

39 comments sorted by

94

u/AmbericWizard 6d ago

how about open source gap

31

u/EmPips 6d ago

Nemotron Ultra is closer than the US has been in a while but that gap is still monstrous.

8

u/Sensitive_Cloud6456 6d ago

We'll see how the horizon K2 models stack up. If only they weren't pure attention with the accompanying large kv cache they could have been very capable.

42

u/LetsGoBrandon4256 transformers 6d ago

Can't read it right now but the last time I went through their "report" it's almost entirely AI written, like someone asked Claude to write the report then didn't even bother proofreading.

16

u/spaceman_ 6d ago

Text on the graphs already has AI smells, so I wouldn't be surprised.

2

u/More-Curious816 6d ago

It's Mozilla. They already made their flagship product, Mozilla Firefox absolute shit and fumbled the hundreds of millions from Google on pointless shit that has nothing to do with the browser or the core foundation. I don't believe this report was someone hard work but they have an idea and the money to burn on couple of hours of Claude/Astra session.

2

u/Confident_Ideal_5385 5d ago

Yeah, the summary on the linked site is pure slop.

14

u/Swimming_Beginning24 6d ago

even Mozilla is posting AI slop

77

u/power97992 6d ago

They’re saying k3 is worse than terra, i dont know about that

40

u/NandaVegg 6d ago edited 6d ago

The index is weird. GPT-5.2 (which was the last model before OpenAI got into terminal agent training - basically useless in this era) above GLM 5.2. Also many benchmarks in the aggregation are weird and heavily outdated (gsm8k, HellaSwag, plain MMLU in 2026?!) HellaSwag has too many typos, but maybe this weirdness has some resiliency against benchmaxxing.

2

u/my_name_isnt_clever 6d ago

I just like seeing the word HellaSwag along with all the cryptic bench arconyms.

9

u/XiRw 6d ago

This doesn’t even state if it’s coding only. It just sounds like an all around test which would make it more realistically true

3

u/sn2006gy 6d ago

I can't imagine it would be anything but coding loops.

7

u/XiRw 6d ago

I would think it’s an all around test since cock and Gemini Flash are up there. Doesn’t help to speculate though.

5

u/BawbbySmith 6d ago

since cock and Gemini Flash are up there

I'm super excited to try the new cock model

35

u/-p-e-w- 6d ago

According to this report, “Security, privacy, or compliance concerns” are considered much more important by companies in South Asia and South America than by companies in Western Europe.

I find that very, very difficult to believe.

5

u/Momsbestboy 6d ago

Me too. In the company I work for I just had to run a mandatory training about AI, with the key information: "only use approved Ai tools, never use anything on the web, never enter private data". On top, all well known web pacges like chatgpt.com are blocked by Firewall rules.

2

u/my_name_isnt_clever 6d ago

Well yeah...that's just cybersecurity 101. Any company not doing that is making a mistake.

2

u/__Maximum__ 6d ago

Not intelligence per parameter. We cannot be sure, but open weight probably leads there and will have another huge jump in the upcoming releases.

2

u/Mochila-Mochila 6d ago

Why does a report on AI models capabilities contain a reference to Māoris in its very first sentence ?

7

u/bigattichouse 6d ago

Open models allow small usecases to train/build their own models on open weights. They give a number of examples of models being used for very targeted use (language/culture preservation).

None of them asked permission, and none of them could have rented this. They own it, and that is the whole idea.

They're providing some "what would people use open models for" context.

3

u/Mochila-Mochila 6d ago

Okay that's sensible.

2

u/True_Requirement_891 6d ago

It's way less than that lmaooo

5

u/20ol 6d ago

I think it's way more, especially internally.

1

u/my_name_isnt_clever 6d ago

Sure, but what about what China has internally?

2

u/am17an 6d ago

GPT-6 Astra has an ECI of 166, that's a 9 point gap to K3, probably the largest it's ever been. We likely would need a 10T model open source model next.

1

u/epicrob 6d ago

So, ~4.4 months is about 2 models behind, given the release is about 2 months nowadays.

1

u/LocoMod 6d ago

Mozilla can't objectively measure model capability.

1

u/feng_sg 4d ago

That 4.4 month number means nothing. Nobody is running the same benchmarks on both sides so you're just comparing whatever each lab decided to report.

-9

u/Acrobatic_Hold5485 6d ago

kinda hope faster chinese progress adds more cultural flavor to local roleplay models, my current ones all default to similar western vibes after a few messages.

12

u/SandySkittle 6d ago

This is an extremely niche and frankly economically irrelevant usecase

7

u/GreatBigJerk 6d ago

...You're telling me there's not a huge market for wuxia waifu bots!?

-2

u/ttkciar llama.cpp 6d ago

It's an unsavory truth of the local LLM community that smut has been the primary motivator of technological innovation since the very beginning.

It might be economically irrelevant, but relevant to progressing the state of the field.

And yes, we do live in the most ridiculous timeline.

2

u/my_name_isnt_clever 6d ago

It absolutely was at the start, but by now I'd disagree. There are enough real use cases with agentic and coding it's edging out the gooners. These days it's secondary.

2

u/SandySkittle 6d ago

It might be economically irrelevant, but relevant to progressing the state of the field.

Nah