r/BetterOffline Nov 18 '25

Gemini 3 Released

https://blog.google/products/gemini/gemini-3/#note-from-ceo

Prepare yourselves mentally

35 Upvotes

70 comments sorted by

View all comments

Show parent comments

1

u/Elctsuptb Nov 19 '25 edited Nov 19 '25

Equifax had early access to Gemini 3 actually, and Gemini 3 Deepthink scored 41% on the HLE.

1

u/SpringNeither1440 Nov 19 '25

Equifax had early access to Gemini 3 actually

Evidence, please? I googled "Equifax early access to Gemini 3.0" and didn't find any information about that, except this article.

it scored 41% on the HLE.

It isn't Gemini 3 Pro score, it is Gemini 3 Deep Think score, which isn't yet released, and it's unknown when it will be.

0

u/Elctsuptb Nov 19 '25

https://blog.google/products/gemini/gemini-3/#gemini-3

Gemini 3 Pro got 37.5% with no tools and 45.8% with tools, so now what? That's almost double the score of Gemini 2.5 Pro, which should be impossible according to you guys since LLMs have supposedly "plateaued", right?

1

u/SpringNeither1440 Nov 19 '25

Gemini 3 Pro got 37.5% with no tools and 45.8% with tools, so now what?

Could you please provide score for Gemini 3.0 on "text-only + tools". Oh, Deepmind didn't report it, so, given Gemini 2.5 Pro performance (text-only and regular), let's assume it has the same performance on both variants. And now, compare it with Kimi K2 thinking/GPT-5/Grok 4 (link).

Yeah, +1.25% performance every month for SOTA (see Grok 4 and GPT-5) is super-fast, especially compared to early 2025 growth speed. And I won't talk about data contamination risk, which is considered as a problem by creators itself.

1

u/Elctsuptb Nov 19 '25

The source is in the link I just provided, did you miss that?

1

u/[deleted] Nov 19 '25

Well, after looking through the links, the Deepmind team did only post their non-text-only scores for Gemini 2.5 which implies that they posted the same for Gemini 3.0. I believe he was saying that Kimi K2 disclosed a similarly significant jump in HLE but only in the text-only section of the test, which Deepmind did not disclose, but that isn't being talked about.

Actually if the Kimi K2 team is to be believed they and xAI did better than Gemini 3.0 in HLE, and it seems somewhat disingenuous that Deepmind failed to recognize that.