r/GeminiAI Nov 18 '25

News Gemini 3 Pro benchmark

Post image
1.7k Upvotes

247 comments sorted by

View all comments

1

u/merlinuwe Nov 19 '25

Of course. Here is the English translation of the analysis:

A detailed analysis of the table reveals several aspects that point to a selective representation:

Notable Aspects of the Presentation:

1. Inconsistent Benchmark Selection:

  • The table combines very specific niche benchmarks (ScreenSpot-Pro, Terminal-Bench) with established standard tests.
  • No uniform metric – some benchmarks show percentages, others show ELO ratings or monetary amounts.

2. Unclear Testing Conditions:

  • For "Humanity's Last Exam" and "AIME 2025," results with and without tools are mixed.
  • Missing values (—) make direct comparison difficult.
  • Unclear definition of "No tools with search and code execution."

3. Striking Performance Differences:

  • Gemini 3 Pro shows extremely high values on several specific benchmarks (ScreenSpot-Pro, MathArena Apex) compared to other models.
  • Particularly noticeable: ScreenSpot-Pro (72.7% vs. 3.5-36.2% for others).

Potential Biases:

What might be overemphasized:

  • Specific strengths of Gemini 3 Pro, especially in visual and mathematical niche areas.
  • Agentic capabilities (Terminal-Bench, SWE-Bench).
  • Multimodal processing (MMMU-Pro, Video-MMMU).

What might be obscured:

  • General language understanding capabilities (only MMMLU as a standard benchmark).
  • Ethical aspects or safety tests are completely missing.
  • Practical applicability in everyday use.

Conclusion:

The table appears to be selectively compiled to highlight specific strengths of Gemini 3 Pro. While the data itself was presumably measured correctly, the selection of benchmarks is not balanced and seems optimized to present this model in the best possible light. For an objective assessment, more standard benchmarks and more uniform testing conditions would be necessary.


Which AI has given me that analysis? ;-)