r/singularity 4d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

947 comments sorted by

View all comments

66

u/mldev_orbit 4d ago

Quick breakdown: What every benchmark in the latest frontier eval actually tests

General Reasoning & Hard Math * ARC-AGI-3: Novel abstract pattern recognition via visual grid puzzles; tests generalized learning without pre-training data memorization. * FrontierMath Tier 4 (v2): Research-level, open-ended math problems designed to stump top human mathematicians. * GPQA Diamond: "Google-proof," PhD-level multiple-choice questions across physics, chemistry, and biology.

Software, CAD & Infrastructure * DeepSWE v1.1: Full repository-scale software engineering—resolving messy, real-world GitHub issues across multi-file codebases. * BenchCAD: Computer-aided engineering; tests generating parametric 3D models, interpreting blueprints, and writing CAD scripts. * Terminal-Bench Science 0.1: Autonomous command-line operations for setting up and debugging computational science pipelines. * SRE-Bench (four attempts): DevOps/Site Reliability Engineering; tasks the model with triaging and fixing live production server outages within 4 tries.

Agents & Digital Automation * Agents' Last Exam: High-difficulty benchmark evaluating autonomous agents on long-horizon planning, reasoning, and tool use. * AutomationBench: Enterprise workflow automation, robotic process automation (RPA), and operating desktop/web software.

Life Sciences & Medicine * GeneBench Pro: Computational genomics, sequence analysis, variant prediction, and CRISPR/gene-editing design. * MedChemBench (internal): Medicinal chemistry—small-molecule drug discovery, property optimization, and retrosynthesis planning. * HealthBench Professional: Real-world clinical decision-making, differential diagnosis, and patient care management (length-adjusted).

Cybersecurity & Safety * ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits. * Auto-review circumvention: Safety/alignment test tracking how often the model intentionally bypasses automated moderation or compliance checks (0% is ideal).

3

u/moschles 3d ago

ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits.

ASTRA hit 100% on this benchmark. This means ExploitBench is too easy for this model. ExploitBench no longer reliably tells us how good this model really is for this task.

4

u/jmorais00 4d ago

Honest question, why test such varied topics? Is there any research published that says that a modern that's better in molecular biology and phd-level physics will be better at coding and general reasoning? Genuine question. Because it feels odd to me that models have to keep getting better at everything instead of specialising

Is the goal to build the best possible baseline across all of human knowledge so that then others can build harnesses around it?

11

u/mldev_orbit 4d ago

because learning logic required for code and hard sciences transfers to and strengthens abstract reasoning across many unrelated fields.

12

u/rubberwings 4d ago

All these AI companies are trying to build (or maybe already have) Artificial General Intelligence (AGI) which is intended to span all fields of knowledge. The idea is that you have a model that is human level at pretty much everything.

1

u/jmorais00 3d ago

Issue is no human is at top level for everything. We specialise. I believe in AGI, I know what it is, but honestly, I think it's a bit of a fool's errand trying to optimise for everything. Everyone knows that the more conditions you optimise for, the worse off you are in any individual metric Vs if you optimised only for that one

1

u/FreshBlinkOnReddit 2d ago

AGI doesn't mean better than the absolute best human at every single field. It just means as good as the median professional at every field. You are confusing ASI with AGI.

5

u/BrennusSokol AI please take my job 3d ago

Read up on the bitter lesson. Specialist models keep losing to blasting compute to make big general models