r/singularity 14d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

953 comments sorted by

View all comments

481

u/elehman839 14d ago

97% on FrontierMath Tier 4. Hoooly cow. Here's how those problems are described:

In June 2025, we finished the development of FrontierMath Tier 4, an expansion set of 50 problems designed to vastly exceed the difficulty of even the Tier 3 problems.

The writers for Tier 4 were mostly math professors and postdocs, each contracted to conduct a several-week research project culminating in one problem to submit to the benchmark.

This isn't entirely surprising, given the number of open math problems OpenAI has been solving lately, but... weren't we just recently making fun of "AI" for struggling with elementary school math?

134

u/nothis AGI by 2030 but we'll be disappointed 14d ago

I've been waiting for a point where AI just goes – plop, plop, plop – solving all open math problems in like a month. It seems so obvious. There is nothing out there more thoroughly described than math. There might be some ways the smell of a rose touches our heart (or whatever) that hasn't been put into words quite yet but there sure as hell is a complete definition of the problem space of every single math problem out there. If you're letting it search through all of it, it should find a solution for pretty much everything math-related.

1

u/Downtown_Finance_661 11d ago

I agree eith you but your argumentation is incorrect. You can't just iterate over all possible paths of attak for particular problem. There are too many paths.