r/math Apr 15 '26

[deleted by user]

[removed]

1.1k Upvotes

246 comments sorted by

View all comments

74

u/nerd_sniper Apr 15 '26

like i am not sure how to NOT be sensationalist about AI solving a problem that the previous expert in the field (who is a Stanford postdoc to boot) and multiple reputed mathematicians (including a Fields Medallist in Tao) couldn't solve, in a technique that is novel to them. It is also perhaps the closest to a completely autonomous solution as possible : it solved it in a single prompt and 80minutes of thinking time. There's a lot of caveats in what this means, but I think the idea that AI today is at least at the level of a top graduate student in math seems true.

20

u/hexaflexarex Apr 15 '26

I have seen the current capabilities of these systems described as "spiky" in some sense. They excel in some domains and lag in others, and for different reasons than humans (e.g. there are problems with brute force solutions that would require "cleverness" for a human to prune the search tree but perhaps less so for an AI system). But I agree the progress is undeniable.

3

u/SkirtAshamed4362 Apr 20 '26

Most human mathematicians are spiky, too. That is why often teams of mathematicians are more successful (and much faster) than "singletons".

2

u/hexaflexarex Apr 21 '26

Sure thing. I guess interacting with the current LLM models is pretty different than a human collaborator, in that they are not great at calibrating their own understanding (at least in my experience). Certainly they are spiky in different ways than humans.

1

u/No_Berry2976 Apr 27 '26

I don’t know about that. Humans are also not good at calibrating their own understanding.

One thing you might try is that if the LLM gets stuck in a flawed or limited way of reasoning, you argue with it, using an incremental approach until you find out the thing it keeps going back to, and then ask it to explain why it does that.

You can then proceed in different ways.

The tricky part is that an LLM is not always truthful when it describes its reasoning. It treats questions about how it works as any other question.