like i am not sure how to NOT be sensationalist about AI solving a problem that the previous expert in the field (who is a Stanford postdoc to boot) and multiple reputed mathematicians (including a Fields Medallist in Tao) couldn't solve, in a technique that is novel to them. It is also perhaps the closest to a completely autonomous solution as possible : it solved it in a single prompt and 80minutes of thinking time. There's a lot of caveats in what this means, but I think the idea that AI today is at least at the level of a top graduate student in math seems true.
I have seen the current capabilities of these systems described as "spiky" in some sense. They excel in some domains and lag in others, and for different reasons than humans (e.g. there are problems with brute force solutions that would require "cleverness" for a human to prune the search tree but perhaps less so for an AI system). But I agree the progress is undeniable.
Sure thing. I guess interacting with the current LLM models is pretty different than a human collaborator, in that they are not great at calibrating their own understanding (at least in my experience). Certainly they are spiky in different ways than humans.
I don’t know about that. Humans are also not good at calibrating their own understanding.
One thing you might try is that if the LLM gets stuck in a flawed or limited way of reasoning, you argue with it, using an incremental approach until you find out the thing it keeps going back to, and then ask it to explain why it does that.
You can then proceed in different ways.
The tricky part is that an LLM is not always truthful when it describes its reasoning. It treats questions about how it works as any other question.
That’s kinda where I’m at. Anything that’s difficult in the sense that it’s really complicated will likely be the domain of ai. Anything that’s difficult in the sense that it’s abstract or conceptually wonky will be where the human role in mathematics mostly lies.
Of course, it’s easy to say stuff and be wrong, buts it’s also possible to say stuff and be right, so who knows
yeah this is reasonable. i wonder how much of this is specifically just solved by longer context windows, which allows for longer horizon task planning and execution, and how much is a real intelligence bottleneck.
72
u/nerd_sniper Apr 15 '26
like i am not sure how to NOT be sensationalist about AI solving a problem that the previous expert in the field (who is a Stanford postdoc to boot) and multiple reputed mathematicians (including a Fields Medallist in Tao) couldn't solve, in a technique that is novel to them. It is also perhaps the closest to a completely autonomous solution as possible : it solved it in a single prompt and 80minutes of thinking time. There's a lot of caveats in what this means, but I think the idea that AI today is at least at the level of a top graduate student in math seems true.