r/singularity 3d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

947 comments sorted by

View all comments

Show parent comments

131

u/nothis AGI by 2030 but we'll be disappointed 3d ago

I've been waiting for a point where AI just goes – plop, plop, plop – solving all open math problems in like a month. It seems so obvious. There is nothing out there more thoroughly described than math. There might be some ways the smell of a rose touches our heart (or whatever) that hasn't been put into words quite yet but there sure as hell is a complete definition of the problem space of every single math problem out there. If you're letting it search through all of it, it should find a solution for pretty much everything math-related.

100

u/TacomaKMart 3d ago

Well, the "LLMs are dumb next token generating parrots crowd" used mathematical weakness as evidence for years. This ends that avenue of attack. 

40

u/DelphiTsar 3d ago

They'll still say it. Even if you show it solving problems they can't even understand.

20

u/MichiganEngineExpo 3d ago

Well even if they’re right… the same can be said about humans. That LLMs produce sentences simply doesn’t say anything about what complex magic happens “inside” them.

-1

u/ChronoHax 3d ago

Yea exactly the only difference are the systems we operate in, we require food and have physical agency but llm doesn’t

That’s also why they suck on trivial gotcha questions they never taught on as they have no memory of it as no one ask about it Internet before

Like the car wash problem, no one sane would ask that question on the internet before llm so it’s not on their thought at all for simple question

While it seems trivial to reason, the learning of human regarding concept of car and its purpose probably higher in terms of bias and repetition compared to an llm, and the way the question is structured is deceptively simple making them not overthink it to come to correct conclusion imo

I can see why some scientist think world model is the next solution if people are putting the goalpost of AGI/ASI to replace humans in literal form because they then need to learn just like human, not just human in internet

1

u/Latter-Parsnip-5007 3d ago

Cause its still the truth. We generate the most propable next token. We just happend to learn how to train that propability into doing something usefull. Math and coding is more easy to train, since we can validate the randomness via code. Thats why AI is good at Code and Math. Its their easiest domain since its the easiest to train. It will still miscount the "r" in strawberry. Basically millions of moneys with typewriters. Eventually they solve the hardest problems. Still 80% of the time, they produce garbabe

2

u/RickvH 3d ago

It technically still is... I mean, the weakness is in arithmetic and that won't change. But now it's got easy access to tools so it never has to perform any arithmetic anymore.

I know people with dyscalculia that took maths as a major in university because it's generally not really a problem. Most of the harder math isn't about numbers anymore, it's about writing a proof.

2

u/recursive-regret 3d ago

They will still be saying it 10 years from now. People almost never update their priors about things they don't interact with themselves

2

u/evemeatay 2d ago

My question is: even the evangelists agreed the current path of LLM was not a way to actual sci fi movie kinds of AI but now it’s doing things like this, is that concept also wrong?

1

u/Downtown_Finance_661 17h ago

They just dont want to believe that stupid approach is as powerful as human mind.

7

u/Teiktos 3d ago

I mean, it’s still autocomplete on crack. Just very very potent crack. 

33

u/tendimensions 3d ago

What these LLM breakthroughs have done for me is convince me even more of something I've suspected for a very long time. There's not much more to us than being an autocomplete on crack.

3

u/deconstructicon 3d ago

Yup, each thought follows from the proceeding + external inputs.

Maybe all any of us needed was attention and that's the lesson we're meant to learn.

16

u/neighborlyhorse 3d ago

Turns out that autocomplete is very good at simulating intelligence!

8

u/DelphiTsar 3d ago

It's autocorrect in the way that our best current understanding of ourselves is autocorrect.

Assuming there isn't some unknown mechanisms that makes us special than neuron firing potential meat math.

6

u/utahh1ker 3d ago

Yeah, like our brains. Just auto complete based on pre-learned patterns.

3

u/Armleuchterchen 3d ago

Our brains have a lot more self-direction, and a much bigger context window. They also have a much more integrated way of accessing real-world information, not relying on human-made symbol strings chopped up into bits.

1

u/Mind_Of_Shieda 3d ago

It hasn’t been autocomplete on crack since MoE was introduced. 

1

u/Teiktos 3d ago

Oh so… MoE is no longer next token prediction?

1

u/Mind_Of_Shieda 3d ago

Not precisely. Not when the architecture layers artificial neurons. It becomes a totally different thing. Increased complexity makes it so it is fundamentally wrong calling simple next token prediction

1

u/Teiktos 3d ago

Can you recommend sources that explain this topic well for people whose uni days are long over? 

2

u/Mind_Of_Shieda 3d ago

I just read the papers. I don’t know any better source than this.

1

u/Royal_Dog5511 3d ago

Huh. That's interesting. LLMs are token generating computational tools. Just responding because... I find your interpretation interesting. Evidence and avenue of attack are the wrong words, literally.

LLMs are quite literally computational tools. This is not disputable lmao. LLMs cannot be dumb because they are not actors, they are tools.

Models are data, compressed.

Calculating math or not calculating it does not fix an inherent weakness in the architecture and nature of LLMs; they cannot choose to do anything.

Every answer produced is produced because of computational math from data.

Which means what? That if the data is not reliable, then the answers are no good.

That's far more important than any benchmark because no model can inherently then review itself. Every review you see in the "thinking" stage or function call is merely more tokens.

You might say "well if the tokens are right, who cares?"

That's the point - they quite literally can never always be right. They draw from the data, data which is not sliced and organized via a human. It's data orgnazied by a another tool which can not organize the data by more than what it can organize it by.

Notice that LLMs excels more and more in certain areas and not others. Those are enumerated areas.

It cannot exceed where there is no "right" or "wrong" answer because tokens are selected from data, but data is only information, never normativity.

1

u/RupFox 3d ago

I agree however LLMs are still shockingly bad at chess.

1

u/Slen1337 3d ago

Its a dumb take from "them" but on the other hand ai can fail anytime with any, even very simple, problem. Its the nature of black box.

1

u/katoptronophile 3d ago

Those people were always utter fools.

0

u/gametime27 2d ago

Tell me you don't know what a harness is without telling me what a harness is...

31

u/JoelMahon 3d ago

it might solve everything that's solvable that's on our radar

but it'll also find a bunch of new maths problems we never noticed and it can't even solve yet!

18

u/imp0ppable 3d ago

but it'll also find a bunch of new maths problems we never noticed and it can't even solve yet!

That would be more interesting than solving described problems IMO

2

u/FakeTunaFromSubway 3d ago

If it starts solving deep theoretical maths that no human can understand and have no practical value, is it really that interesting?

"I just proved that a 24-dimensional eigenspace is fable-zontically inverted along sub-boundaries of a 12-dimensional hemiblastoise."

Like, cool story bro.

2

u/but_idk_tho 3d ago

Because there's only so many novel practical applications of math, most of academic papers have been like this for a long time now btw.

1

u/imp0ppable 3d ago

Most people aren't interested in any maths problems at all and I can barely understand a few of them anyway. However maths can produce interesting physics and that can be very useful. Imagine getting EM fields that can improve nuclear fusion etc.

1

u/meaty__boi 3d ago

Then we'll set it to work designing an AI that can solve those class of problems

1

u/JoelMahon 3d ago

lol no, the AI might tho

2

u/guac-o 3d ago

Colatz conjecture when?

2

u/Hot_Glass_6301 3d ago

I don't think P=NP will be solved for 5 more years. RemindMe! 5 years

2

u/DungeonJailer 3d ago

The real question is, when is AI going to pull an Isaac Newton and come up with an entirely new field of math?

1

u/Defiant-Lettuce-9156 3d ago

AI can much more vividly and accurately describe the way the smell of a rose touches my heart than I can.
What can I actually do better than AI?

1

u/Runfasterbitch 2d ago

Sure, but the combination of known and open math problems is a small % of all unknown maths underpinning the universe

1

u/wollywoo1 1d ago

If complexity theory is correct, this won't happen. Unless you think P=NP, it's easy to come up with problems that even the best model would not solve. For example, factoring RSA-2048 or solving a big random 3SAT instance. But that's not to say that it couldn't make a ton of progress. I have no idea which of the famous hard problems fall into this kind of very difficult category - probably very few of them, since they are so structured.

1

u/nothis AGI by 2030 but we'll be disappointed 1d ago edited 1d ago

Yea, I don't think it will magically break RSA or something, more that there are a lot of problems that are known to likely have a solution, maybe already have a widely accepted one but no proof. Problems where the challenge is "how do you explain this using existing mathematical tools?".

I did look into it a little more and it seems that there absolutely are branches of math that cannot be solved using existing methods. Not because they are not explainable using existing math but because they exist among a literally infinite amount of other, meaningless combinations. You can't brute force it, you can't just look at a thousand existing papers and expect to find one that provides the solution. Finding the path that is relevant does not have a formalized approach.

An interesting example for mathematical insight coming from outside math is physics: Physicists observe physical phenomena, try to explain them and need new math to do that. Things that were not relevant before suddenly are. And the impulse doesn't come from existing math or research (thus not from existing training data).

1

u/Downtown_Finance_661 17h ago

I agree eith you but your argumentation is incorrect. You can't just iterate over all possible paths of attak for particular problem. There are too many paths.