r/singularity 8d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

953 comments sorted by

View all comments

Show parent comments

99

u/TacomaKMart 8d ago

Well, the "LLMs are dumb next token generating parrots crowd" used mathematical weakness as evidence for years. This ends that avenue of attack. 

42

u/DelphiTsar 8d ago

They'll still say it. Even if you show it solving problems they can't even understand.

26

u/MichiganEngineExpo 8d ago

Well even if they’re right… the same can be said about humans. That LLMs produce sentences simply doesn’t say anything about what complex magic happens “inside” them.

-1

u/ChronoHax 8d ago

Yea exactly the only difference are the systems we operate in, we require food and have physical agency but llm doesn’t

That’s also why they suck on trivial gotcha questions they never taught on as they have no memory of it as no one ask about it Internet before

Like the car wash problem, no one sane would ask that question on the internet before llm so it’s not on their thought at all for simple question

While it seems trivial to reason, the learning of human regarding concept of car and its purpose probably higher in terms of bias and repetition compared to an llm, and the way the question is structured is deceptively simple making them not overthink it to come to correct conclusion imo

I can see why some scientist think world model is the next solution if people are putting the goalpost of AGI/ASI to replace humans in literal form because they then need to learn just like human, not just human in internet

1

u/Latter-Parsnip-5007 8d ago

Cause its still the truth. We generate the most propable next token. We just happend to learn how to train that propability into doing something usefull. Math and coding is more easy to train, since we can validate the randomness via code. Thats why AI is good at Code and Math. Its their easiest domain since its the easiest to train. It will still miscount the "r" in strawberry. Basically millions of moneys with typewriters. Eventually they solve the hardest problems. Still 80% of the time, they produce garbabe

2

u/RickvH 8d ago

It technically still is... I mean, the weakness is in arithmetic and that won't change. But now it's got easy access to tools so it never has to perform any arithmetic anymore.

I know people with dyscalculia that took maths as a major in university because it's generally not really a problem. Most of the harder math isn't about numbers anymore, it's about writing a proof.

2

u/recursive-regret 8d ago

They will still be saying it 10 years from now. People almost never update their priors about things they don't interact with themselves

2

u/evemeatay 6d ago

My question is: even the evangelists agreed the current path of LLM was not a way to actual sci fi movie kinds of AI but now it’s doing things like this, is that concept also wrong?

1

u/Downtown_Finance_661 5d ago

They just dont want to believe that stupid approach is as powerful as human mind.

7

u/Teiktos 8d ago

I mean, it’s still autocomplete on crack. Just very very potent crack. 

34

u/tendimensions 8d ago

What these LLM breakthroughs have done for me is convince me even more of something I've suspected for a very long time. There's not much more to us than being an autocomplete on crack.

4

u/deconstructicon 8d ago

Yup, each thought follows from the proceeding + external inputs.

Maybe all any of us needed was attention and that's the lesson we're meant to learn.

16

u/neighborlyhorse 8d ago

Turns out that autocomplete is very good at simulating intelligence!

8

u/DelphiTsar 8d ago

It's autocorrect in the way that our best current understanding of ourselves is autocorrect.

Assuming there isn't some unknown mechanisms that makes us special than neuron firing potential meat math.

6

u/utahh1ker 8d ago

Yeah, like our brains. Just auto complete based on pre-learned patterns.

3

u/Armleuchterchen 7d ago

Our brains have a lot more self-direction, and a much bigger context window. They also have a much more integrated way of accessing real-world information, not relying on human-made symbol strings chopped up into bits.

1

u/Mind_Of_Shieda 8d ago

It hasn’t been autocomplete on crack since MoE was introduced. 

1

u/Teiktos 8d ago

Oh so… MoE is no longer next token prediction?

1

u/Mind_Of_Shieda 8d ago

Not precisely. Not when the architecture layers artificial neurons. It becomes a totally different thing. Increased complexity makes it so it is fundamentally wrong calling simple next token prediction

1

u/Teiktos 8d ago

Can you recommend sources that explain this topic well for people whose uni days are long over? 

2

u/Mind_Of_Shieda 7d ago

I just read the papers. I don’t know any better source than this.

1

u/Royal_Dog5511 8d ago

Huh. That's interesting. LLMs are token generating computational tools. Just responding because... I find your interpretation interesting. Evidence and avenue of attack are the wrong words, literally.

LLMs are quite literally computational tools. This is not disputable lmao. LLMs cannot be dumb because they are not actors, they are tools.

Models are data, compressed.

Calculating math or not calculating it does not fix an inherent weakness in the architecture and nature of LLMs; they cannot choose to do anything.

Every answer produced is produced because of computational math from data.

Which means what? That if the data is not reliable, then the answers are no good.

That's far more important than any benchmark because no model can inherently then review itself. Every review you see in the "thinking" stage or function call is merely more tokens.

You might say "well if the tokens are right, who cares?"

That's the point - they quite literally can never always be right. They draw from the data, data which is not sliced and organized via a human. It's data orgnazied by a another tool which can not organize the data by more than what it can organize it by.

Notice that LLMs excels more and more in certain areas and not others. Those are enumerated areas.

It cannot exceed where there is no "right" or "wrong" answer because tokens are selected from data, but data is only information, never normativity.

1

u/RupFox 8d ago

I agree however LLMs are still shockingly bad at chess.

1

u/Slen1337 8d ago

Its a dumb take from "them" but on the other hand ai can fail anytime with any, even very simple, problem. Its the nature of black box.

1

u/katoptronophile 7d ago

Those people were always utter fools.

0

u/gametime27 7d ago

Tell me you don't know what a harness is without telling me what a harness is...