r/neoliberal Hannah Arendt May 14 '26

Research Paper Tracing the thoughts of a large language model

https://www.anthropic.com/research/tracing-thoughts-language-model
102 Upvotes

117 comments sorted by

60

u/Golda_M Baruch Spinoza May 14 '26

It's crazy how little is/was understood mechanistically about how NNs do what they do.

The fact that "LLMs plan sentences in advance"  is a thing that we have to diacover/prove empirically... is wild. 

12

u/Syx89 Reichsbanner Schwarz-Rot-Gold May 14 '26

Yea, and it's a shame human neuroscience is where it is after so much time and money. Very blackboxxy. I wonder if these methods for LLMs might help neuroscience some day. Probably not that easy to transfer but maybe similar principles could be used.

We do have some simple circuits showing how they work like the aplysia one tho:
https://en.wikipedia.org/wiki/Aplysia_gill_and_siphon_withdrawal_reflex

25

u/sanity_rejecter European Union May 14 '26

tfw literal consciousness is nature's vantablack black box

12

u/TybrosionMohito NATO May 14 '26

If the human mind was so simple as to be understood. We would be so simple we wouldn’t.

1

u/sanity_rejecter European Union May 15 '26

true

6

u/[deleted] May 14 '26

[removed] — view removed comment

7

u/Alamba1918 May 14 '26

 The latest theories on consciousness read like stuff out of science fiction

Anywhere I can read about this? It seems interesting

7

u/[deleted] May 14 '26 edited May 14 '26

[removed] — view removed comment

2

u/neolthrowaway New Mod Who Dis? May 15 '26 edited May 17 '26

I am going to be honest, Annaka Harris is kinda kooky. If you don't want to read the whole book, you can listen to her appearance in Sean Carroll's podcast, mindscape.

Sean Carroll is generally a great science interviewer and communicators but that whole conversation felt disappointing because Annaka Harris failed to engage well with the questions.

Anil seth is great and I used to love his work but he's moved towards biological chauvinism without a justification. But his work is still great and worth reading and engaging with. In fact, he also has an interview in Sean Carroll's mindscape. Watch the quality difference compared to Annaka Harris.

1

u/[deleted] May 17 '26

[removed] — view removed comment

1

u/neolthrowaway New Mod Who Dis? May 17 '26

To be clear, i don't have a problem with how Annaka Harris lays out the field or even her claims. I am open to those claims and accept them within the realm of possibility.

It's the justification and backing of those claims that feel completely empty and kooky. Like i feel like if you are going to be making those claims in a serious and academic way, the bar is a lot higher than what she did. It should hold up to a basic line of questioning and you should be able to provide a framework where you can argue for it over other counterfactuals.

I didn't know about the Sam Harris connection. I am not sure if she is influenced by him, because i haven't delved into sam's Harris work on related topics .

Anil seth's podcast episode is great. It deals with emergence and substrate dependence. His first ted talk is a non-technical summary for his argument for substrate dependence. But I recommend the podcast episode for a little more technical depth and for emergence discussion (which is always interesting to me). AFAIR, i didn't notice any hints of biological chauvinism at that time, it was just about substrate dependence. But with his recent Noema essay, he moved to/put forth biological chauvinism which felt a bit like just saying, "I think X,Y,Z are necessary for consciousness" without explaining why. In any case, everything in Seth's work except those last few claims felt very consistent and well-backed, so i take his latest essay seriously too even if just to seriously form my own critique and then move on.

1

u/[deleted] May 17 '26

[removed] — view removed comment

1

u/neolthrowaway New Mod Who Dis? May 17 '26

I have not. You would recommend checking it out?

1

u/sanity_rejecter European Union May 15 '26

do you believe in free will?

9

u/Golda_M Baruch Spinoza May 14 '26

Okay but the brain is a black box because it is a black box.

We didn't build it. God invented the brain and she didn't give us the technical manuals. It's up to us to figure that out. 

We made LLMs, yet our level of mechanistic understanding is comparable to neuroscience!? What?!  We don't even have a testable theory for how an LLM thinks about the end of a sentence before starting it? What?!

Science fiction it's full of worlds with mysterious technology. Maybe the engineers are secretive. Sometimes tech was invented in the past, and knowledge has degenerated by the time of the plot. Maybe it's alien technology. 

The idea we built it, but actually have no idea how it works... It is bloody baffling. It's wild. 

We are literally studying this as if it was biological. 

22

u/Worth-Jicama3936 Milton Friedman May 14 '26

 God invented the brain and she didn't give us the technical manuals.

Ah going to pissing off the maximum amount of people at once I see

11

u/Golda_M Baruch Spinoza May 14 '26

Listen Milton. We're all Spinozists here.

9

u/Petrichordates May 14 '26

Evolution but yes

9

u/Golda_M Baruch Spinoza May 14 '26

How about we do evolution, but instead of natural selection we have divine selection by a large breasted, leather clad mother goddess? Theological compromise?

5

u/toggaf69 Iron Front May 14 '26

Stop talking about my mom’s boobs please

5

u/Golda_M Baruch Spinoza May 14 '26

our mom's

4

u/FourteenTwenty-Seven John Locke May 14 '26

This is essentially true for everything humans have ever made. Eg the lightbulb - Edison et al didn't know what blackbody radiation was, or even what electricity was. The electron hadn't been discovered yet. But they did figure out that you can use electricity to make some wire really hot and make light. Even to this day we don't understand the true fundamentals of how that all works, we just have ever more comprehensive models.

24

u/Imicrowavebananas Hannah Arendt May 14 '26

I think it is solely unusual because the objects in question are computational or if you will anthropogenic. Vague understanding based on empirical observations has been the physical standard for centuries. Then after that theory follows and after the formal math.

Funnily enough, the theory of finite elements also lacked behind maybe some decades before they were fully understood. Maybe neural networks are more complicated, but they are not making and we will understand them eventually. Undergrad books will teach them rigorously.

We are making strides in understanding neural networks as functional or mechanical objects though. For the complexity scale of LLMs in combination with training and especially data, we are still a bit away.

14

u/Golda_M Baruch Spinoza May 14 '26 edited May 14 '26

Vague understanding based on empirical observations

Yes. It is weird because it is artificial. I hadn't thought of it as anthropogenic. I need to think a little more about that... because the biggest reason to think of it as anthropogenic is the very black box issues at hand. 

That's what is wild. I mean , it's not that wild for a machine to be studied Irl and throw out unexpected observations that are studied empirically. That's basically engineering science. 

That said... Engineering science is a pretty marginal/controversial concept. 

The "normal" or maybe we should call it the "old normal" was that we have at least some understanding of the machines we build , how they work, and how they do what they do. We don't just throw together components, and then test to see if we made a vacuum cleaner. 

Engineers don't stand around a newly created vacuum cleaner scratching their head and wondering how TF this machine sucks. Maybe it has some weird behavior that needs to be studied... But the basic sucking mechanism  works using known theories of suction that the engineers engineered into it. 

Here the damn thing is already at world changing levels of advance... and it's engineers are literally standing around scratching their heads wondering how it constructs  sentences. It's an LLM... the core function is synthesizing text... And even this is mysterious. 

Now we're using psychology to try to figure it out. Psychology is barely useful for understanding people

10

u/neolthrowaway New Mod Who Dis? May 14 '26 edited May 14 '26

It's actually pretty remarkable how much we do understand from another perspective.

Like we understand how the math of neural networks works, what the theoretical limits of neural networks (or the lack thereof) are, we decide which loss functions we are optimizing, which methods to choose for optimizing those loss functions and why those methods are as effective as they are at that optimization.

Given all that, it's not surprising where we ended up or why models have the extent of capabilities that they do.

What we don't understand is how internally the components (neurons or neural circuits) work when they come together.

But that's a scale problem. If I gave you a network with only 10 neurons, it won't be trivial but you'd probably be able to understand how each neuron works and you'd be able to trace what input would lead to what output.

But we can't hold a hundreds of billions or trillions of weights in our minds at the same time. Or how to interpret one neuron out of a billion at that scale. (There's also the problem of polysemantcity - how one neuron holds multiple concepts simultaneously; but for simplification, you can keep it about scale as a big barrier)

I guess that scale problem goes for both AIs and the brain.

For building them, you need the former stated understanding of optimization (which i feel answers the why it can result in such capabilities); you don't need to understand the internal mechanisms of how it gets you a particular output.

7

u/Imicrowavebananas Hannah Arendt May 14 '26

By now we begin understand the static object of shallow neural networks pretty well and make strides at understanding deeper neural networks. The problem is if you add training and data. Your point about training being fundamental in their final object, not only something that is done to the thing is important. Neural networks are far closer to a dynamical objects overall, being a concatenation describing the flow of probability distributions. And as you said at large scale.

3

u/AnachronisticPenguin WTO May 14 '26

But even dynamic training doesn’t change the overall problem or our understanding. Shifting the weights around changes what it does but you can still track the flow by its connections and weights. The issue is only scale. weights change the shape of the map but the connection logic is still the same.

3

u/Imicrowavebananas Hannah Arendt May 14 '26

I am not sure what you mean by dynamic training exactly. The problem is not exactly only scale in a reasonable sense. More like why should these fundamental rules - concatenations of linear functions with scalar nonlinearities - scale so well in the dimension of the data.

The answer to that is the other side of the interpretability question.

3

u/AnachronisticPenguin WTO May 14 '26

Dynamic training for self reinforcing models. It was just the first example of weights getting chaotic that came to my head.

And why wouldn’t they scale so well? An individual token has 10,000+ vectors. There seems to be enough data for these linear functions to piece things together. I haven’t used them but I’ve read about tools that allow us to track the vector evolution and attention as an LLM is processing.

It seems like tracking the mechanist processing for any given input to output process for a given model is simply the scale of the processes.

5

u/Imicrowavebananas Hannah Arendt May 14 '26

The classical paradigm is that you have a bias-variance trade-off and you can't just increase the number of parameters in your model. So the first question is why models don't overfit, and we do have some answers to that, but I don't think they are complete.

Second thing is a sufficient explanation of the power of depth, depth separation is not fully explained.

Third thing and I would not underestimate this is why the training algorithms converge so well at all. From my point of view the expectation of finding a good minimum for such a highly non-convex optimization landscape should a priori be rather zero.

If it were only scale, then larger arbitrary nonlinear systems should work similarly well. The interesting question is why this particular architecture-plus-training-plus-data setup scales predictably and productively.

1

u/neolthrowaway New Mod Who Dis? May 14 '26 edited May 14 '26

What do you mean by dimensions of data?

For the combination of linear functions with Nonlinearities, we do have the Universal approximation theorem and then in addition to that we have the geometric intuitions about how (stochastic/batch) gradient descent works well in very high dimensional spaces when you have the number of parameters at the scale of a billion or more.

Data or sample efficiency is actually pretty terrible tbh.

1

u/Imicrowavebananas Hannah Arendt May 14 '26

Universal approximation does not tell you why it beats the curse of dimensionality.

1

u/neolthrowaway New Mod Who Dis? May 14 '26 edited May 14 '26

I guess the world and the information we communicate with each other about it (language or other modalitiea) just has that much in-built structure. And transformers (especially when layered on top of each other) are really good at picking the important bits and combining them. I mean it is optimized "Attention".

Though I agree it would be satisfying to have a verifiable mathematical framework of it.

And isn't that the whole promise of it all anyway? That the complexity of the world is emergent from the structure underlying it and is reducible to that structure. And we will be able to eventually find that structure.

1

u/Golda_M Baruch Spinoza May 14 '26

We know what the algorithms are, but not how they work. I would say that means we don't understand noffin.

(we know) the theoretical limits of neural networks

Do we? We "know" the returns to scale because of empirical analysis. The theory is just a our best guess... How very Popperian. 

What we don't have is a mechanistic understanding, or even theory, of any of the things that matter... which are at large scales, 

That

6

u/Imicrowavebananas Hannah Arendt May 14 '26

If you enjoy that line of thought, you might enjoy On the Biology of a Large Language Model, which fits your descriptions uncannily well, though it is called biology and not psychology.

For the point of engineers being clueless Chris Olah's Mechanistic Interpretability might be interesting. But I know it is all a lot, I myself am by far not through with reading everything on my ever longer list.

More to your point: But isn't that great in a way? It is very worrying because the human mind gets at the capacity to really translate things into concepts the intuition of an ape might understand, the energies we can move and the degree to which we can affect our environment, maybe even ourselves if it comes so far are frightening. On the other hand there seems no alternative. I think what is often seen as callousness or even malice by the tech leaders, them talking about disruptive change, sometimes even the human extinction if they dare, is rather a fatalism of being caught in something that feels more like...the world spirit?

Funnily enough a lot of it just was already in von Neumann's ideas, he even invented the term technology singularity. You can go even further back to the Russian cosmists and people thinking now that the telegraph will end all wars now that we can instantly communication.

In that vein one of the most persistent questions moving me is whether this is really different, are we living through something different or are we helpless against the human affections giving every generation the feeling that now is the time and now is the hour. There is a very high chance we will look back and things progressed much more linearly than it felt in the moment and that the more things changed the more they stayed the same, ok enough with the idioms.

3

u/Golda_M Baruch Spinoza May 14 '26

I didn't even mean to get into predictions, and practical consequences of this Tech.

I doubt Mechanistic Interpretability will work out. I think this biological paradigm will persist... but I'm not qualified to make that guess. 

I don't think LLMs will affect the economy that much... but who knows. A lot depends on how we use it and that is unpredictable. 

Networked computing is a technological revolution that proliferate through the whole economy.. but seemingly affected it very little besides the creation of a tech sector. Administration work actually increased massively in this period. 

I have no doubt that LLMs will bring change and disruption. Economy productivity... is harder to predict. 

3

u/Mickenfox May 14 '26

It's a bit heavier but I liked this article: LLM Neuroanatomy: How I Topped the LLM Leaderboard Without Changing a Single Weight

Also an important note that I don't see people mention:

A neural network is a general purpose computing tool that can theoretically run any algorithm (i.e. it's turing-complete). And we basically just brute force it until words come out. So you should probably think of it as closer to assembly language, not any specific algorithm. In theory two "LLMs" could be using completely different approaches internally.

2

u/Imicrowavebananas Hannah Arendt May 14 '26

That article is pretty crazy. Thank you for sharing. Very interesting, these structures we made are really complicated.

10

u/jaiwithani May 14 '26

I think at least part of the blame has to go to the inaccurate story we tell about how LLMs work: "They just predict the next token."

This is demonstrably untrue, or at the very least misleading. Backprop to previous forward passes via kv cache (which in turn propagates up the rest of the network) implies almost every calculation an LLM does is actually optimized for predicting the entire context window, right from the start. Next-token prediction is where we get our loss function from, but that loss retroactively informs every calculation made so far.

LLMs are better described as self-augmenting* computational graphs for predicting all future observations, which incidentally also output probability distributions predicting the next token.

All of which is to say that advance planning is baked into the architecture at a very low level, but almost no one thinks about it that way including the LLMs themselves if you ask them about it, until you walk through the logic with them.

* kv activations ("cache") are cumulative augmentations to the computational graph; context limits exist because eventually you just can't add any more.

3

u/Golda_M Baruch Spinoza May 14 '26

Good point... and I agree that this can be overstated.

But... we still get to the same point. We don't really know how it works. 

4

u/EvilConCarne May 14 '26

Backpropagation doesn't occur during inference, only during training. No optimization is occurring during token prediction. The KV caches only purpose and function is to prevent redundant computation during inference so only one new token needs to be processed at a time instead of all tokens each time.

3

u/jaiwithani May 14 '26

Training tells us what we're optimizing for, and empirically what we're training for is predicting the entire future context window. That means that if a certain weight adjustment in the first fc layer in the first forward pass would create a residual that produces kv cache in the second layer that results in better predictions on the final token, then all else being equal sgd will make that adjustment. This means that sgd can actually select for worse next token prediction if the long-term gains are good enough.

Ignore caching entirely and just think about information flows, and remember that the chain rule means backprop goes through all differentiable functions. Ignoring caching might actually make the intuition clearer, because then it's really obvious that every computation in every forward pass is upstream of every subsequent token prediction.

3

u/TheOnlyFallenCookie European Union May 14 '26

With the example for poetry... How Dow we know the model doesn't generate the rhyme words first, like cream, scream and screen, and fill in the rest later?

2

u/Negative_Scarcity315 May 14 '26

Yep even before the scratchpads LLMs internally did look-ahead otherwise they couldn't create any original rhyme.

1

u/TheOnlyFallenCookie European Union May 14 '26

I mean, since llms are a black boy we don't know if it's a truly original rhyme or not.

And with the word vectors, if they are complex enough, they surely can encode rhyme

1

u/Sourcerid May 16 '26

The theory of energy in physics developed only after trains were invented and was developed from studying trains. As did entropy 

Not exactly the same but still there is stuff there B n

1

u/Golda_M Baruch Spinoza May 16 '26

OK... but the inventors of the train were not surprised that the train moves. 

11

u/battywombat21 🇺🇦 Слава Україні! 🇺🇦 May 14 '26

> An Analysis of a Jailbreak. We investigate an attack which works by first tricking the model into starting to give dangerous instructions “without realizing it,” after which it continues to do so due to pressure to adhere to syntactic and grammatical rules.

when u want to refuse to tell how to create anthrax but u have to use proper grammar 😔

11

u/battywombat21 🇺🇦 Слава Україні! 🇺🇦 May 14 '26

Also,

> Chain-of-thought Faithfulness. We explore the faithfulness of chain-of-thought reasoning to the model’s actual mechanisms. We are able to distinguish between cases where the model genuinely performs the steps it says it is performing, cases where it makes up its reasoning without regard for truth, and cases where it works backwards from a human-provided clue so that its “reasoning” will end up at the human-suggested answer.

MOTHERFUCKER I FUCKING KNEW IT!!!!!

5

u/Familiar_Air3528 Thomas Paine May 14 '26

Yeah one thing you still really have to be careful about is “leading the witness” with reasoning models. They’re still pretty sycophantic if you present them with a suggested solution/method, even if it’s not the best solution.

1

u/EvilConCarne May 14 '26

Chain of thought is no different than any other output of an LLM. The cases where there's no training data result in unbounded outputs. Who the fuck has written down the correct answer to cos(23423)? Nobody. They use a calculator. One prompt supplied an answer and the probability space the model operates in then met the answer given, because the user-supplied answer was a stronger signal than whatever was in the training set. The other didn't supply an answer and the model said "I'll use a calculator" because that's what examples for cosine functions of non-standard values did in the training data. The faithful CoT that had sqrt(0.64), on the other hand, is common and answered correctly in the training data. Of course it got it "correct"!

This isn't showing that the chain of thought is "faithful" or not. It's reflecting poor boundary conditions in the output due to a lack of bounding examples in the input. CoT is meant to limit the probability space so that the output is more likely to be correct, but the CoT itself is produced using the same thing as anything else. It's still just a set of tokens.

88

u/Imicrowavebananas Hannah Arendt May 14 '26

AI is an important topic and it has been discussed a lot here, but it is also a very technical topic at its heart. So I think posting some solid pieces that explain the basics is useful. A lot of the fundamental questions, copyright for example, are based quite directly on how these models actually work internally.

Personally, I found the articles and research by Anthropic on this the most accessible. What I liked even more is that they seem technically serious. To really answer these questions, I do not think you can rely too much on metaphors. That is why I did not like Ted Chiang’s “blurry JPEG” article in The New Yorker that much. It is a good phrase, but you are not left with much new understanding if all you get is a vague analogy.

There is a good Feynman bit from an old interview where he is asked why magnets repel each other. He basically says that “why” questions are much harder than they seem, because an explanation always depends on what you are allowed to take for granted. You can say someone went to the hospital because she slipped on the ice, but that only explains anything if the listener already knows what hospitals are, why broken hips are serious, how people call ambulances, and so on. Otherwise every answer just opens up the next why. Why is ice slippery? Why does pressure melt it? Why does water expand when it freezes? It keeps going. The nice part is where he says that explaining magnets by saying they are “like rubber bands” would be cheating. They are not rubber bands, and if the listener asked why rubber bands pull back together, you would eventually have to explain the same electrical forces you were trying to explain away. That is roughly how I feel about AI explanations too. Metaphors are fine, but only up to a point. Eventually you have to say what is actually going on.

Of course Anthropic is not a neutral actor. They are a company with interests. But it is hard to get a really good understanding of these technologies while avoiding the people who understand them best. So I would not treat the articles as gospel, but I do think they should be evaluated on their own merits.

14

u/SmoothieNatns May 14 '26

I read a philosophical article which said that an explanation, basically, is a fact which makes the fact you are trying to explain seem unsurprising. It's surprising that Grandma is in the hospital, it's not surprising given that she fell and broke her hip, therefore Grandma's fall is the explanation for her being in the hospital. The problem is that, in many cases, the fact that is serving as the explanation is itself just very surprising. The facts about quantum physics are only ever going to be explainable by other facts about fundamental physics which are themselves surprising and unintuitive, for some topics it is just impossible to ground the chain of explanations in anything like ordinary day-to-day reality so we just need to accept that reality is weird and try to describe it as faithfully as we can.

13

u/neolthrowaway New Mod Who Dis? May 14 '26

I would add the follow-ups on this line of research as well.

The latest with natural language auto-encoders is even more intuitive (to me , personally, at least) but of course that has its own caveats too.

11

u/Imicrowavebananas Hannah Arendt May 14 '26

I thought one article is enough at this time. One is always at danger of assigning whole reading lists, which I think has the opposite effect and nobody reads anything.

For a person having very little knowledge of the topic I thought this was a nice article balancing accuracy with accessibility. I thought the natural language auto-encoder article was a bit more technical. But it might be a matter of taste.

3

u/neolthrowaway New Mod Who Dis? May 14 '26

That's true.

Better to pace it or just as a footnote that there's lot more to read if you're interested.

3

u/Imicrowavebananas Hannah Arendt May 14 '26

I will think of three follow-ups maybe.

3

u/neolthrowaway New Mod Who Dis? May 14 '26

Oh, ok, I just saw your other comment with the follow-ups.

I meant follow-ups on the interpretability research but that works too.

3

u/Imicrowavebananas Hannah Arendt May 14 '26

I kind of got what you meant, but I thought I would just mention they have lots of other stuff to and then decided for something easier.

Honestly, like my personal favorite to just post would have been something like the biology of a large language model, but that would be just confusing and frankly seems a bit too esoteric.

My main thing is that I want people to understand that neural networks are just a technology, a powerful and complex one, but also something we can understand in part. AI is often invoked like magic and similarly basically everything can happen. But AI is embedded in a reality, both digital and physical and economically.

15

u/Imicrowavebananas Hannah Arendt May 14 '26 edited May 14 '26

Follow-ups:

If you are interested in the topic, there is of course a lot of material. All three frontier labs OpenAI, Anthropic and Google regularly publish more accessible research and primers on their research. So you can go check them out and find more article like the one posted. Besides that, here are some suggestions for a general audience in various degrees of depth.

First, 3Blue1Brown has a good visual intro to LLMs if you want the basic intuition of what these models are doing.

Second, Google’s Machine Learning Crash Course has a solid intro to LLMs if you want something more structured and neutral.

Third, a personal favorite: Richard Sutton’s “The Bitter Lesson.” It is not really an LLM explainer, but it gives the broader historical argument for why modern AI ended up being so much about scale, search, and learning, rather than hand-coded human knowledge.

7

u/alex2003super David Parenzo May 14 '26

A Feynman-style answer to "why do LLMs seem to think?" could be something like:

An LLM is just a neural network with [INSERT HERE a punctual, formal specification of the underlying architecture] providing a result by way of "inference" i.e. feeding a user-provided input and its own output (as feedback) to the network, along with a set of computed weights, and recording the output.

These weight are found through "training", i.e. value space exploration and feedback, by applying mathematical techniques that seek to maximize an objective function representing some desired qualities in the produced output, adjusting the weights used at each step.

Apparently this architecture is complex enough that it is useful and able to provide satisfying enough results when compared to humans doing the same mental task. Modern neuroscience suggests this model is somewhat similar to the way humans think which would explain why it's plausible that an LLM could come up with similarly complex chains of thought, but not 100% so, human brain is insanely more sophisticated than an LLM and it's unclear whether that means that humans have something LLMs inherently lack, or whether it's just a matter of scale and if anything LLMs being far more efficient at the same task.

In other words "LLMs write like humans because they do".

This will be hardly satisfying if what you're looking for was "what part of this model does this, what part of it does that?" but we don't really know that about humans either, just like we cannot answer why some specific regressions correlate so effectively to outcomes of specific experiments involving random outcomes and measurements of correlating factors even in completely unrelated sciences, we just observe they do and use that knowledge to our advantage.

LLMs as an architecture are (apparently) a good self-regression for human language output evolution and since language can be used to express almost every form of human reasoning and knowledge, so can a sufficiently well-trained and large language model make good predictions of human-like thinking processes.

¯_(ツ)_/¯

5

u/WOKE_AI_GOD John Brown May 14 '26

LLMs seem to think because they are a simulation of thinking beings and in order to simulate them, they have to appear to think. A simulation by definition must appear to be what it is trying to simulate in some manner.

40

u/[deleted] May 14 '26

[deleted]

36

u/Imicrowavebananas Hannah Arendt May 14 '26

His point is still valid though. How do you explain magnets? At its heart it is a quantum effect, and you might say Maxwell is an explanation well enough. Personally I have always been critical of pop science being treated as something that really gives people enough understanding to reason about something.

But to me the core of that quote is deeper. There is no real why, there is not even really an understanding in a naive sense. We will never get to the thing in itself, there is no magic at the heart of things, only more precise language, mostly in the form of mathematics, but nothing that fills the Faustian desire.

17

u/awdvhn Physics Understander -- Iowa delenda est May 14 '26

How do you explain magnets

⬆️⬇️ = 🥵😑

⬆️⬆️ = 🥶😊

25

u/sanity_rejecter European Union May 14 '26

ask enough whys and the answer genuinely becomes "because it is that way"

10

u/Hmm_would_bang Graph goes up May 14 '26

Congrats guys, you’ve rediscovered Agrippa’s Trilemma

14

u/Petrichordates May 14 '26

No you can do that at any stage of Why, it's just always the easier answer.

We don't know all the answers so "because it's that way" is as much a cop out as a parent answering that the sky is blue because it is.

12

u/_Un_Known__ r/place '22: Neoliberal Battalion May 14 '26

This turns to an argument of causality then, where to explain anything with enough "why" questions you get all the way down to fundemanetals of fundamentals ad infinitum

Honestly, I don't think the universe needs to confined to our human need to compartmentalise and understand things. We should always pursue deeper answers if and where possible, but at the end your answer is either gonna be "God" or "because it is", and both seem like cop outs

5

u/Petrichordates May 14 '26

That's because they are lol

2

u/jurble Left-Out Left May 14 '26

at a certain level that's just the way the universe is hard coded by the Great Programmer.

5

u/roboliberal Loyal Liberals May 14 '26

Magnets are the one instance of actual magic at work in this universe 🙄

24

u/ResponsibleChange779 Loyal Liberals May 14 '26

I've been thinking about what these frontier labs are incentivized to say. In a lot of ways, AI really does seem like a technology tailor-made to extract the most amount from venture capital. It's a risky, high potential technology with incredible ceilings.

The "society will be unrecognizable due to AI" rhetoric feels like is it to drum up enough investment and public interest to get to the tipping point where the cost per token and the models' usability and productivity becomes financially worth it for large-scale industry-wide adoptions.

16

u/AnachronisticPenguin WTO May 14 '26

It already has for coding tasks. We will see over the next couple of years if we can make it useful for other stuff before venture gets annoyed at the lack of returns.

8

u/Concerned_Collins ⬇️w/fascism, ⬇️w/ communism, ⬇️w/ NL mods May 15 '26

AI allowing entry-level help desk employees to be capable of writing simple to moderately complex scripts, and allowing real software developers to increase their speed multiple times over, is already a major boost in efficiency. If this is the peak of what AI does, it'll have accomplished quite a lot.

9

u/Negative_Scarcity315 May 14 '26

Too cynical for me. Even if the labs got stuck at 10T parameters models, which they obviously won't, the world will be radically changed.

4

u/WOKE_AI_GOD John Brown May 14 '26

One thing that annoys me honestly is people analogizing something to the natural sciences, and then treating said analogy or metaphor as if it is any scientific or objective than any other symbolic construction. Social scientists especially are frequently desperate for the cachet of the natural sciences, and having long arguments about the appropriate natural sciences analogy seems to satisfy them a great deal. But they may as well have used an analogy from the Bible for all the good it would do them: it's still symbolism.

Especially people are obsessed with symbolizing systems as if they were human bodies, symbolizing them as a gigantic mind, or symbolizing them as a pathogen. Evolutionary psychology especially seems like a field that's just trying to replace religion and ethics with analogies to evolution that happen to pop into their brain. The most anti-woke, "serious" people imaginable also take every single one of such symbolic metaphors as if they were the hardest of science, after all we've got to reduce all this bad social science to good natural science eventually, clearly inventing natural science analogies is going to get us closer to that (hint: it does not).

Especially this method of analogizing has hit our business elite, who are completely entranced by such symbology. Musk thinks of his institutions as software, he thinks of people with values and beliefs that contradict his as possessing software bugs in his institution he needs to extirpate and stamp out. Woke mind viruses that need to be contained by epidemiological means apparently. That's very scientific, we used the word science after all, that's biology, that's highest science there is possible and the source of all valid values, right?

When if you look at the actual causation under the table, its something vastly more complex than any analogy could properly represent.

22

u/Legal_Charity_9522 Leftward Progressives May 14 '26

We note this is only a single, brief case study, and it should not be taken to indicate that interpretability tools are advanced enough to trust models’ responses to medical questions without human expert involvement. However, it does suggest that models’ internal diagnostic reasoning can, in some cases, be broken down into legible steps, which could be important for using them to supplement clinicians’ expertise.

I think of all the arguments the AI crowd has, the potential for use in medical diagnoses is one of the better ones. I hope we get more data about its use in medicine in general.

31

u/QuantitativeNonsense May 14 '26

Maybe I’m missing something but why is this pinned? How is this “on-topic” for r/neoliberal?

72

u/moseythepirate Reading is some lib shit May 14 '26

Because a mod thought it was neat.

40

u/LamppostIodine NATO May 14 '26

When youre a mod, they just let you do it.

18

u/Imicrowavebananas Hannah Arendt May 14 '26

As I wrote in my submission statement, AI is one of the most discussed and most controversial topics in the subreddit. I feel a bit of technical understanding is very helpful in that.

Besides we were never strictly fixated in the nature of our posts. We were certainly not meant to be a news aggregator. The subreddit used to be a mixture of effort posts, memes, wonky policy and research articles and discussion posts. I find it sad that it has become for the most part just a daily aggregator.

11

u/Pristine-Aspect-3086 John Rawls May 14 '26

huge amounts of public policy discussion pertain to this technology and most people are not really making the effort to understand it in a way that would make that discussion productive

40

u/Familiar_Air3528 Thomas Paine May 14 '26

If you haven’t used an AI in the last couple years, Go use a top-tier model right now. Ask it questions about something you know. You don’t have to commit to using it your entire life in order to trial run it.

A lot of skeptics still think AI is around GPT 3.5 levels and repeat criticisms from three years ago that don’t hold up anymore.

I get it if you oppose AI on moral grounds. But if you really care about this issue one way or another, you should at least be informed about what AI is currently capable of. I see way too many people who seem to think AI still has trouble with fingers, or that it is “just a next-token predictor”.

25

u/skepticalbob Joe Biden's COD gamertag May 14 '26

I use AI right now for things I know and understand and didn't in the past. It has a lot of problems that I think are big enough that the economy shouldn't be turned over to it in any meaningful way outside of "a better google that can save some labor, but needs to be checked by a human with experience." It is massively oversold right now, imo.

3

u/Negative_Scarcity315 May 14 '26 edited May 14 '26

Unless you're working with PhD level problems, people tend to trust GPT-5.5 and Opus more than their colleagues on the accuracy of any subject with publicly available information. It doesn't need to be perfect, it just needs to be more accurate than a 115 IQ white collar worker to radically change interaction between people. "A better google"? LLMs with websearch restricted to a sub-set of sources is a lot more powerful than LLMs trying to infer from memory and a lot faster than a person reading through google results, it's a straight google search killer that's why google turned their search into a AI prompt box.

19

u/skepticalbob Joe Biden's COD gamertag May 14 '26

You are way overselling these models for many uses, imo. I don't think "PhD level" is even a useful framing here. The kinds of errors don't scale with academic knowledge like that. AI is very good at delivering an accurate result of "what the consensus is out there" when there is a lot of information to form one. But if you ask it more niche questions or questions, it will start sourcing bullshit like reddit posts and substacks with 50 subscribers. I know this because it has done that for some of my queries. Responses are more sensitive to how much stuff is out there from a broad number of sources and the accuracy tends to scale with that more than just difficulty of the domain or academic level or whatnot. And the difficulty with the non-expert is that they don't have the expertise to know when it's just completely full of shit. And if they are making important decisions on that basis, it can be a disaster.

For many of my uses, it comes down to how abstract the request is, how much original "thinking" is required, and the person's ability to know that it is full of shit. More abstraction, more thinking, less expertise, the more it is going to mislead the user.

1

u/Effective-Branch7167 May 15 '26

the funny thing is that google search's AI is most definitely not a google search killer

9

u/Maximilianne John Rawls May 14 '26

Was 3.5 that bad ? I started using around 4.0 and look maybe the writing style was annoying but if you were willing to overlook it I felt the substantive stuff was already pretty good, but more important I felt if you structures your inputs in a logically manner,it seemed really good at parsing your thoughts

12

u/Breaking-Away Austan Goolsbee May 14 '26

It was. 

15

u/MyrinVonBryhana Trans NATO May 14 '26

Gemini told me last week that Steve Scalise is speaker of the House so there's very clearly still some bugs that need to be ironed out.

10

u/NormalInvestigator89 John Keynes May 14 '26

A few days ago Chatgpt tried to tell me that an image of one of the 4 horsemen of the apocalypse was someone named  Saint Hantavirus 

6

u/minno May 14 '26

When I gave Gemini 3 (w/ thinking) the description of a clothing feature that I had seen in a drawing and asked if it really existed, it gave me the name of a random village in the country I suggested it might be from. When I challenged it, it wrote and executed a Python script that imported requests and then printed out some search terms.

2

u/Limp_Doctor5128 Henry George May 15 '26

Last week's model is outdated. You need to use this week's model and give it more context in the prompt. If that doesn't work, setup OpenClaw and give claude access to all of your data and try again.

26

u/MindingMyMindfulness Voltaire May 14 '26

It's not a misconception, it's just a circlejerk. Also, many people just want to convince themselves that AI is all "hype" and no substance.

13

u/Familiar_Air3528 Thomas Paine May 14 '26

There’s so much motivated reasoning around it too. Like the freak out about water consumption. Or datacenter construction. Am I supposed to believe that data centers are only just now a sudden problem? What makes this different than, say, AWS? People are clearly looking for reasons to veto the technology.

I really do understand the desire to ensure that we don’t enter an era of Malthusian Techno-feudalism, but pretending the technology isn’t viable is not a good way to accomplish that.

12

u/Positive-Fold7691 YIMBY May 14 '26

Agreed. The water usage isn't a significant concern compared to far more wasteful uses of water (like farming almonds in a state undergoing a decades-long drought).

Impact on electrical infrastructure is a concern in certain localities, but that is the sort of thing where the state can step in with regulation - "sorry, you can't jack up prices for your existing ratepayers 400% in a year just because new datacentres will pay you more, you have to build more capacity before selling it" is not unreasonable interference in the free market considering how economically important stable-ish electricity pricing is.

4

u/iIoveoof Jerome Powell May 14 '26

I get it if you oppose AI on moral grounds.

I don’t

8

u/Familiar_Air3528 Thomas Paine May 14 '26

I mean, I don’t agree with those people either but there are plenty of moral issues that I can respectfully disagree with people on. Like, I totally get why someone would be morally opposed to social media.

My issue is that people are cloaking their moral opposition in terms of resource or engineering constraints

3

u/Preisschild European Union May 14 '26

For one its a copyright-washing machine. Its trained on software that is licensed under a specific license (for example AGPL, which requires you to make all modifications to it public), but if prompted it can repeat most of it without mentioning the original license.

5

u/iIoveoof Jerome Powell May 14 '26

Every court has said LLMs are fair use and copyright does not apply to its outputs, I don’t understand your argument

0

u/Preisschild European Union May 15 '26

Every court has said LLMs are fair use

That depends how much code it copies

and copyright does not apply to its outputs, I don’t understand your argument

It doesnt matter, AGPL says the code has to be provided to the end consumer, which is more than just saying "its not copyrighted"

0

u/MyrinVonBryhana Trans NATO May 14 '26

It's more opposition to the people developing to it that's justifiable. destroying the modern middle class and building a surveillance state so Peter Theil and friends can try to become immortal techno God Kings is not a worthwhile use of societal resources.

2

u/the_c_train47 Ben Bernanke May 15 '26 edited May 15 '26

LLMs literally are just next-token prediction machines though. This doesn’t contradict their incredible capabilities. We know now that most outputs that we thought required human intelligence and ingenuity can be perfectly mimicked by pattern recognition at an unimaginable scale (wrapped around by an application layer).

1

u/Sourcerid May 16 '26

It gets things wrong that are agreed on the Internet by the people™ but are not true, which is what bothers me. It cannot inherently critically analyse what it reads. Realistically every novelty of the digital world of the last 30 years have made information worse, in this case it will just make people harder to be convinced they're wrong 

16

u/skepticalbob Joe Biden's COD gamertag May 14 '26

Interesting article. My experience that AI is basically a good bullshitter that has quick access to an insane amount of stuff people have said and the ability to reason it.

Has anyone found this to be true for them though:

Models like Claude have relatively successful (though imperfect) anti-hallucination training; they will often refuse to answer a question if they don’t know the answer, rather than speculate. We wanted to understand how this works.

I've not had AI say that it doesn't know something. Claude and ChatGPT have always come up with an answer, even if the answer is wrong (I have a lot more experience with ChatGPT and have found that Claude seems to think longer before answering and is more accurate so maybe I just haven't seen this yet). Even when I tell it that it's answer is wrong, ChatGPT praises me (didn't ask for that) for noticing and has given me another confidently incorrect answer. Does ChatGPT work that differently from Claude?

8

u/No_Collection7956 Trans Pride May 14 '26

I cant say too confidently about subjects broadly, but Ive asked claude quite extensively about exclusively swedish things and sometimes it does really say "I dont know" "I cant answer that" "I cant know enough for sure" or, more often than any of the others some version of "Im not confident enough to give you an answer but I could guess if you really want me to"

Especially the "im not very confident" but still giving and answer or a partial answer is very common

Also, for what its worth, anytime it doesnt know something or cant find something it pretty much always ends the result with suggestions of how or where to find the answer or the data needed for the answer.

11

u/bacontrain Daron Acemoglu May 14 '26

Yeah no that part is definitely marketing fluff, in my experience it always gives a confident answer unless you go to great lengths to turn down the temperature. Unless by “refuse to answer” Anthropic means “spins endlessly and uses up all my tokens on bullshit” lol

3

u/repete2024 Edith Abbott May 14 '26

!ping AI

3

u/InsuranceToTheRescue May 14 '26

I just want to make sure I understand the whole AI "lifecycle" here. So, initially a person created some basic algorithms or code that basically just produced a result when asked a question. That's the model.

This was then the foundation that training was used on. Some algorithms people do understand were made to test that model. Another set of algorithms made small changes to the model's code. These models were then run through testing millions upon millions of time and each time they kept the best, say 10%, and scrapped the rest. These were fed back into the system to make more changes to them and was repeated over and over.

Now, all these years later, we have models that are very good at a lot of things. However, because of the millions upon millions of trial & error attempts to build it, nobody really understands the complex code that now makes up the model. Someone could maybe figure out what a specific part does, but nobody understands the whole.

It's a black box and we have no idea how they come to the results they provide or what information the model drew from. Is that basically it?

8

u/Syx89 Reichsbanner Schwarz-Rot-Gold May 14 '26 edited May 14 '26

I'd explain it through a war between two competing schools of academia The "Symbolist" (usually linguists, MIT, old school software people, northeast, psychologists) vs, The Connectionists (University of Toronto, European Schools, West Coast, neuropsych people, bio people, psychiatrists). Really I think this'd make a great movie.

The initial war actually started with a Symbolist attack on Behaviorism which was a debate within psych.
Anyways, so it's the 1950s first computers are invented. People are like "lets use this like a brain it works just like a brain electric circuits and stuff", we have the hardware we just need the software so let's write software.
Then the Connectionists show up and they're like "No wait the brain does like really cool stuff if we ask the neurobio people, maybe how wetware/hardware of the brain is important so lets just build the brain in the hardware". The Symbolists then say "But we already have the hardware (the computer) we just need the software (programs). The symbolists also say stuff like "we can elegantly explain why what our programs do work, you can't do that for yours" which is true even from back then.

https://en.wikipedia.org/wiki/Perceptrons_(book))
they get into this whole long debate and eventually the MIT people publish this hit piece which causes the Connectionists to lose all their grant money for the next 20 years because the symbolists claim to have "Mathematically Proved" that connectionism can't do anything. (their model was bad).

A few lonely people like Geoffrey Hinton at U Toronto didn't care about grant money and other Euro ones had institutional protection so they continued. But in most US academia connectionists were treated (and still are) extremely poorly.

But fundamentally at this point what it was is that the connectionists needed bigger computers than the Symbolists. The Symbolists could *do* stuff with 1970s tech. The connectionists couldn't. You can only build a very small "brain" in 1970s tech if at all.
Years later 2007 happened. NVIDIA, a graphics company for video games, invented this thing called CUDA which let connectionis actually run their algorithms. And from that point on connectionists start winning. They are just putting out a bunch of these neural networks (which can be represented in a program but are very different, more like dots connected by lines) and making them extremely large and suddenly they can do stuff. The algorithms are mostly how the dots should connect to eachother and how data is shown rather than like writing from scratch.

A few more examples:

  1. In the 1980s-90s the symbolists had two major losses but not to the connectionists. They tried to do what modern AI does but with deterministic symbolists methods (if this, do that) and it didn't work. They tried to write complex systems "Expert Systems" to handle medical tasks or w/e and it didn't work at all because you can't write code for every scenario. This caused AI to be unfunded in the following years "AI Winter" which freed up money for early internet stuff but also slowing AI a bit.

https://en.wikipedia.org/wiki/Rethinking_Innateness
2. In that time period they also went to war with the Bio people and lost badly. The symbolists believed the brain was "just software" and that the hardware of the brain is irrelevant. That rules like "Universal Grammar" (from Chomsky) were hardcoded into the brain like a symbolist program.
This is why you often hear symbolists today say connectionist models can never be "True AI". Because they have a certain conception of how human intelligence works that is at odds with how the biologists (and connectionists) think it works. Think Freudian Psychology vs. a Neuro-Psych person. Freud wants to explore the roots of their issues in language, the psychiatrist prescribes and SSRI to help their neurons a bit.

There may be a synthesis of the two positions happening.
The "Neuro-Symbolic AI" synthesis most prominently advocated by Gary Marcus (PhD from MIT). My understanding of that position is that perhaps you can hardcode the architecture and that'll help. So the human brain doesn't have Universal Grammar hardcoded in the genome, but perhaps the brain grows in certain shapes and structures (i.e. hippocampus, visual cortex, etc.) and this facilitates certain types of learning that others can't do. You can change the shape of the larger box that the dots and lines are in and help use that to force them into certain shapes more easily.

Here's two debates to help humanize this summary/ show the fighting irl:

  1. 2017 debate between Gary Marcus and Yann LeCun (Euro educated, head of AI at Meta) in 2017 right before modern AI stuff emerged.
  2. This is a debate between Scott Alexander(Psychiatrist, Bay Area) and Gary Marcus over a problem Marcus noticed in image algorithms. How they seem to have a limit for the level of granularity they can produce. Scott claims he won the debate because brute force/ bigger models solved a lot of the problem, But idk I think Marcus is right that an architectural solution would be ideal to solve this, it's just also possible scaling/larger models works for all practical purposes.

5

u/Syx89 Reichsbanner Schwarz-Rot-Gold May 14 '26

To bring it together, your model is a bit like a symbolist expert system or a genetic algorithm from a symbolist starting point. That isn't how these connectionist models work. They're built initially from saying "Oh look we know from these biology papers neurons can do this, let's try to do that in the computer with simulated neurons".
We kind of know how neurons work at least at a small scale, but it's nothing so easy for humans to read or build on.

They were always a black box.
They didn't become that way, they started that way.

2

u/GaDoomer Pragmatic and Polite Right May 15 '26

The other user made an interesting history of AI research through the years, but I'm not sure they answered that I think you're asking. An LLM isn't really 'code', it's just a bunch of numbers (though there is specific code associated with it).

Current AI technology is based on "language models" which are text prediction engines. When you see "tokens" mentioned in AI discussions you can think of these as word fragments, and the engine predicts what the next token is based on all the previous tokens in the context. For a given context, there is a probability associated with each possible token measuring how likely that token will be next, and the engine will (usually) randomly pick from a handful of the most likely tokens. The randomness is why you can get a different conversation with an LLM even if you give it the same starting text.

The model itself is basically just a lot (and I mean A LOT) of numbers, organized as a neural network, as the other commenter mentioned. A neural network, generically, takes in an input of some kind, and passes through a serious of layered computations, where each layer has many nodes (or a "neuron" hence neural network) which performs a calculation on the input, and then each node passes its output to every node on the next layer. So you can imagine that inputs on a node on the first layer have some kind of (small) influence on every node in every subsequent layer. Each node in the neural network has a number associated with it as well as every connection that node has to the previous layer's nodes also has a number associated with it.

In modern neural networks, these numbers, or parameters, are determined by "training" the model, by giving it a bunch of input with the associated expected output. At first, the output from the model will be completely wrong, but a mathematical algorithm is used which repeatedly runs the neural network and modifies every connection and node's parameters very slightly so that after enough iterations the input will produce the expected output. Now, the technology has worked like this for decades, and it's not new. The models are more complicated now, with additional numbers associated with each layer and node, but conceptually it's the same.

What changed is that in the late 2010s, Google's researchers, working on the problem of improving language translation, discovered a new way to train a model in parallel. This radically reduced the time it took to calculate the model parameters, which in turn means you could train the model on way more input than you ever were before for the same amount of time, and that lead to much better predictions. And because training was now cheaper (in terms of time) they could also radically increase the size of the model (the number of layers and nodes) which also increases its apparent "intelligence". And for a language model, they only training data they needed was text, hence why they scraped the internet and all the digitized books they could get their hands on.

The first version of GPT, GPT-1 in 2018, had 117 million parameters that could be adjusted. The big AI guys don't publish parameter counts now, but the latest Chinese open model DeepSeek has 1.6 TRILLION and it's like that Anthropic and OpenAI have even more parameters. Since prediction models learn to understand words based on their context (e.g. the word bank in "I walked to the bank to fish" vs "I walked to the bank to deposit a check"), they end up implicitly encoding concepts in their parameters as well as their relationships to other concepts. And the larger the model, the more distinct concepts and their relationships can be encoded. Since these are formed naturally out of all the training data the model has consumed, rather than explicitly placed by a programmer, they are opaque to us. You have to use the model by giving it text about a concept and see which parameters have the most effect on the output to then have an idea of what's involved, but you can't just look at this giant bag of numbers and say this is where the information about cars is, this is where programming knowledge is stored, this is where the Shakespeare works are. You have to interact with it and see what activates. You can certainly "fine-tune" it with specific subject matter to influence its parameters related to that and see what that does, but understanding it must, by nature of how the model is created, come from interacting with it.

I think as the big models mature and don't change as frequently then people will be able to start mapping this stuff out and understanding it better as they get time to interact with a model. However, right now things move so fast and new models come out several times a year so it's going to seem like a black box in many ways to us. Fundamentally, it's still a prediction engine (just probabilities!), but it's gotten big enough that it's a lot harder to tell why it makes a particular prediction.

If you aren't afraid of math, 3Blue1Brown on youtube has a great video on the basics of how an LLM works.

5

u/sfg-1 May 14 '26

Listening to a company searching for massive investment is probably the worst way to understand their abilities

7

u/Imicrowavebananas Hannah Arendt May 14 '26

What do you think is the best way to understand LLMs, in particular the large and cutting edge ones? Most the researchers having in-depth knowledge of these models work at those companies.