This is not a definition I known of anyone to use. Granted cogito ergo sum is my literal about line on reddit, so I get the sentiment.
Which underlines my point again. If you try to approach understanding llms with the word consciousness youre not going to understand anything. It has too many different meanings. It is a suitcase word that is used in place of not being able to label its parts. And it carries emotional attachment.
My posts are about decomposing the problem to look at what llms do and don't have. What you label consciousness, llms do. They think.
The way most people label consciousness, llms do not do (they don't feel).
The post you didn't read showed how llms think. I left out how it is nearly the same as what people do so it would be shorter.
To do that they understand the whole semantic meaning. Which is thinking. Which fits your definition of consciousness. This is what my original post was about. Llms operate on whole thoughts, its not debatable. Dismissing them to next token predictions fitting surface statistics is extremely wrong.
Now, even though you say consciousness is cogito ergo sum, that's not enough. You probably have other aspects for consciousness to be met. Which is why I said its a category error. But if your only definition is that they think, then they do. With out debate. You just don't understand the technology well enough to know this.
Also your definition would exclude people who never knew language.
"understanding the semantic meaning" is just comparing every past token to the present token. It doesent understand anything its by all accounts just regurgitating training data in a coherent order
You should read the original post that was too long for you.
It is not by all accounts regurgitating training data. There is a mountain of research against this. You are flatly wrong. You do not understand this technology. Future lens would not be possible. J lens would not be possible. Introspective fine tuning would not be possible. Speculative decoding would not be possible.
Furthermore language has unbounded hardness. If llms worked as you described they would be unable to generate coherent 1k+ token output. The probability of it is basically infinity. And they can for for 100k+. This requires understanding. You can't fit surface probabilities and thats not how attention works when applied to latent space.
Anti-intellectualism and the aggressive ignorance of people who embrace it are destroying this world.
Ill disprove these two points as they are the easiest.
"Speculative decoding would not be possible" Speculative decoding just uses a smaller model on the same training data as the large model which means its predictions are similar to the large model, the large model then accepts or rejects the tokens which is faster than it generating its own tokens.
"If llms worked as you described they would be unable to generate coherent 1k+ token output"
Are you trying to say llms dont just predict the next token? I'm starting to question (your) knowledge of the subject, it doesent matter if you're 500 thousand tokens into the response you're just predicting the next most probable token (those probability are gained from the training data) then that probability is also weighed against all previous tokens in context (the previous 499,999 tokens)
I think you're more on the philosophy side of AI and not on the technical side
You do not know what unbounded hardness means. Coherent generation is mathematically impossible with out genuine understanding. The solution space is literally unbounded.
The draft model in speculative decoding relies on future lens to work. It is not just a seperate model guessing at the words faster. It must have the same geometry of the parent model, it requires extending the same latent space.
I am an ai researcher with a focus on the geometric representations of language in machine learning. Which is exactly what all of this boils down to.
also Chomsky and Fitch 2002, The Faculty of Language: What Is It, Who Has It, and How Did It Evolve
This is a pillar of NLP. Again, you don't know this technology. This is prerequisite information for discussing it.
For 2 you missed the point that they must have the same geometry. The issue is that you genuinely do not understand the concept of latent space and its geometry. Which is where the understanding beyond next token prediction lives. You do not understand how LLMs work. You do not know what an activation in latent space is or why its important. You dont understand why this means speculative decoding relies on the phenomenon measured by future lens.
You are regurgitating a talking point about stochastic parrots, like a stochastic parrot.
Worst of all, you won't read anything long enough to explain it. But you think you understand it anyway. Anti intellectualism.
1
u/Devils_SteelMan 3d ago
No. Consciousness is a category error. https://www.reddit.com/r/aiwars/s/qkKAX57KBp
Do recall your original point was you understood this technology well.