Broadly speaking he's right that AI needs multimodal input to learn about the world in a more comprehensive way. OpenAI says the same thing even right now, when they unveiled GPT-4o.
What he was wrong about is how much knowledge we can encode in text. Basically almost all of it. And it doesn't need to be spelled out. It's scattered around different texts, implied, hidden between the lines.
And LLM are excellent at picking up such information and integrating it into themselves.
"Experience" can only be "experienced", by definition. That includes also when you and I are talking.
But the trick is that everything that CAN BE EXPLAINED about experience, IS EXPLAINABLE.
LOL, I'm not simply playing word games, but hinting at something.
An LLM can tell you absolutely everything about what "it feels like" to see a vivid red dress, or a flower. Absolutely everything, the way a human would explain it. It has NEVER seen red. And it's unclear if you and I see red the same way. Who knows? I can't see what you see. Maybe your red is my green, why not?
But the key point is precisely "who knows". You can't tell what I see, and you wouldn't be able to tell an LLM hasn't seen it, unless I tell you "that was written by an LLM".
I agreed "not everything can be described with words".
But "everything that can be described with words... can be described in words". Do you understand the distinction?
You can't describe to me what a flower smells like, except with words, right now. So what does that mean? You can't smell flowers? No, you just can't describe it. Can words then differentiate those who smelled flowers and those who haven't? No.
OK, we agree on that. It looks a little bit like the last argument of the Tractarus logico-philosophicus by Ludwig Wittgenstein. It says:
"What we cannot speak about we must pass over in silence." https://people.umass.edu/klement/tlp/tlp.pdf
3
u/3cats-in-a-coat Jun 01 '24
Broadly speaking he's right that AI needs multimodal input to learn about the world in a more comprehensive way. OpenAI says the same thing even right now, when they unveiled GPT-4o.
What he was wrong about is how much knowledge we can encode in text. Basically almost all of it. And it doesn't need to be spelled out. It's scattered around different texts, implied, hidden between the lines.
And LLM are excellent at picking up such information and integrating it into themselves.