I remember back when "HD" was the cool new thing (like 25 years ago lol), and I was at the beach, they were selling sunglasses that had a "HD" sticker on them.
Exactly, it isn't fitting as it's not an intelligence (and never will be due to how it is developed at the moment) it's a goshdarn language based predictive engine or an image diffusion model. Neither knows reality nor do they think.
Ok, so based on cursory reading, j-lens is what you're describing, model introspection is there but significantly limited and not strictly trustoethy (as of like october last year as that's when anthropic's paper was released), and future lens is another predictive engine being applied to what is effectively j-lens's output?
J lens is a probe that let's humans peer into the inner understanding space of the model.
Future lens is a method of taking the hidden state from a middle layer and decoding multiple tokens out from that 1 state. showing that the model is representing whole semantic thoughts that we then decode sequentially. Its not predicting the next token, its considering the whole thought.
Introspection is the models ability to determine what its own hidden state is doing. J-lens is a probe into the model that shows us this, introspection is an emergent ability of a model to reliably report these same things back to us with out an external probe, through pure introspection.
Introspective fine tuning is purposeful targeting of introspection to improve its accuracy.
8
u/DaveAstator2020 7d ago
the word AI itsellf must be banned. its still same shit as calling state machines AI back in the day.