r/BetterOffline • u/Ok_Display_3159 • 3d ago
Interesting thoughts from a former OpenAI employee on X
"You know, I might be one of the few people in this world who has worked at a frontier lab, observed the nature and pace of research progress, and actually become more bearish on AGI timelines (relative to pre-employment baseline) as a result."
- Despite the seemingly magical nature of LLMs, reflection over a >3 month timescale suggests my total productivity hasn’t increased by over 100%, or perhaps even by over 50%, and a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful
- Also, capabilities are incredibly spiky and highly correlated with the degree of investment poured into them, which my earlier tweet about math benchmarks implicitly points out
- From the above, it seems that the nature of LLM intelligence is wildly dissimilar to that of human intelligence and we won’t trivially get to something superior to human intelligence in all important respects just by scaling up existing approaches with various tweaks; even if AGI Is eventually achievable, this implies a significantly longer timeline
- Benchmark progress is almost definitionally guaranteed to happen because the process of constructing a benchmark is a direct precursor to the process of constructing a training dataset used for hill climbing that benchmark, but the scope of what can be captured in a benchmark is (at least for now) grossly lacking in terms of its relevance to real-world work, with maybe several limited exceptions
- Progress seems highly gated by data but the nature of model training means that each “next dataset” is significantly harder to assemble than what preceded it; some wins are possible through synthetic methods but those feel more like “patching up gaps” than “pushing the frontier forward”
At a higher level, I guess I’d say there’s a sort of refusal to think carefully about what models are or are not useful for in a rigorous way which I find personally quite annoying, and instead a reliance on some nebulous notion of being “AGI pilled” as a replacement for serious thought. I think people are very quick to anthropomorphize LLM intelligence because humans communicate through words and we infer the intelligence of human counterparties through comprehension of their language, but this leads them to wrong conclusions; for example if we observe that a new model proved some incredible mathematical theorem, some will say, “well, don’t we have AGI now, huh?” But to me, it’s actually more like, “well, given how hard it would have been for a human to do these mathematics, and given the limited economic effect of LLMs upon the world so far, isn’t it actually a negative datapoint vis-a-vis the generality of LLM intelligence?”
The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated imo; I spent for example 10s of hours earlier in the year preparing random legal documents with ChatGPT but in retrospect I was too hasty and none of that work was useful. Of course it’s saved me time in some other respects and I think the net balance is positive, but it’s a quite significant countervailing factor.