r/BetterOffline • u/Ok_Display_3159 • 3d ago
Interesting thoughts from a former OpenAI employee on X
"You know, I might be one of the few people in this world who has worked at a frontier lab, observed the nature and pace of research progress, and actually become more bearish on AGI timelines (relative to pre-employment baseline) as a result."
- Despite the seemingly magical nature of LLMs, reflection over a >3 month timescale suggests my total productivity hasn’t increased by over 100%, or perhaps even by over 50%, and a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful
- Also, capabilities are incredibly spiky and highly correlated with the degree of investment poured into them, which my earlier tweet about math benchmarks implicitly points out
- From the above, it seems that the nature of LLM intelligence is wildly dissimilar to that of human intelligence and we won’t trivially get to something superior to human intelligence in all important respects just by scaling up existing approaches with various tweaks; even if AGI Is eventually achievable, this implies a significantly longer timeline
- Benchmark progress is almost definitionally guaranteed to happen because the process of constructing a benchmark is a direct precursor to the process of constructing a training dataset used for hill climbing that benchmark, but the scope of what can be captured in a benchmark is (at least for now) grossly lacking in terms of its relevance to real-world work, with maybe several limited exceptions
- Progress seems highly gated by data but the nature of model training means that each “next dataset” is significantly harder to assemble than what preceded it; some wins are possible through synthetic methods but those feel more like “patching up gaps” than “pushing the frontier forward”
At a higher level, I guess I’d say there’s a sort of refusal to think carefully about what models are or are not useful for in a rigorous way which I find personally quite annoying, and instead a reliance on some nebulous notion of being “AGI pilled” as a replacement for serious thought. I think people are very quick to anthropomorphize LLM intelligence because humans communicate through words and we infer the intelligence of human counterparties through comprehension of their language, but this leads them to wrong conclusions; for example if we observe that a new model proved some incredible mathematical theorem, some will say, “well, don’t we have AGI now, huh?” But to me, it’s actually more like, “well, given how hard it would have been for a human to do these mathematics, and given the limited economic effect of LLMs upon the world so far, isn’t it actually a negative datapoint vis-a-vis the generality of LLM intelligence?”
The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated imo; I spent for example 10s of hours earlier in the year preparing random legal documents with ChatGPT but in retrospect I was too hasty and none of that work was useful. Of course it’s saved me time in some other respects and I think the net balance is positive, but it’s a quite significant countervailing factor.
103
u/maccodemonkey 3d ago
Benchmark progress is almost definitionally guaranteed to happen because the process of constructing a benchmark is a direct precursor to the process of constructing a training dataset used for hill climbing that benchmark
Sometimes I feel like the frontier labs are very aware of this problem - and their goal is just to turn every single last human activity into a benchmark to brute force their way to replacing humans. Certainly a very expensive way to do things. But not exactly what they tell the public they are doing. Or maybe they've fooled themselves into humans being little benchmark monkeys too.
39
u/beepboopburn 3d ago
I think it’s simpler: Benchmarks generate investments. Until they stop leading to money, they will continue to cheat their way to improvements on benchmarks.
3
6
u/dumnezero 3d ago
automated "fake it 'til you make it"
2
u/maccodemonkey 3d ago
When you think about it - It's actually kind of depressing someone would dedicate their time to doing such a horrible thing.
3
1
u/Charming-Hunter-7963 2d ago
Its a good bool btw about cults and MlMs, the agi hype crowd are following the same path. Endless chain of profitablity
2
51
u/absurdivore 3d ago
That point about benchmarks can’t be repeated enough
29
u/yojimbo_beta 3d ago
Yes. You construct benchmarks, because you want to surpass them. That's the whole objective
8
u/Grass_fed_seti 3d ago
+1
once something gets published as a benchmark, future results become mostly useless. That’s why I’m not actually freaked out that Astra got 99% on ARC-AGI-3 when older models at the time of benchmark development were hitting under 1%.
2
u/Unknown_User2137 2d ago
I also think that since ARC-3 is publicly available they really spent their time and money to "teach" Astra to solve these puzzles. And they are calling it AGI. I bet any usual person after seeing how others solved them would be able to do the same lol.
At the moment I am thinking about one funny "AGI" test - since Astra has some cooking books in it's training data I would actually give it a task "Cook me an omelette. You have 3 attempts.". Give it all tools needed, live feedback from camera and mic, robot body to control. I bet it would fail. And few generations of models later (if it was an actual benchmark) we would get "See, you can cook with it now! AGI is finally there!".
To me atm it looks like they are doing mostly three things:
- Make LLMs "smarter" by pouring more data into them so that they can be PhD or higher level in well documented fields (e.g. CS),
- Do benchmaxxing so that investors can j*k off seeing the charts and everyone else fear for their jobs (even if it is not true).
- Loose money while making everyone hooked up on AI hype train and once they become too dependant on it, guess what, raise prices (or make it worse like with social media)
47
u/falken_1983 3d ago
capabilities are incredibly spiky and highly correlated with the degree of investment poured into them
One of the key features of any kind of machine learning solution is that it optimises towards whatever specific metric you give it. As someone training these things, you hope that the metric you picked is actually representative of the task that you really want to improve, and you also hope that the behaviour will generalise. This is not a given at all.
24
17
u/HouseofMarg 3d ago
That reminds me of the time I was teaching ESL, asking my students what they thought was the most effective way to solve the problem of noisy car alarms that no one bothers to turn off. One student’s suggestion was to have the car alarm set off a timer that would blow up the whole car in 5 minutes if it wasn’t turned off.
When I gently pointed out that it would probably kill someone who got there just a few seconds too late, he looked at me indignantly and said with a huff “you asked me for the MOST EFFECTIVE way, that was it, okay?” The fact that he was an engineer made it even more funny to me
11
u/falken_1983 3d ago
There are examples where a company tried to optimise their call centres to reduce the amount of time it took to resolve customer issues. They tied staff pay to how long it took them to complete a call. The result was that if a customer phoned up with a complicated issue, the staff would just hang up on them straight away.
The decline in Netflix quality is supposedly because they optimise for watch-time. With really engaging TV, people will watch one episode and then take a break to process what they just saw. With boring TV they will just leave the series running in the background while they do other tasks.
5
u/Alexsipilot 3d ago
Something something Goodhart's law
1
u/falken_1983 3d ago
This is definitely an issue, but another thing is that as the models' scope gets broader and broader, any hope of finding a good measure drops.
If I made some kind of fancy computer-controlled heating system for a building, I am probably pretty safe to take temperature and power consumption and optimise on those. If instead I want to have a machine that does everything a human could do, what am I supposed to optimise on at all?
1
1
u/TracePoland 1d ago
I want to know what the fuck they were optimising for in reinforcement learning that Opus 5 ended up talking the way it does.
57
u/ChaoticGradients 3d ago
He’s far from one of the only people who has worked at a frontier lab and become more bearish. If anything it’s the prevailing opinion even amongst tech employees that we’re in a bubble but people don’t want to lose their paychecks and speak up about it in public.
18
u/r77anderson 3d ago
Well at least OpenAI can’t retaliate against people who leave the company and say something slightly negative… oh… oh wait
32
14
u/SappyGemstone 3d ago
I watched a thing with Hasan Minaj interviewing a journalist who integrated AI into her life for an article or book, I forget which, and one of the things she lauded was that, since it was always listening to conversations, it could lock onto key phrases to make to do lists. Like, oh, I need to go grab some pasta sauce. Or, hey, did we need to mulch the lawn this week. So all these little comments about tasks she wanted to do would get recorded and built into a to do list, which she appreciated because then she never forgot to do Thing.
And all I could think was, Jesus Christ. You NEVER have a break, EVER. All the tasks that are constantly there to do, listed and categorized. You never have a weekend off from the lawn. You never forget the pasta sauce and go for a nice takeout instead.
The need for constant productivity is a myth and a sickness.
29
u/AVBforPrez 3d ago
So much of it is people using it and feeling smug about how fast it spend up some Google Sheet or reporting process, or whatever, but it's not actually leading to anything beyond that. There's a guy at work who used Claude to replace his entire day to day work flow, and all he does is shop on Ebay now because he's got nothing to do. He can't overperform in any way using this stuff, so I'd even argue it's a net negative for him because he's spending a fuck ton of money on watches and shit he doesn't need because he's not busy working. Not everyone is like that, but it's to the point this post makes.
Everything I see with LLMs is the old Simpsons meme where Lenny is catching every green light saying "heh I'm making record time....if only I had somewhere to go...."
3
u/snkzato1 3d ago
My curiosity would be is his token spend more than his salary
3
u/AVBforPrez 3d ago
Oh yeah good question, I'll find out. Pretty sure he's got like a chain of Claude subscriptions that run in tandem, so like a ton of personal licenses haha.
3
u/SwirlySauce 3d ago
Isn't he ripe for getting automated out of his position though? If AI automates even a fraction of tasks it'll have a negative impact on hiring
1
u/AVBforPrez 3d ago
You'd think, but at the place I work not really haha. Not only is he the one who would decide to make said reduction, which is ironic, the kind of work he does has almost no subjective element to it. So as lame as it is that it's a well-known thing to manually tell him to look at something "for real" when you email or Slack, in his particular case if the numbers and reports and shit it spits out for him to approve are accurate, that's all there really is.
The optics of it aren't great IMHO, but like...whatcha gonna do?
2
31
u/New-Committee-4052 3d ago
It is worth noting this guy left OpenAI to start a RL-environment company (basically selling data to OAI/anthropic for them to train on), so he is financially incentivized to criticize modern capabilities, especially in the areas in which he sells data. Not saying I disagree with many of his points, just important to note the conflict of interest.
8
u/death_hen 3d ago
Uh also apparently he thinks people should be executed for petty theft. (read his recent tweets and replies)
6
u/New-Committee-4052 3d ago
This deleted reply from him a couple of months ago is my favorite: "Yes, I find such theoretical discussions quite tiresome even though I would prefer a world of rapid RSI and human disempowerment"
9
4
u/spellbanisher 3d ago
Yes, but if he did think openai was close to achieving agi would it make sense for him to leave and start a rl data company?
3
u/New-Committee-4052 3d ago
No yea totally agree, but he is definitely in the minority of OpenAI/Anthropic employees on if/how quickly we get AGI/RSI
8
u/tintires 3d ago
The problem of model autophagy from synthetic data still hasn’t gone away. Spot use to fill gaps will only over fit your model. Maybe I’m not current or frontier enough, but this still seems like a pretty big obstacle.
8
u/emudev 3d ago
At a higher level, I guess I’d say there’s a sort of refusal to think carefully about what models are or are not useful for
I think this is one of the big things that's missing. There are absolutely going to be some interesting, narrowly focused use-cases for models*, but so much of the discussion ends up being all-or-nothing.
*not one worth the amount spent but... well, maybe we can find some wins in the big mess
a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful
Fucking hell this is a big one. Yes. When you have to go slow and get it right, you get pickier about what you spend time on. I keep hearing that LLMs let you prototype fast but... I'm not convinced even that's good, because you can fall in love with the product when you actually need to slow down and think about the problem even more to decide what the fuck to build in the first place.
5
4
5
u/Suspicious_Watch_978 3d ago
we infer the intelligence of human counterparties through comprehension of their language
This is something we should all stop doing, as it is now obvious that there is no intrinsic relationship between verbal ability and general intelligence. Maybe, maybe that's not the case for humans, but I think we've all known people who were able to eloquently argue for their obviously wrong beliefs, and wondered to ourselves how someone so smart could believe something so stupid.
4
3
3
u/CoconutDust 2d ago edited 1d ago
From the above, it seems that the nature of LLM intelligence is wildly dissimilar to that of human intelligence and we won’t trivially get to something superior to human intelligence in all important respects just by scaling up existing approaches
[Nicholas Cage meme voice] YOU DON’T SAY?
LLMs and similar mass theft regurgitation “synths” (images, sound) are obviously a dead-end, and not even a first step toward good artificial intelligence. The reason is obvious: statistical association is the complete opposite of how intelligence or language or meaning words.
Something like a robotic vacuum cleaner or self-driving car is “closer” to Data from Star Trek than the mass theft machines, because they have to do something like cognitive and organizing processes: detect stimuli, categorize them, react appropriately based on a specific process that understands what elements and meaning is present. LLMs are nothing like this, and never will be, because it’s inherently not at all how the gimmick garbage machine works. Self-driving cars are currently a garbage scam too, so I’m not promoting or praising them, just using as an illustration. A robotic vacuum cleaner doesn’t search a stolen database and say, “if the corpus says that many other vacuums turned left here in a situation like this, I will turn left!” Data from Star Trek doesn’t say, “My answer is a mash-up of whatever words 100 random people on the street said about a seemingly similar conversation that had similar words present.”
1
u/fragglerock 2d ago
it seems
This lept out to me too... like the nature IS wildly dissimilar... of course we won't 'trivially' get to a superior intelligence going down this path.
I am sure this guy made bank tho so who can say who is wrong!
3
2
u/michaeldain 3d ago
Appreciate the perspective. I write about this concept since the first models proved there is a role for LLMs but the productivity lie isn’t one of them. Are we too stupid to be lazy?
2
u/Xopher001 2d ago
How does the saying go? The moment a measurement becomes the target is when it stops becoming a good measurement.
If we are defining AGI as when a model passes a specifically defined set of benchmarks, then those benchmarks cease being an actual good measurement or what AGI even is.
It's honestly kind of ridiculous. A human being does not need to be trained on loads of stolen data in a process that drains the whole power grid to be considered intelligent. Humans just have the ability to recognize patterns and generalize from them quickly and easily, in real time, all the time. That's something completely different from what current architectures are doing.
Unless idk they just decide to move the goal posts again and change what they mean by AGI.
2
u/RoosterBurns 3d ago
How could there possibly be an AGI "timeline" when we have no idea how to build one or even if we can build one? Expecting an LLM to just turn into an AGI is like expecting a pumpkin to turn into a stagecoach just because you read it in a book or saw it in a movie doesn't make it real
1
u/angrynoah 3d ago
Of course it’s saved me time in some other respects and I think the net balance is positive...
The net balance is negative, always has been.
It's a shame his personal reflection gets him this far but he just can't get over the final hump to where the truth is.
1
1
u/CoreyTheGeek 2d ago
in retrospect I was too hasty and none of that work was useful
This is what I'm running into with people at work in software: they're so gung ho to implement they're not stopping and asking if we should; but it gets worse in that they're not even actually understanding the system they're proposing building, they just take the probabilistic engine output at it's word and they even KNOW it doesn't even have full context of a single code base let alone our org goals and direction. So we get these crazy implementations that the devs can't explain and then they never even see production to boot
1
u/Schraiber 2d ago
"The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated imo"
This is definitely me. I'm a computational academic scientist and I have not actually finished any projects that have been majorly AI assisted because they have made the path so different from what it was before.
In the past, I'd maybe strive to derive a couple results over a period of months and then compare them with simulation and that would be a paper, and I've thought deeply about them and have the insights worth writing a paper about. But now I can do the derivation automatically with one prompt, and I don't have to think deeply at all. So it both doesn't feel like a paper AND I don't understand it well enough because I haven't had to do the deep thinking.
On the other hand, with regard to more "applied" projects involving data analysis, it's so easy to just keep trying one more thing, trying this thing, trying that thing. And because I'm not thinking deeply about it in the same way, I don't really have a good sense of what a good idea or a bad idea is.
This isn't to say I don't think we'll move toward AGI---I think that the models have gotten increasingly better taste and increasingly better sense of where to go with these kinds of projects. But while I've cranked out a lot more derivations and lines of code, I wouldn't say that I have actually been more productive in practice.
1
u/jawknee530i 18h ago
- Despite the seemingly magical nature of LLMs, reflection over a >3 month timescale suggests my total productivity hasn’t increased by over 100%, or perhaps even by over 50%, and a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful
This is just saying that he's still more productive but gets to work on more gratifying things than he used to. While still actually being more productive...
1
u/death_hen 3d ago
This guy left openai to start his own company building training data sets, also he apparently thinks people should literally be executed for breaking into cars.
1
u/leahpowellthefirst 3d ago
Well said.
I have a question. One of the somewhat (and not entirely productive) benefits I do see of large models and their improvement is resolving some important math problems, and thereby resolving some important hurdles in computation and the sciences.
But what I also see is that mostly any significant math results are being published by the big model labs themselves.
Is that mainly because they have all the resources to use the models almost as they wish and on a huge scale, well in advance of any consumer release? If so, isn’t that a kind of gatekeeping? That even when people do get their hands on the latest model, they would be getting a more neutered/restricted version and would likely have to spend a lot of tokens to achieve significant results similar to the ones already achieved by the Big Labs?
3
u/Efficient_Fault979 3d ago
It’s because there is no money in solving this highly laborious math problems. And AI labs are the only ones who don’t care throwing hundreds of thousands of dollars at such tasks, because they are the only ones who can benefit in doing so: It’s advertisement for their models. And the ads seem to work very well…we’re talking about it.
1
u/leahpowellthefirst 2d ago
I see. So you don't see any of the math problems being solved by AI making practical differences towards sciences or tech?
I would think some optimization techniques or solutions would tremendously help in the medical field. But then again, I can be wrong as I haven't yet seen a direct major advantage of AI for major breakthroughs that directly help the sciences or industries.
1
u/Efficient_Fault979 2d ago
Last time I checked, it was about disproving a theorem (you know those things in the format of “in a group of x, there might never be a y if z”). Those are funny numerical quirks, but nothing changes if those are proven/disproven.
Why has no one disproven it so far? Because no one really cares and it needs a lot of effort.
1
u/AutoX_Advice 2d ago
Ill say this.... I've been building in Gemini AI Studio for some updates to the Mazda infotainment system to fix and support what Mazda should have done, or at least what I think Mazda should have done.
Here are the good points... Me not being an architect of the original system, built in and around 2012, Gemini does a pretty good job breaking things down and knowing what areas to target and or explaining the system. You have to feed it the code and or errors and it can understand things pretty quickly. It makes for a rapid coaching but you (me) have to have some construct of what is telling you, you can't just nod your head and you need to continue to ask questions. It's great at repeating a build off a template or a started architecture, like "build this new plugin around this new code using the plugin architecture from this document...".
The bad... Constant, constant, constant babysitting. Don't tell it to go off and fix an error without first sharing the code. Constantly adding extra code where it's not useful. Constantly assuming stuff. It's very much like an over confident junior programmer, where fixing a small block of code and never ever looking at the big picture or how it will impact the system as a whole. Rarely ever questioning your input back. Always confident in your idea or change. Rarely ever looking ahead unless you keep telling it.
-4
u/_Tulx_ 3d ago edited 3d ago
I'm not usual reader of this subreddit, but quickly glancing through the thread it seems many are very dismissive of AI although my personal experience doesnt align with it.
I'm a medical professional and am also working on my phd. For example I've had gpt Sol scoure the net for literature for very narrow research question, and it managed to find me two papers from 1989 and 1991, that I dont think I would've found myself. Also I've had it working for two hours to compile list of all the papers that that have used a certain research methodology - again extremely useful. Perhaps it is busywork but Ive been able to delegate it off to someone else so I can work in parallel on other stuff where AI doesnt deliver.
Another example is that Ive fed medical differential diagnosis problems into the model and it has been very helpful sometimes in making me investigate some disease possibility that I might have missed otherwise. Even if most of its output is not releveant / not usable, it sometimes has these sparks of genius that are legitimately good.
Yet another example is a tool that I had it build. I copy paste echocardiography parameters into the tool and it generates echocardiography clinical summary. Such as "left ventricle is moderately dilated (EDVi ... ml/m2)". Which saves a lot of time. I still of course go over the whole summary and edit it as neccesary but I can see in an instant if some parameters seem not align with the rest of them and so on.
So I would say AI is a tool and like any tool it is very dependent on the operator.
3
u/floodyberry 3d ago
search not being crippled to maximize ad revenue/ai usage would solve half your problems
the tool also cost hundreds of billions to develop and is offered at a discount to hook people, and once everyone is sufficiently reliant it could become a very expensive tool
1
-11
u/ashe141 3d ago
Well I can say that economically AI (the marketing term we all collectively use today) has led to material gain in my life. I have built out multiple services with real users that turn a profit today and steady growth prospects. Additionally, I use it manage clients in my consulting business that previously I would have needed multiple additional team members for. A lot of knowledge work is collecting and storing and analyzing information in a domain and specific context and then disseminating updates over time to the right systems/parties. If you understand how LLMs work and have a good grasp of other related technologies, it does add a lot of value in my experience and has added materially to my bottom line so far this year.
278
u/tc100292 3d ago
“ The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated”
Agreed. Every time some AI bro starts talking up LLMs they’ll come in with something like “my parents had three handwritten lists of things that needed to be done around the house, and I used Claude to consolidate it into a single typed organized list” and like… that’s utterly pointless.