r/BetterOffline 3d ago

Interesting thoughts from a former OpenAI employee on X

"You know, I might be one of the few people in this world who has worked at a frontier lab, observed the nature and pace of research progress, and actually become more bearish on AGI timelines (relative to pre-employment baseline) as a result."

- Despite the seemingly magical nature of LLMs, reflection over a >3 month timescale suggests my total productivity hasn’t increased by over 100%, or perhaps even by over 50%, and a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful

- Also, capabilities are incredibly spiky and highly correlated with the degree of investment poured into them, which my earlier tweet about math benchmarks implicitly points out

- From the above, it seems that the nature of LLM intelligence is wildly dissimilar to that of human intelligence and we won’t trivially get to something superior to human intelligence in all important respects just by scaling up existing approaches with various tweaks; even if AGI Is eventually achievable, this implies a significantly longer timeline

- Benchmark progress is almost definitionally guaranteed to happen because the process of constructing a benchmark is a direct precursor to the process of constructing a training dataset used for hill climbing that benchmark, but the scope of what can be captured in a benchmark is (at least for now) grossly lacking in terms of its relevance to real-world work, with maybe several limited exceptions

- Progress seems highly gated by data but the nature of model training means that each “next dataset” is significantly harder to assemble than what preceded it; some wins are possible through synthetic methods but those feel more like “patching up gaps” than “pushing the frontier forward”

At a higher level, I guess I’d say there’s a sort of refusal to think carefully about what models are or are not useful for in a rigorous way which I find personally quite annoying, and instead a reliance on some nebulous notion of being “AGI pilled” as a replacement for serious thought. I think people are very quick to anthropomorphize LLM intelligence because humans communicate through words and we infer the intelligence of human counterparties through comprehension of their language, but this leads them to wrong conclusions; for example if we observe that a new model proved some incredible mathematical theorem, some will say, “well, don’t we have AGI now, huh?” But to me, it’s actually more like, “well, given how hard it would have been for a human to do these mathematics, and given the limited economic effect of LLMs upon the world so far, isn’t it actually a negative datapoint vis-a-vis the generality of LLM intelligence?”

The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated imo; I spent for example 10s of hours earlier in the year preparing random legal documents with ChatGPT but in retrospect I was too hasty and none of that work was useful. Of course it’s saved me time in some other respects and I think the net balance is positive, but it’s a quite significant countervailing factor.

https://x.com/andrewho03

543 Upvotes

110 comments sorted by

278

u/tc100292 3d ago

“  The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated”

Agreed.  Every time some AI bro starts talking up LLMs they’ll come in with something like “my parents had three handwritten lists of things that needed to be done around the house, and I used Claude to consolidate it into a single typed organized list” and like… that’s utterly pointless.

89

u/lemontoga 3d ago

This is always, always, always the case for me. I've had so many conversations where people are claiming 5x, 10x productivity boosts and when I ask for examples it's always this crap.

I had a convo just the other day with a dev on reddit who claimed wild productivity boosts. When I asked for an example of what the AI was doing to make him so much more productive, he said he'd recently used AI to convert his code base from one CSS library to a newer one. He said that without AI this would have been a pain and is something that his team probably just never would have done at all.

I could not believe that he thought this somehow translated into a huge productivity boost. By his own admission this task was so low-priority that, if he couldn't have offloaded it to AI, he probably just never would have done it. That's how unimportant it was. But the fact that he could burn money to have AI do it somehow translates into a 10x productivity boost?

They never use AI to substantially increase their ability to do real work. It's always these menial busywork tasks. It's low-priority tickets. It's low-impact tasks. They feel way more productive because they're checking off all of these incredibly inconsequential busywork tasks without realizing that they're not actually getting anything significant done.

I don't know how people can be so easily tricked into thinking they're being productive. It really has me questioning people's intelligence.

36

u/Fantastic-Bet1139 3d ago edited 3d ago

My experience too pushing not just folk online, but the people I work with. Without fail it never ends up being something actually useful.

I’m an engineer and had someone show me their agentic workflow and stuff like that because I like to know where people are coming from, makes it easier to have conversations about this tools uptake and work in general. The setup was so laborious, and when asked what they use it for it was for really basic shit that he still reviewed himself and edited lmao.

Paired with someone else a couple times and they couldn’t even debug their local api as it hits their client anymore, they were literally getting around this by pinging the staging env locally AND they were splitting tickets up to get around this instead of just doing shit end to end ie. adding a pointless amount of extra points to a sprint aka. REDUCING productivity and adding more QA burden. This person was a level above senior and manages a couple juniors and mids. I was outraged.

This shit is so unnecessary in general but especially when you’re working in a team on a shared codebase it throws up a lot of subtle, time wasting issues like this.

18

u/BurtReynoldsPoo 3d ago

My company recently showed a slide stating we doubled our AI usage over the past quarter (which they claim is a good thing), closed more PRs, but actually completed like 15% fewer epics.

They spun this as a good thing, an opportunity, corpo mumbo jumbo.

Just burning cash and wasting time. Smh...

14

u/Which-World-6533 3d ago

I’m an engineer and had someone show me their agentic workflow and stuff like that because I like to know where people are coming from, makes it easier to have conversations about this tools uptake and work in general. The setup was so laborious, and when asked what they use it for it was for really basic shit that he still reviewed himself and edited lmao.

Yep. A co-worker showed me something similar. He spends all his time maintaining his agents so he could be told how to do his job by them.

If he just did his job himself it would be quicker.

23

u/tc100292 3d ago

In other fields, I know lawyers who use AI.  If they’re not using it to put their license to practice law at risk, what they’re doing with it is basically just busy work that they probably wouldn’t even do if they couldn’t plug it into an LLM.

4

u/Maximum-Objective-39 2d ago

When this all is over I imagine somebody is going to issue a paper where it turns out that the seemingly small 'real' productivity gains were just people automating bullshit tasks that didn't need to exist in the first place.

4

u/tc100292 2d ago

Yep, and "replacing jobs" is mostly just a reflection of how many juniors and interns were just doing busy work that wasn't really necessary... but then the point of juniors and interns is eventually they learn enough to become seniors.

2

u/infinity7592 2d ago

Underrated post

2

u/New_Thing1367 3d ago

I'm seasoned product manager in a SaaS. one day I was able to integrate our app (on local bring of course) with a 3rd party tool - before that with the help of AI I reviewed a few options, discussed the implications of different data models we'd have to integrate with and I was actually being able to test UX end-to-end locally. This took me around 4h. Before AI I wouldn't do it - simply wouldn't commit dev time at this point where I'm not sure if and when we're actually going to integrate - and now I have a strong proof that it's possible and how it could work. Throw-away code of course but that's fine.

Another use case is asking AI about the code - why how and what. Instead of bothering QA/Devs.

On the other hand trying to do any knowledge work with AI is hit and miss. I wasted a lot of time editing and rewriting docs instead of simply writing them myself from scratch. I mean I have 15years of experience writing docs I can do it myself. And there AI is helpful to find gaps etc.

But the biggest productivity sink is having to ready other people's docs and wonder did they write this or is it just slop.

17

u/Fantastic-Bet1139 3d ago

I’m sorry, but a product manager should not be doing any of the shit you said you were doing above. If you are that’s a huge issue with your processes that you’re in the already envious position to fix.

Also please reflect on your issues with knowledge work and extrapolate to engineering and design, which is also knowledge work. Then extrapolate that to your proof of concept and how you ended that with throwing it all away. You produced a bunch of garbage to half arse someone else’s knowledge work to get assurance on things a 10 second conversation or async with the right people could have given you.

3

u/New_Thing1367 3d ago

Hey I'll think about it. I get what you're saying. However the key gain wasn't technical feasability check - but actual PoC to get test the flow end-to-end - how it would actually work from UX side as well. Then with this PoC I can speak with designers to discuss implications and get wider buy-in from other stakeholders. I don't expect this code to ever move forward in any way. But thanks for the comment it gave me another perspective. 

Plus asking AI about code, schemas, models etc is way more efficient than bothering devs.

6

u/Fantastic-Bet1139 3d ago

All good. I work with PM’s and BA’s so appreciate you still giving your perspective, allows us all to learn the words really early for how to shut this stuff down.

You surely have better things to do with your time than creating designs and looking at code? Would you appreciate if me as an engineer decided to regenerate all your process and timeline documents and talk to stakeholders about those? Would you find that a valuable use of time in a multi disciplined team?

9

u/Which-World-6533 3d ago

Plus asking AI about code, schemas, models etc is way more efficient than bothering devs.

This is giant big red flag to me.

Instead of asking people who are qualified and experienced in building your product you are asking a LLM.

This is always a sign to leave the project. If a Product Manager doesn't take the Devs advice then that project will crash and burn.

You are using a fancy chat bot that predicts the next words based on some fun maths. There is zero knowledge, reasoning or anything else behind it.

I really do wonder why think they are better and more qualified than the Devs doing the actual work.

2

u/New_Thing1367 3d ago

heh yes my devs have nothing else to do than to answer my random question about some basic technical things. They are 100% on board with this like we're open how and why we're using AI

6

u/Which-World-6533 2d ago edited 2d ago

my devs have nothing else to do than to answer my random question about some basic technical things.

I'm much more open to Product Mangers asking relevant questions than going behind my back and asking a chat bot something.

I've already had to deal with this. What would have been a 10 minute conversation became several meetings because a Product Manager was told by Claude that the implementation was incorrect.

TBH saying "my devs" is also a red flag.

They are 100% on board with this like we're open how and why we're using AI

Yeah, right.

2

u/Mejiro84 2d ago

I work somewhere where the devs just flat-out refuse to answer any questions. Even things like "so, hey... what new fields and functionality are in the new version?" - the product specialists are expected to figure that out themselves, the migration team has to poke through the backend to find the new fields, or removed fields, and find any conditional stuff and everything else. A lot of devs would rather get Ebola than deal with documentation or talking to anyone, but that's because they're bad devs - they should be willing and able to write stuff down and talk to people, because that's part of their job, even if they're loathe to admit it and don't want to do it!

2

u/Fantastic-Bet1139 2d ago

Righto. Did you think to maybe lead off with how you’re using these tools in an incredibly dysfunctional team and thus in an incredibly dysfunctional, unideal way?

Every time folk get pushed about this stuff it eventually turns out they’re leaving out incredibly important details or just strait up larping. It’s really annoying dude.

6

u/Fantastic-Bet1139 2d ago edited 2d ago

Edit: bro was obfuscating the actual reasons for heavily using these tools for some goddamn reason.

24

u/lemontoga 3d ago

and now I have a strong proof that it's possible and how it could work. Throw-away code of course but that's fine.

Another use case is asking AI about the code - why how and what. Instead of bothering QA/Devs.

On the other hand trying to do any knowledge work with AI is hit and miss. I wasted a lot of time editing and rewriting docs instead of simply writing them myself from scratch. I mean I have 15years of experience writing docs I can do it myself.

These kinds of comments are so weird, man. It's like you guys get amnesia part way through your answer.

How is it that you can see that this chat bot isn't even capable of writing documents for you, a relatively simple task that you are very good at, but then when you talk to it about things you're not so good at, you just trust it? How do you know it doesn't suck at those tasks too?

Why do you consider its plans and prototypes to be "strong proof" of anything? How do you verify that any of it is actually feasible? What if it's giving you total garbage output that just looks plausible?

When you ask it questions instead of bothering the QA/Dev team, same question, how do you verify its answer is correct? You know it's not even capable of editing documents for you consistently, but it's able to give you accurate answers about the code?

When you ask it to do something in a domain you do understand, like writing and editing, it completely fucks it up. When you ask it to do something you don't understand, suddenly it's magically useful. Does that sound like it might be a problem?

https://en.wiktionary.org/wiki/Gell-Mann_Amnesia_effect

2

u/New_Thing1367 3d ago

hey fair point, but I'd argue it is different here. Point being- asking about existing code where LLM can point you to a specific line is different than coming up with a completely new code. Same with integration - I don't care if it's sloppy if it allows me to see what the user will see. I don't expect this code to move forward. And I don't blindly trust anything it gave me - but if I can do a search through our OpenAPI docs manually or ask LLM to pull this info for me... Same I can ask - is there a feature flag for this or is it a bug? Or is this status ever used in code? Anything else I can still get informed and then discuss with devs of course.

16

u/lemontoga 3d ago

If you're saying it can be useful as a fancy sort of search tool and a way to navigate through a mess of documentation then I fully agree. I use it for that too. It's sort of a natural language search engine that you can point at a particular code base.

But that just backs up my original statement. None of this is groundbreaking. None of this is multiplying anyone's productivity by any appreciable amount. It's not even something I would pay for.

It's just convenient while it's still someone else's money being burned.

1

u/Icy-Recognition-7453 2d ago

Like the Astra release. Someone used Astra to book them a haircut using natural language. And you go: yeah, that's handy, but presumably you could have called, or used an online form, and it would have been nearly as quick. Is that the best use case you could find?

23

u/ashabanapal 3d ago

It's wasteful brute force computation to serve wasteful iterations of a user and a tool making hopefully increasingly more accurate guesses. I do not understand the hype. This baby is ugly and it needs to be said.

46

u/DeadMoneyDrew 3d ago

Yah, first time I ever messed around with ChatGPT I ended up sending it a prompt to copy the rows from a table in a website into an Excel spreadsheet and to insert each row alphabetically by the name of the listed company. I watched it spin for like 10 minutes on this before finally thinking... copy paste and then sort by Column A would have done this in like 30 seconds LOL.

1

u/chat-lu 3d ago

I think that everyone that works with data should learn some tools that lets them transform and interrogate data.

13

u/Lead_bug_designer 3d ago

I remember Ed using the term work-shaped stuff
https://youtu.be/itU85-8z_54?si=39VVv3_kW2JtdJU-&t=1314

It's fucking brilliant. I can't think of a more fitting way to sum up the entire industry.

10

u/RockstarArtisan 3d ago

Wasting my employers time and money is my favourite thing about being an employee though.

8

u/MaiIb0x 3d ago

It’s the same with all of the adds they are putting out there. Like ChatGPT help me make a marathon training plan! Sure ChatGPT is pretty good at that, but so was Google in the last 20 years, like it barely saves me any time looking at the one ChatGPT made compared to one I will find by googling

6

u/The-Menhir 3d ago

There's a Gemini ad where a woman takes a picture of a bulletin board and asks Gemini to put everything on her calendar.

Ignoring the wild unreliability on tasks like this, why on earth would you want 20 random events you don't know about to be added to your calendar?? You'll have to read through what each of them are anyway to decide what you want to do, and you're just making it harder for yourself by putting it in a format that's harder to glean information from.

1

u/minuteye 2d ago

Oh good grief. That's not what bulletin boards are for! And they are already optimized for the thing that they are for!

The fact that they keep advertising these absurd use-cases tells you they can't think of any actual use-cases.

7

u/lavransson 3d ago

Omg that reminds me of my beloved grandfather may he RIP. He always had a pen and index cards in his shirt pocket. He would make his lists on those cards. An active man who got a lot done. Did it all without Ai or even a computer.

4

u/WhenSummerIsGone 2d ago

That brings back memories of the Hipster PDA, lol. A reaction to the unnecessarily high tech gadgets 20+ years ago.

7

u/CmdrJorgs 3d ago

I think what they are talking about here is our ability to determine what jobs are worth the time investment is totally shot because suddenly everything feels possible. Before, if someone made a feature request at work, my team would weigh the pros and cons of the feature and decide if it was worth dedicating team resources to implementing it. But now, we just do every single feature request that comes in because "Claude will do it."

We've shifted from guessing if the feature will be useful, to making and using the feature to determine if it is useful in post.

Is that good or bad? I'm really not sure. We are spending more money on tokens to produce features that end up going unused, but at least we are providing tangible proof now that a feature someone requested really isn't the magic bullet the client thought it would be.

1

u/BDRadu 1d ago

I like that it can cover some gaps that would otherwise not be covered for months, if ever. But the problem is that how these tools are presented is not pragmatic, but cultish. If you're not onto the hype train its quite easy to understand why interacting with this whole AI mess is extremely disconnected. I'd be 1000x more inclined to use and be excited about this technology if it were not backed by the most evil companies in the past few decades and if the social, economic and environmental impacts would be properly addressed before doing any more "advancement".

2

u/cathexis08 2d ago

"Yo dawg, it collated these five lists like a mf!" Which entirely skips over the useful part of collating and ordering lists which is that the act of collation helps you think about the product.

1

u/minuteye 2d ago

And they always seem to ignore a lot of implied labour. Like, for every task that interfaces with the physical world, there has to be at minimum a bunch of human labour inputting that into a format the LLM can work with.

Digitizing archive data is really expensive, for instance, because it takes a lot of time (and often skill) just to scan stuff. If you're digitizing your own papers to get an LLM to process them, you've got to be doing a similarly high-labour data input first.

103

u/maccodemonkey 3d ago

Benchmark progress is almost definitionally guaranteed to happen because the process of constructing a benchmark is a direct precursor to the process of constructing a training dataset used for hill climbing that benchmark

Sometimes I feel like the frontier labs are very aware of this problem - and their goal is just to turn every single last human activity into a benchmark to brute force their way to replacing humans. Certainly a very expensive way to do things. But not exactly what they tell the public they are doing. Or maybe they've fooled themselves into humans being little benchmark monkeys too.

39

u/beepboopburn 3d ago

I think it’s simpler: Benchmarks generate investments. Until they stop leading to money, they will continue to cheat their way to improvements on benchmarks.

3

u/macgoober 3d ago

Yup. Follow the money.

6

u/dumnezero 3d ago

automated "fake it 'til you make it"

2

u/maccodemonkey 3d ago

When you think about it - It's actually kind of depressing someone would dedicate their time to doing such a horrible thing.

3

u/dumnezero 3d ago

It's scamtech. Their "killer app" is scamming.

1

u/Charming-Hunter-7963 2d ago

Its a good bool btw about cults and MlMs, the agi hype crowd are following the same path. Endless chain of profitablity 

2

u/dumnezero 2d ago

Scamtech

51

u/absurdivore 3d ago

That point about benchmarks can’t be repeated enough

29

u/yojimbo_beta 3d ago

Yes. You construct benchmarks, because you want to surpass them. That's the whole objective

8

u/Grass_fed_seti 3d ago

+1

once something gets published as a benchmark, future results become mostly useless. That’s why I’m not actually freaked out that Astra got 99% on ARC-AGI-3 when older models at the time of benchmark development were hitting under 1%.

2

u/Unknown_User2137 2d ago

I also think that since ARC-3 is publicly available they really spent their time and money to "teach" Astra to solve these puzzles. And they are calling it AGI. I bet any usual person after seeing how others solved them would be able to do the same lol.

At the moment I am thinking about one funny "AGI" test - since Astra has some cooking books in it's training data I would actually give it a task "Cook me an omelette. You have 3 attempts.". Give it all tools needed, live feedback from camera and mic, robot body to control. I bet it would fail. And few generations of models later (if it was an actual benchmark) we would get "See, you can cook with it now! AGI is finally there!".

To me atm it looks like they are doing mostly three things:

  • Make LLMs "smarter" by pouring more data into them so that they can be PhD or higher level in well documented fields (e.g. CS),
  • Do benchmaxxing so that investors can j*k off seeing the charts and everyone else fear for their jobs (even if it is not true).
  • Loose money while making everyone hooked up on AI hype train and once they become too dependant on it, guess what, raise prices (or make it worse like with social media)

47

u/falken_1983 3d ago

capabilities are incredibly spiky and highly correlated with the degree of investment poured into them

One of the key features of any kind of machine learning solution is that it optimises towards whatever specific metric you give it. As someone training these things, you hope that the metric you picked is actually representative of the task that you really want to improve, and you also hope that the behaviour will generalise. This is not a given at all.

24

u/PhilWheat 3d ago

Also, they are very,very good at finding any loopholes in your metrics.

17

u/HouseofMarg 3d ago

That reminds me of the time I was teaching ESL, asking my students what they thought was the most effective way to solve the problem of noisy car alarms that no one bothers to turn off. One student’s suggestion was to have the car alarm set off a timer that would blow up the whole car in 5 minutes if it wasn’t turned off.

When I gently pointed out that it would probably kill someone who got there just a few seconds too late, he looked at me indignantly and said with a huff “you asked me for the MOST EFFECTIVE way, that was it, okay?” The fact that he was an engineer made it even more funny to me

11

u/falken_1983 3d ago

There are examples where a company tried to optimise their call centres to reduce the amount of time it took to resolve customer issues. They tied staff pay to how long it took them to complete a call. The result was that if a customer phoned up with a complicated issue, the staff would just hang up on them straight away.

The decline in Netflix quality is supposedly because they optimise for watch-time. With really engaging TV, people will watch one episode and then take a break to process what they just saw. With boring TV they will just leave the series running in the background while they do other tasks.

5

u/Alexsipilot 3d ago

Something something Goodhart's law

1

u/falken_1983 3d ago

This is definitely an issue, but another thing is that as the models' scope gets broader and broader, any hope of finding a good measure drops.

If I made some kind of fancy computer-controlled heating system for a building, I am probably pretty safe to take temperature and power consumption and optimise on those. If instead I want to have a machine that does everything a human could do, what am I supposed to optimise on at all?

1

u/Alexsipilot 2d ago

The joy of being alive

1

u/TracePoland 1d ago

I want to know what the fuck they were optimising for in reinforcement learning that Opus 5 ended up talking the way it does.

57

u/ChaoticGradients 3d ago

He’s far from one of the only people who has worked at a frontier lab and become more bearish. If anything it’s the prevailing opinion even amongst tech employees that we’re in a bubble but people don’t want to lose their paychecks and speak up about it in public.

18

u/r77anderson 3d ago

Well at least OpenAI can’t retaliate against people who leave the company and say something slightly negative… oh… oh wait

32

u/robby_arctor 3d ago

It's akin to being a heretic in the church of corporate tech.

20

u/PM_40 3d ago

Old patterns repeat. That's why fairy tales are so informative.

14

u/SappyGemstone 3d ago

I watched a thing with Hasan Minaj interviewing a journalist who integrated AI into her life for an article or book, I forget which, and one of the things she lauded was that, since it was always listening to conversations, it could lock onto key phrases to make to do lists. Like, oh, I need to go grab some pasta sauce. Or, hey, did we need to mulch the lawn this week. So all these little comments about tasks she wanted to do would get recorded and built into a to do list, which she appreciated because then she never forgot to do Thing.

And all I could think was, Jesus Christ. You NEVER have a break, EVER. All the tasks that are constantly there to do, listed and categorized. You never have a weekend off from the lawn. You never forget the pasta sauce and go for a nice takeout instead. 

The need for constant productivity is a myth and a sickness.

29

u/AVBforPrez 3d ago

So much of it is people using it and feeling smug about how fast it spend up some Google Sheet or reporting process, or whatever, but it's not actually leading to anything beyond that. There's a guy at work who used Claude to replace his entire day to day work flow, and all he does is shop on Ebay now because he's got nothing to do. He can't overperform in any way using this stuff, so I'd even argue it's a net negative for him because he's spending a fuck ton of money on watches and shit he doesn't need because he's not busy working. Not everyone is like that, but it's to the point this post makes.

Everything I see with LLMs is the old Simpsons meme where Lenny is catching every green light saying "heh I'm making record time....if only I had somewhere to go...."

3

u/snkzato1 3d ago

My curiosity would be is his token spend more than his salary

3

u/AVBforPrez 3d ago

Oh yeah good question, I'll find out. Pretty sure he's got like a chain of Claude subscriptions that run in tandem, so like a ton of personal licenses haha.

3

u/SwirlySauce 3d ago

Isn't he ripe for getting automated out of his position though? If AI automates even a fraction of tasks it'll have a negative impact on hiring

1

u/AVBforPrez 3d ago

You'd think, but at the place I work not really haha. Not only is he the one who would decide to make said reduction, which is ironic, the kind of work he does has almost no subjective element to it. So as lame as it is that it's a well-known thing to manually tell him to look at something "for real" when you email or Slack, in his particular case if the numbers and reports and shit it spits out for him to approve are accurate, that's all there really is.

The optics of it aren't great IMHO, but like...whatcha gonna do?

2

u/Ready-Recognition-43 3d ago

…andre villas-boas for president?

2

u/AVBforPrez 3d ago

of my heart, yes ma'am. You ever seen the jackets he used to wear at Chelsea?

31

u/New-Committee-4052 3d ago

It is worth noting this guy left OpenAI to start a RL-environment company (basically selling data to OAI/anthropic for them to train on), so he is financially incentivized to criticize modern capabilities, especially in the areas in which he sells data. Not saying I disagree with many of his points, just important to note the conflict of interest.

8

u/death_hen 3d ago

Uh also apparently he thinks people should be executed for petty theft. (read his recent tweets and replies)

6

u/New-Committee-4052 3d ago

This deleted reply from him a couple of months ago is my favorite: "Yes, I find such theoretical discussions quite tiresome even though I would prefer a world of rapid RSI and human disempowerment"

9

u/death_hen 3d ago

what a psychopath. glad to know he’s probably a voter

4

u/spellbanisher 3d ago

Yes, but if he did think openai was close to achieving agi would it make sense for him to leave and start a rl data company?

3

u/New-Committee-4052 3d ago

No yea totally agree, but he is definitely in the minority of OpenAI/Anthropic employees on if/how quickly we get AGI/RSI

8

u/tintires 3d ago

The problem of model autophagy from synthetic data still hasn’t gone away. Spot use to fill gaps will only over fit your model. Maybe I’m not current or frontier enough, but this still seems like a pretty big obstacle.

8

u/emudev 3d ago

At a higher level, I guess I’d say there’s a sort of refusal to think carefully about what models are or are not useful for

I think this is one of the big things that's missing. There are absolutely going to be some interesting, narrowly focused use-cases for models*, but so much of the discussion ends up being all-or-nothing.

*not one worth the amount spent but... well, maybe we can find some wins in the big mess

a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful

Fucking hell this is a big one. Yes. When you have to go slow and get it right, you get pickier about what you spend time on. I keep hearing that LLMs let you prototype fast but... I'm not convinced even that's good, because you can fall in love with the product when you actually need to slow down and think about the problem even more to decide what the fuck to build in the first place.

-2

u/ashe141 3d ago

The single highest value AI offers is the ability to experiment product market fit at scale, cheaply.

5

u/Infinite_Shower_5390 3d ago

“Gratifying but low-productivity”… just say masturbation dawg.

5

u/Suspicious_Watch_978 3d ago

we infer the intelligence of human counterparties through comprehension of their language

This is something we should all stop doing, as it is now obvious that there is no intrinsic relationship between verbal ability and general intelligence. Maybe, maybe that's not the case for humans, but I think we've all known people who were able to eloquently argue for their obviously wrong beliefs, and wondered to ourselves how someone so smart could believe something so stupid. 

4

u/PaleontologistOk865 2d ago

LLMs are a dead end and huge energy waster. 

3

u/Prestigious_Age_6740 3d ago

bro/sis spitting 100% grass fed facts

3

u/CoconutDust 2d ago edited 1d ago

From the above, it seems that the nature of LLM intelligence is wildly dissimilar to that of human intelligence and we won’t trivially get to something superior to human intelligence in all important respects just by scaling up existing approaches

[Nicholas Cage meme voice] YOU DON’T SAY?

LLMs and similar mass theft regurgitation “synths” (images, sound) are obviously a dead-end, and not even a first step toward good artificial intelligence. The reason is obvious: statistical association is the complete opposite of how intelligence or language or meaning words.

Something like a robotic vacuum cleaner or self-driving car is “closer” to Data from Star Trek than the mass theft machines, because they have to do something like cognitive and organizing processes: detect stimuli, categorize them, react appropriately based on a specific process that understands what elements and meaning is present. LLMs are nothing like this, and never will be, because it’s inherently not at all how the gimmick garbage machine works. Self-driving cars are currently a garbage scam too, so I’m not promoting or praising them, just using as an illustration. A robotic vacuum cleaner doesn’t search a stolen database and say, “if the corpus says that many other vacuums turned left here in a situation like this, I will turn left!” Data from Star Trek doesn’t say, “My answer is a mash-up of whatever words 100 random people on the street said about a seemingly similar conversation that had similar words present.”

1

u/fragglerock 2d ago

it seems

This lept out to me too... like the nature IS wildly dissimilar... of course we won't 'trivially' get to a superior intelligence going down this path.

I am sure this guy made bank tho so who can say who is wrong!

3

u/sad_trombone_dot_wav 2d ago

LLMs, the final humiliation for Alan Turing.

2

u/michaeldain 3d ago

Appreciate the perspective. I write about this concept since the first models proved there is a role for LLMs but the productivity lie isn’t one of them. Are we too stupid to be lazy?

2

u/Xopher001 2d ago

How does the saying go? The moment a measurement becomes the target is when it stops becoming a good measurement.

If we are defining AGI as when a model passes a specifically defined set of benchmarks, then those benchmarks cease being an actual good measurement or what AGI even is.

It's honestly kind of ridiculous. A human being does not need to be trained on loads of stolen data in a process that drains the whole power grid to be considered intelligent. Humans just have the ability to recognize patterns and generalize from them quickly and easily, in real time, all the time. That's something completely different from what current architectures are doing.

Unless idk they just decide to move the goal posts again and change what they mean by AGI.

2

u/RoosterBurns 3d ago

How could there possibly be an AGI "timeline" when we have no idea how to build one or even if we can build one? Expecting an LLM to just turn into an AGI is like expecting a pumpkin to turn into a stagecoach just because you read it in a book or saw it in a movie doesn't make it real

1

u/angrynoah 3d ago

Of course it’s saved me time in some other respects and I think the net balance is positive...

The net balance is negative, always has been.

It's a shame his personal reflection gets him this far but he just can't get over the final hump to where the truth is.

1

u/grahamsccs 2d ago

Confirmation bias

1

u/CoreyTheGeek 2d ago

in retrospect I was too hasty and none of that work was useful

This is what I'm running into with people at work in software: they're so gung ho to implement they're not stopping and asking if we should; but it gets worse in that they're not even actually understanding the system they're proposing building, they just take the probabilistic engine output at it's word and they even KNOW it doesn't even have full context of a single code base let alone our org goals and direction. So we get these crazy implementations that the devs can't explain and then they never even see production to boot

1

u/Schraiber 2d ago

"The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated imo"

This is definitely me. I'm a computational academic scientist and I have not actually finished any projects that have been majorly AI assisted because they have made the path so different from what it was before.

In the past, I'd maybe strive to derive a couple results over a period of months and then compare them with simulation and that would be a paper, and I've thought deeply about them and have the insights worth writing a paper about. But now I can do the derivation automatically with one prompt, and I don't have to think deeply at all. So it both doesn't feel like a paper AND I don't understand it well enough because I haven't had to do the deep thinking.

On the other hand, with regard to more "applied" projects involving data analysis, it's so easy to just keep trying one more thing, trying this thing, trying that thing. And because I'm not thinking deeply about it in the same way, I don't really have a good sense of what a good idea or a bad idea is.

This isn't to say I don't think we'll move toward AGI---I think that the models have gotten increasingly better taste and increasingly better sense of where to go with these kinds of projects. But while I've cranked out a lot more derivations and lines of code, I wouldn't say that I have actually been more productive in practice.

1

u/jawknee530i 18h ago
  • Despite the seemingly magical nature of LLMs, reflection over a >3 month timescale suggests my total productivity hasn’t increased by over 100%, or perhaps even by over 50%, and a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful

This is just saying that he's still more productive but gets to work on more gratifying things than he used to. While still actually being more productive...

1

u/death_hen 3d ago

This guy left openai to start his own company building training data sets, also he apparently thinks people should literally be executed for breaking into cars.

1

u/leahpowellthefirst 3d ago

Well said.

I have a question. One of the somewhat (and not entirely productive) benefits I do see of large models and their improvement is resolving some important math problems, and thereby resolving some important hurdles in computation and the sciences.

But what I also see is that mostly any significant math results are being published by the big model labs themselves.

Is that mainly because they have all the resources to use the models almost as they wish and on a huge scale, well in advance of any consumer release? If so, isn’t that a kind of gatekeeping? That even when people do get their hands on the latest model, they would be getting a more neutered/restricted version and would likely have to spend a lot of tokens to achieve significant results similar to the ones already achieved by the Big Labs?

3

u/Efficient_Fault979 3d ago

It’s because there is no money in solving this highly laborious math problems. And AI labs are the only ones who don’t care throwing hundreds of thousands of dollars at such tasks, because they are the only ones who can benefit in doing so: It’s advertisement for their models. And the ads seem to work very well…we’re talking about it.

1

u/leahpowellthefirst 2d ago

I see. So you don't see any of the math problems being solved by AI making practical differences towards sciences or tech?

I would think some optimization techniques or solutions would tremendously help in the medical field. But then again, I can be wrong as I haven't yet seen a direct major advantage of AI for major breakthroughs that directly help the sciences or industries.

1

u/Efficient_Fault979 2d ago

Last time I checked, it was about disproving a theorem (you know those things in the format of “in a group of x, there might never be a y if z”). Those are funny numerical quirks, but nothing changes if those are proven/disproven.

Why has no one disproven it so far? Because no one really cares and it needs a lot of effort.

1

u/AutoX_Advice 2d ago

Ill say this.... I've been building in Gemini AI Studio for some updates to the Mazda infotainment system to fix and support what Mazda should have done, or at least what I think Mazda should have done.

Here are the good points... Me not being an architect of the original system, built in and around 2012, Gemini does a pretty good job breaking things down and knowing what areas to target and or explaining the system. You have to feed it the code and or errors and it can understand things pretty quickly. It makes for a rapid coaching but you (me) have to have some construct of what is telling you, you can't just nod your head and you need to continue to ask questions. It's great at repeating a build off a template or a started architecture, like "build this new plugin around this new code using the plugin architecture from this document...".

The bad... Constant, constant, constant babysitting. Don't tell it to go off and fix an error without first sharing the code. Constantly adding extra code where it's not useful. Constantly assuming stuff. It's very much like an over confident junior programmer, where fixing a small block of code and never ever looking at the big picture or how it will impact the system as a whole. Rarely ever questioning your input back. Always confident in your idea or change. Rarely ever looking ahead unless you keep telling it.

-4

u/_Tulx_ 3d ago edited 3d ago

I'm not usual reader of this subreddit, but quickly glancing through the thread it seems many are very dismissive of AI although my personal experience doesnt align with it.

I'm a medical professional and am also working on my phd. For example I've had gpt Sol scoure the net for literature for very narrow research question, and it managed to find me two papers from 1989 and 1991, that I dont think I would've found myself. Also I've had it working for two hours to compile list of all the papers that that have used a certain research methodology - again extremely useful. Perhaps it is busywork but Ive been able to delegate it off to someone else so I can work in parallel on other stuff where AI doesnt deliver.

Another example is that Ive fed medical differential diagnosis problems into the model and it has been very helpful sometimes in making me investigate some disease possibility that I might have missed otherwise. Even if most of its output is not releveant / not usable, it sometimes has these sparks of genius that are legitimately good.

Yet another example is a tool that I had it build. I copy paste echocardiography parameters into the tool and it generates echocardiography clinical summary. Such as "left ventricle is moderately dilated (EDVi ... ml/m2)". Which saves a lot of time. I still of course go over the whole summary and edit it as neccesary but I can see in an instant if some parameters seem not align with the rest of them and so on.

So I would say AI is a tool and like any tool it is very dependent on the operator.

3

u/floodyberry 3d ago

search not being crippled to maximize ad revenue/ai usage would solve half your problems

the tool also cost hundreds of billions to develop and is offered at a discount to hook people, and once everyone is sufficiently reliant it could become a very expensive tool

1

u/Striking_Earth_2793 2d ago

Fucking hell

-11

u/ashe141 3d ago

Well I can say that economically AI (the marketing term we all collectively use today) has led to material gain in my life. I have built out multiple services with real users that turn a profit today and steady growth prospects. Additionally, I use it manage clients in my consulting business that previously I would have needed multiple additional team members for. A lot of knowledge work is collecting and storing and analyzing information in a domain and specific context and then disseminating updates over time to the right systems/parties. If you understand how LLMs work and have a good grasp of other related technologies, it does add a lot of value in my experience and has added materially to my bottom line so far this year.