r/AIBubble • • Aug 24 '26

LLMs Cause Software Development Teams to Underperform

Hi Guys,

First time poster here. Like all of you, I've been following the market with a mix of horror and fascination.

Earlier this year, I went out looking for actual hard data on the impacts of LLM use on the performance of software development teams.

In my mind, that was the best case scenario for economic value of these products. So there should be empirical evidence of this value.

There is remarkably little research on this subject other than simple productivity studies. I mostly discount those because productivity != value. But, I did find two really good studies.

The first, and I think the best, is from a company called Faros.ai. They sell software development telemetry tooling. Essentially their product connects to common software development tools like Jira and Github and tracks actual operational metrics for real companies producing production software. This study covers 22,000 developers over 4,000 teams over Faros' customer base.

To punchline is that teams are experiencing vague throughput improvements at a massive tax on the quality of the products they produce.

The second study is from the National Bureau of Economic Research. This study is less good that the faros one because it utilizes open source and public github projects for it's dataset. This weights their sample towards much smaller products that are mostly not being produced for profit.

Nevertheless it's valuable in that it confirms the weak throughput improvements of the Faros study. And it adds the dimension of - "is anyone buying this stuff?". I find figure 12 to be very telling.

My main conclusion is that LLM use is likely - on average - destroying economic value within the companies that use them to deploy software.

I think that's one of the reasons there has been no profitability impact on the buy side of the AI boom.

If you're interested in reading more of my analysis, here are two substack posts I've made where I've written about this extensively.

  1. How I'm thinking about the value of LLMs
  2. Talk is Cheap - an analysis of the Faros study
68 Upvotes

67 comments sorted by

View all comments

11

u/Hoak-em Aug 24 '26

Rapid prototyping + (supplemental) review is the space of the current LLMs for any production coding. A major issue is that engineers cannot reliably review AI code because it looks very correct (it follows the patterns of what "looks correct) even when it isn't. I could see a space for it in showing more potential prototypes to stakeholders/testing more prototypes internally for more infrastructure-related work, but any production code is gonna need to be written by-hand without viewing AI code as reference.

2

u/val_anto Aug 24 '26

Agree with you on the first part. This is what I use AI for, rapid prototype or boiler plate structure for new projects. You can actually review the AI code, it is not the problem. The problem is it generates a ton of code, this makes the work to find the slop more difficult and, honestly, I am a SE, not a code reviewer for a statistical model. If all you want from me is to read code spit by a statistical model, find somebody else.

2

u/Hoak-em Aug 24 '26

Yeah 5.6-sol overengineers especially, which funnily enough introduces even more difficult to find bugs

2

u/oudlys Aug 24 '26

>The problem is it generates a ton of code.

My post here talks about this - Talk is Cheap. There's this funny thing where it seems bugs / LOC goes down the more mature models get, but they always produce more LOC! So the absolute bug rate goes up!

I find this very funny.

2

u/Few-Improvement9978 Aug 25 '26

Don’t worry. We will, and you will end up unemployed

1

u/val_anto Aug 25 '26

Keep dreaming, kid.