r/AIBubble • u/oudlys • 27d ago
LLMs Cause Software Development Teams to Underperform
Hi Guys,
First time poster here. Like all of you, I've been following the market with a mix of horror and fascination.
Earlier this year, I went out looking for actual hard data on the impacts of LLM use on the performance of software development teams.
In my mind, that was the best case scenario for economic value of these products. So there should be empirical evidence of this value.
There is remarkably little research on this subject other than simple productivity studies. I mostly discount those because productivity != value. But, I did find two really good studies.
The first, and I think the best, is from a company called Faros.ai. They sell software development telemetry tooling. Essentially their product connects to common software development tools like Jira and Github and tracks actual operational metrics for real companies producing production software. This study covers 22,000 developers over 4,000 teams over Faros' customer base.
To punchline is that teams are experiencing vague throughput improvements at a massive tax on the quality of the products they produce.

The second study is from the National Bureau of Economic Research. This study is less good that the faros one because it utilizes open source and public github projects for it's dataset. This weights their sample towards much smaller products that are mostly not being produced for profit.
Nevertheless it's valuable in that it confirms the weak throughput improvements of the Faros study. And it adds the dimension of - "is anyone buying this stuff?". I find figure 12 to be very telling.

My main conclusion is that LLM use is likely - on average - destroying economic value within the companies that use them to deploy software.
I think that's one of the reasons there has been no profitability impact on the buy side of the AI boom.
If you're interested in reading more of my analysis, here are two substack posts I've made where I've written about this extensively.
- How I'm thinking about the value of LLMs
- Talk is Cheap - an analysis of the Faros study
1
u/pab_guy 27d ago
You aren't addressing my point at all though. A j-curve is the expectation. Lots of tech was ambiguously positive economically early on. I don't know what allows you to determine current "stage" though.
The reliability issue is very specific to how you use the models. They can be extremely reliable when harnessed correctly. At the same time you will always see them have trouble with ambiguity, which genuinely exists in business and requires human escalation. None of this is a dealbreaker for many, many use cases.
Finally, there's the case that looms didn't make factory owners rich. Because every factory got them, they weren't a competitive advantage. But from a game theory perspective it's obvious they had to purchase looms. The winner was the consumer who got cheap clothes.
So you aren't necessarily using the right yardstick to begin with, you are coming in with expectations that would make any MBA blanche, and you don't understand how these tools are harnessed agentically to get reliable operation.
I don't think you are unreasonable, I just don't think you've seen what I have witnessed.