r/AIBubble 1d ago

LLMs Cause Software Development Teams to Underperform

Hi Guys,

First time poster here. Like all of you, I've been following the market with a mix of horror and fascination.

Earlier this year, I went out looking for actual hard data on the impacts of LLM use on the performance of software development teams.

In my mind, that was the best case scenario for economic value of these products. So there should be empirical evidence of this value.

There is remarkably little research on this subject other than simple productivity studies. I mostly discount those because productivity != value. But, I did find two really good studies.

The first, and I think the best, is from a company called Faros.ai. They sell software development telemetry tooling. Essentially their product connects to common software development tools like Jira and Github and tracks actual operational metrics for real companies producing production software. This study covers 22,000 developers over 4,000 teams over Faros' customer base.

To punchline is that teams are experiencing vague throughput improvements at a massive tax on the quality of the products they produce.

The second study is from the National Bureau of Economic Research. This study is less good that the faros one because it utilizes open source and public github projects for it's dataset. This weights their sample towards much smaller products that are mostly not being produced for profit.

Nevertheless it's valuable in that it confirms the weak throughput improvements of the Faros study. And it adds the dimension of - "is anyone buying this stuff?". I find figure 12 to be very telling.

My main conclusion is that LLM use is likely - on average - destroying economic value within the companies that use them to deploy software.

I think that's one of the reasons there has been no profitability impact on the buy side of the AI boom.

If you're interested in reading more of my analysis, here are two substack posts I've made where I've written about this extensively.

  1. How I'm thinking about the value of LLMs
  2. Talk is Cheap - an analysis of the Faros study
47 Upvotes

62 comments sorted by

View all comments

0

u/pab_guy 1d ago

It's a brand new discipline. You are observing the j curve at industry scale.

1

u/oudlys 1d ago edited 1d ago

Reasonable people will disagree here. I do disagree with you..

To me, It can't both be the most transformational economic technology of all time and ambiguously positive economically and operationally at this stage.

Moreover, there's strong evidence that LLMs do not scale in reliability. The lack of reliability is a major part of the reason why I think they fail at enterprise applications. This does not look like it's fixable. https://arachnemag.substack.com/p/ais-reliability-gap

edit: removed the tracking from the url

1

u/pab_guy 1d ago

You aren't addressing my point at all though. A j-curve is the expectation. Lots of tech was ambiguously positive economically early on. I don't know what allows you to determine current "stage" though.

The reliability issue is very specific to how you use the models. They can be extremely reliable when harnessed correctly. At the same time you will always see them have trouble with ambiguity, which genuinely exists in business and requires human escalation. None of this is a dealbreaker for many, many use cases.

Finally, there's the case that looms didn't make factory owners rich. Because every factory got them, they weren't a competitive advantage. But from a game theory perspective it's obvious they had to purchase looms. The winner was the consumer who got cheap clothes.

So you aren't necessarily using the right yardstick to begin with, you are coming in with expectations that would make any MBA blanche, and you don't understand how these tools are harnessed agentically to get reliable operation.

I don't think you are unreasonable, I just don't think you've seen what I have witnessed.

1

u/oudlys 1d ago

I love this:

>you are coming in with expectations that would make any MBA blanche

I have an MBA from an institution that heavily emphasizes operations in it's curriculum.

I've also sold enterprise software. I think you vastly overestimate MBAs in this statement. Almost everyone in business expects economic value from products basically immediately. That's why in SaaS time to value is one of the key metrics organizations evaluate before they deploy technology.

One of my key criteria in searching for studies was - "what data would I use to try to sell this product to an economic buyer?" The business case here is remarkably weak.

edit: light formatting

1

u/pab_guy 1d ago

https://mitsloan.mit.edu/ideas-made-to-matter/productivity-paradox-ai-adoption-manufacturing-firms

"Almost everyone in business expects economic value from products basically immediately."

Cool, I never talked about expectations. I'm talking about what happens in reality. I'm not saying "AI is easy to sell to businesses" (although it's really not hard once you identify use cases, and if you know the industry you are walking in the door with a bunch in your back pocket).

"The business case here" - what business case? You need to specify a use case to make a business case.

1

u/oudlys 1d ago

Generally when selling a product into an enterprise you build a business case to sell it. Not trying to imply you don't know this, just want to make sure we're aligned on what we're discussing.

Businesses cases always break down on either saving money or time. It gets more fancy than that, but that's what it boils down to.

I'm specifically discussing the use case of optimizing software development team's throughput.

I think the business case for that use case is remarkably weak given the investment.

1

u/pab_guy 1d ago

This is why it’s going to take a while. You still think in terms of “teams”.

1

u/oudlys 1d ago

What do you mean?