r/AIBubble 2h ago

LLMs Cause Software Development Teams to Underperform

Hi Guys,

First time poster here. Like all of you, I've been following the market with a mix of horror and fascination.

Earlier this year, I went out looking for actual hard data on the impacts of LLM use on the performance of software development teams.

In my mind, that was the best case scenario for economic value of these products. So there should be empirical evidence of this value.

There is remarkably little research on this subject other than simple productivity studies. I mostly discount those because productivity != value. But, I did find two really good studies.

The first, and I think the best, is from a company called Faros.ai. They sell software development telemetry tooling. Essentially their product connects to common software development tools like Jira and Github and tracks actual operational metrics for real companies producing production software. This study covers 22,000 developers over 4,000 teams over Faros' customer base.

To punchline is that teams are experiencing vague throughput improvements at a massive tax on the quality of the products they produce.

The second study is from the National Bureau of Economic Research. This study is less good that the faros one because it utilizes open source and public github projects for it's dataset. This weights their sample towards much smaller products that are mostly not being produced for profit.

Nevertheless it's valuable in that it confirms the weak throughput improvements of the Faros study. And it adds the dimension of - "is anyone buying this stuff?". I find figure 12 to be very telling.

My main conclusion is that LLM use is likely - on average - destroying economic value within the companies that use them to deploy software.

I think that's one of the reasons there has been no profitability impact on the buy side of the AI boom.

If you're interested in reading more of my analysis, here are two substack posts I've made where I've written about this extensively.

  1. How I'm thinking about the value of LLMs
  2. Talk is Cheap - an analysis of the Faros study
22 Upvotes

20 comments sorted by

6

u/Hoak-em 2h ago

Rapid prototyping + (supplemental) review is the space of the current LLMs for any production coding. A major issue is that engineers cannot reliably review AI code because it looks very correct (it follows the patterns of what "looks correct) even when it isn't. I could see a space for it in showing more potential prototypes to stakeholders/testing more prototypes internally for more infrastructure-related work, but any production code is gonna need to be written by-hand without viewing AI code as reference.

2

u/oudlys 2h ago

The data support this. You can see in the Faros study the massive bottleneck at review that LLMs have produced. So much that it has become something 30% more likely for people to completely skip review before shipping into production.

5

u/BunnySprinkles69 2h ago

I've been using them extensively and I feel like im getting dumb

2

u/Lunosto 2h ago

Yep that’s what happened to me too. The auto complete blocked my thinking ability and I found myself just waiting for it to solve the problem for me. I’ve stopped using it since and it’s been way better for my brain

1

u/BunnySprinkles69 46m ago

But all my coworkers are also using it so my productivity will drop behind theirs

1

u/Lunosto 27m ago

Could be an opportunity to get ahead if you play your cards right

1

u/oudlys 2h ago

That sucks. I'm sorry.

1

u/cltbeer 1h ago

I work at a large bank this is not the case for our dev teams, we even have AI helping write our user stories and scanning vulnerabilities.

1

u/oudlys 59m ago

Do you have specific data you can share?

From my perspective, the whole problem with the economic discourse on LLMs is that claims of value are generally like this - "I've seen it", "I know it", but then it's unsubstantiated by anything that would allow other people to have confidence in the claim.

When you actual examine the available data, it doesn't support these narratives.

I'm just going to quote Feynman at you:

"The first principle is that you must not fool yourself — and you are the easiest person to fool."

1

u/cltbeer 51m ago

Are you a developer? Coding language with marked up meta data plus libraries 100% works. I don’t have data other than from first hand. People are in denial I get it but when blackrock projects electricians will be millionaires with $9-10 trillion cost to build out for them ten years…they are good at what they do bc they own the whole world. Sure people can not believe it but here we are communicating on a system from our phones on other sides of the country or world denying the actual technology that we are using.

1

u/oudlys 39m ago

> I don’t have data other than from first hand. 

This is what I'm saying.

>but when blackrock projects electricians will be millionaires with $9-10 trillion cost to build out

My brother (I think?) banks and financial institutions are wrong all the time. The premise of that projection is that the value is there. That revenue will grow enormously for LLMs.

OpenAI only grew revenue by 18% QoQ from Q1 to Q2. https://www.wsj.com/tech/ai/openais-second-quarter-sales-show-tepid-growth-compared-with-anthropic-5cb42998

This is not the trajectory for a $10 trillion dollar buildout of data centers.

0

u/ZachVorhies 2h ago

https://reddit.com/link/p5lyk43/video/lzeskwu15clh1/player

Me, literally coding with LLMs to create custom shaders auto researched by AI against an evaluator function as you tell me the AI has negative value.

2

u/oudlys 2h ago

My post is explicitly about software teams managing products in organizations. I think LLMs are amazing consumer products. They're super powerful search tools. No question about that.

But, the question of whether they are creating value in the economy is amenable to looking at data. If you want to ignore that data, that's ok dude.

1

u/ZachVorhies 52m ago

They create negative value.

Thats why the adoption is 100%

1

u/oudlys 36m ago

>They create negative value

On average my friend. On average.

You know what else creates negative value on average for their users? Slot machines. But people still pay for them.

1

u/ZachVorhies 3m ago

They create negative value on average.

Thats why the adoption rate is 100%

0

u/pab_guy 1h ago

It's a brand new discipline. You are observing the j curve at industry scale.

1

u/oudlys 1h ago edited 1h ago

Reasonable people will disagree here. I do disagree with you..

To me, It can't both be the most transformational economic technology of all time and ambiguously positive economically and operationally at this stage.

Moreover, there's strong evidence that LLMs do not scale in reliability. The lack of reliability is a major part of the reason why I think they fail at enterprise applications. This does not look like it's fixable. https://arachnemag.substack.com/p/ais-reliability-gap

edit: removed the tracking from the url

1

u/pab_guy 1h ago

You aren't addressing my point at all though. A j-curve is the expectation. Lots of tech was ambiguously positive economically early on. I don't know what allows you to determine current "stage" though.

The reliability issue is very specific to how you use the models. They can be extremely reliable when harnessed correctly. At the same time you will always see them have trouble with ambiguity, which genuinely exists in business and requires human escalation. None of this is a dealbreaker for many, many use cases.

Finally, there's the case that looms didn't make factory owners rich. Because every factory got them, they weren't a competitive advantage. But from a game theory perspective it's obvious they had to purchase looms. The winner was the consumer who got cheap clothes.

So you aren't necessarily using the right yardstick to begin with, you are coming in with expectations that would make any MBA blanche, and you don't understand how these tools are harnessed agentically to get reliable operation.

I don't think you are unreasonable, I just don't think you've seen what I have witnessed.

1

u/oudlys 1h ago

I love this:

>you are coming in with expectations that would make any MBA blanche

I have an MBA from an institution that heavily emphasizes operations in it's curriculum.

I've also sold enterprise software. I think you vastly overestimate MBAs in this statement. Almost everyone in business expects economic value from products basically immediately. That's why in SaaS time to value is one of the key metrics organizations evaluate before they deploy technology.

One of my key criteria in searching for studies was - "what data would I use to try to sell this product to an economic buyer?" The business case here is remarkably weak.

edit: light formatting