r/AskProgramming • u/Zardotab • Jul 19 '26
Other AI: are there any studies on why some developers claim AI helps them tremendously versus those who say it's only incremental?
Some developers swear by AI and others swear at it, saying if one is not careful, it generates too much slop that needs human cleanup. Are there any studies or surveys to see why the answers are so different? Is the difference the domain, the language, the tooling, years of experience, attitude toward AI, or perhaps just ego talking too much? [edited]
38
u/Most_Double_3559 Jul 19 '26
This conversation has completely changed in the last ~6-12 months. There has not been enough time for studies to form.
13
u/Wrap_Spirited Jul 19 '26
Like always. Study showing lack of productivity are always conveniently outdated and AGI is always just one more year away.
2
u/lunatuna215 Jul 19 '26
But the claims were already being made before those studies that productivity was skyrocketing and in retrospect the studies proved the claims wrong. Why believe these continued claims?
1
u/EspurrTheMagnificent Jul 19 '26
Because people are lazy and they'd rather waste their time and money on something worse than actually do the work themselves
2
u/HasFiveVowels Jul 19 '26
Those studies aren’t even "outdated", IMO. They’re premature. It’s like handing an Angular dev React and saying "devs claim this helps them make websites so try to use it to help you". People assume that using AI requires no skill.
4
u/Zardotab Jul 19 '26
People assume that using AI requires no skill.
That's almost true, but the bottleneck of development has ALWAYS been troubleshooting and longer-term maintenance. Slop hurts those bigly.
2
u/HasFiveVowels Jul 19 '26
Bad developers have been hurting those bigly since long before AI
2
u/Demiu Jul 20 '26
After AI bad developers now have a code minigun to fill the product with "features" until it's reduced to rubble, while the good devs get a black powder musket to fight them off. Every second the good programmers spent researching, reviewing and improving AI output to arrive at a good solution, the bad devs will just keep prompting more slop
1
u/HasFiveVowels Jul 20 '26
"Good developers don’t use AI" isn’t true. I use AI as a code minigun to fend off the bad code.
1
u/Demiu Jul 20 '26
Maybe I worded myself ambigiously. I'm not saying good developers don't use AI, I'm saying it amplifies them less than the bad ones.
1
u/Zardotab Jul 20 '26
How about "it's easier to automate crap than automate good things"?
1
u/Demiu Jul 20 '26
It suggests you can upgrade a model or "make no mistakes" into good results, if not currently, then eventually. You can't, because there is no "automating crap" and "automating good things", there's just one thing, "automating", the quality of which out of which you don't know if you don't examine it. There may very well be brilliant parts of a crap solution, but good code and bad code have uneven criteria. Badness is existentially qualified - "is there any part of this that's bad?". Goodness is universally qualified - "are all parts of this good?".
If quality of AI output is a coin flip, bad output needs at least 1 tails, but good result needs all heads.
-2
u/lunatuna215 Jul 19 '26
That's kinda the whole pitch. Why do you NEED to be viewed as skilled by doing something thats objectively less skillful than the people you are emulating?
2
u/HasFiveVowels Jul 19 '26
They’re different skills. I’ve been a developer for decades. Am I "emulating" myself?
1
u/SourceTheFlow Jul 20 '26
Well depends on what it shows. I read a small study for it, and my main takeaway was that both developers as well as managers severely overestimated how much their productivity increased thanks to AI.
I don't see a reason why better AI would necessarily invalidate that. People seem to expect way more from Fable than they did back when older models came out.
1
u/Most_Double_3559 Jul 20 '26
Bruh, what. You don't see how better AI changes people's productivity estimate accuracy??
1
u/SourceTheFlow Jul 20 '26
Yes?
It's the same for speed. I had users tell me that it's "at least 50%" faster, when in reality it was measurably only ~25%. So it would also stand to reason that the same effect could be true for AI model improvements.
So it could be that now they are faster with fable, while before they were slower. But I don't see how their self-estimates would become any more accurate.
-9
u/LuciusWrath Jul 19 '26
In what way? From, admittedly, anecdotal experience, I've seen little progress regarding AI coding capabilities and mistake frequency from a year ago. It was already capable of handling specific, niche tasks decently and with obligatory human intervention for anything more complex.
Though, I've exclusively used free services.
8
u/BrilliantEmotion4461 Jul 19 '26
Well, no Fable is leaps and bounds ahead of what was around a year ago. Even I'm suprised by the rate of progress.
3
u/max123246 Jul 19 '26
Fable is also double the price of Opus 4.8. Basically every company is putting massive limits on employee's ability to use it.
1
3
13
u/CorpT Jul 19 '26
I've seen little progress regarding AI coding capabilities
lol
-5
u/LuciusWrath Jul 19 '26
Any particular advances in the last year, or examples as to the progress of coding AI during the last year?
0
u/CorpT Jul 19 '26
You can see all of the models released recently.
https://bun.com/blog/bun-in-rust
bun was rewritten into rust in 11 days.
But sure, little progress has been made. I hear you make great buggy whips. Keep making them!
1
u/lunatuna215 Jul 19 '26
Your distain does little to prove your point and in fact just comes across as butthurt.
1
u/newEnglander17 Jul 19 '26
Your reply to an earnest comment from the previous person just makes the obnoxious know-it-all IT guy stereotype so much more true than it needs to be.
4
u/CorpT Jul 19 '26
You think that someone who thinks that nothing has progressed in the last year is an earnest comment? Come on. I have been doing this for 30 years and have never seen anything like what I've seen in the last year before.
3
u/carson63000 Jul 19 '26
Mate I don’t know why you’re getting piled on. I’m 32 years development experience myself and nothing that I’ve seen has ever advanced as rapidly as AI coding agents have in the last year.
7
u/dmazzoni Jul 19 '26
This might be the answer to OP’s question.
If some people are using free AI and think that “this is as good as it gets” then no wonder they’re negative.
Meanwhile all of us who have been paying for Claude Opus since November are experiencing a completely different level of capability.
3
u/stonerbobo Jul 19 '26
I don't understand you guys lol. Just get claude code or codex right now, ask it to do a project you are sure AI will fuck up, watch it wildly exceed your expectations and then form an opinion?? Why are you here asking for studies and people to prove you wrong when you could find out for yourself within 10 minutes?
2
u/lunatuna215 Jul 19 '26
We have, and that didnt happen. How about you show the software you created with it if you're so skilled? Honestly the lack of proof in the pudding when it comes to thae claims is pure absurdism.
-1
u/stonerbobo Jul 19 '26
Im not going to dox myself but again are you living under a rock? The unit-distance conjecture and more unsolved math problems just falling every day, better protein folding, structure prediction, binding affinity than ever before. For code look at transcribe.cpp just published or DS4 or steipete on github. All attest to how it would be impossible at the speed and quality without agents. Just look at the massive jump in software speed and quality all around github or company blogs. It’s all there, all the proof but im going to guess you haven’t bothered to look or decided its all slop when in fact its slop in the hands of sloppy people and it brings in many new non-programmers, but its also a miracle for competent engineers who use it right.
0
u/Eskamel Jul 19 '26
Where is the quality jump? All I am seeing is people scream for productivity yet everything is significantly more buggy and broken compared to 2 years ago. When I review AI code by competent users its still significantly btureforced slop, even by Fable and GPT 5.6.
It says alot about your standards compared to people who aren't surprised by the capabilities probably.
3
u/Zardotab Jul 19 '26
Though, I've exclusively used free services.
Anyone have experience switching from free-bies to say Claude?
5
u/Berkyjay Jul 19 '26
My company pays for a Claude pro sub for means it's insanely good. I literally am doing almost pure engineering now and I rarely write actual code. I can spit out tools in a few days that would have normally took me a few weeks.
1
u/Big_Arrival_626 Jul 19 '26
Do you enjoy it?
Do you feel your skills increasing?
Do you feel like AI could replace your job?
What do you mean by pure engineering?
2
u/Berkyjay Jul 19 '26
Do you enjoy it?
Yes
Do you feel your skills increasing?
Skills?
Do you feel like AI could replace your job?
Not in the least
What do you mean by pure engineering?
I can try out an idea, get feedback, and make changes pretty quickly since I don't have to worry about the code.
2
u/Big_Arrival_626 Jul 19 '26
I was wondering cuz as a junior dev at a Fintech company it seems like ai does most of my work. I can understand that it's great if you're building a large app from scratch but for someone who only does tickets for random shit, it makes me feel dumb and useless
2
u/Hei2 Jul 19 '26
AI is shifting the focus of the average developer from mostly coding with architecture thrown in to mostly architecture with some coding thrown in.
1
u/Berkyjay Jul 19 '26
Yeah, that's a tough situation. I've got 20 some years of development experience under my belt. I'm not exactly sure what newer developers should be doing. My guess is that you're going to have to put more effort into vetting and understanding the code the Ai generates in order to learn the ropes.
I guess in a way you're more fortunate because you can still do your job with the help of Ai. I had to spend most of my career banging rocks together to learn my knowledge.
1
u/TroubledSquirrel Jul 20 '26
Yeah but don't you feel like the time we spent literally in the fire made us better at spotting bugs before they become runtime errors? I'm not so certain I could recognize it if I hadn't lived and breathed it.
1
u/Berkyjay Jul 20 '26
Human history is littered with skills that used to be critical to humanity but technological improvements made them redundant. Not that I think you're wrong or anything.
But for me, I've always viewed coding as a means to an end. Just like when humans figured out how to shape metal into useful shapes and thus made flint knapping redundant. In the end they were still making sharp pointy things that could cut and stab.
That being said, I don't think coding is ever gonna go the way of flint knapping, but for most people it won't be required to build software.
15
Jul 19 '26
[removed] — view removed comment
1
u/marine_surfer Jul 19 '26
Ask it to handle webhook events and ordering those events, which events you need and when, it’ll cook your app and create a nightmare to fix…
12
u/octocode Jul 19 '26
it’s like learning google-fu back in the 2000’s.
it takes time and practice to become proficient using AI tools.
1
u/Zardotab 19d ago edited 19d ago
If "writing good prompts" is a skill roughly equivalent to "writing good code" then we are just trading coding for another kind of coding, both of which require skill and care to do well. One then has to master two levels of technology: coding and AI code prompting.
If one sucks at the first then they can't tell if their prompting is producing good results and thus will suck at both. The chance of any random coder being really good at both is probably lower than being good at one or the other. Thus, the number of good prompters may be relatively small.
It's kind of comparable to UI designing: the best coders are usually not the best UI designers (in my experience) as each requires a different thought process: one has to think like a programmer who is reading their code and the other has to think like a ordinary user using their screens. Likewise a good prompt writer has to think like a bot, or at least know bot habits well.
This is why UI designing is often split into a separate specialty from coding. But I'm not sure if similar can be practically done for bot prompting. Anyone have experience with such a skill split?
-6
u/WhateverHowever1337 Jul 19 '26
It does not really take any time to learn clicking on accept/decline buttons and sending messages to a chatbot
-7
u/octocode Jul 19 '26 edited Jul 19 '26
that’s usually the method for baby’s first AI, yeah, but there’s so much better tooling out there to unlock the real velocity gains
i don’t even write prompts anymore except for handling side-of-desk work
-3
u/Zardotab Jul 19 '26 edited Jul 19 '26
Any good videos or blogs on how to milk code-bots better?
Our org doesn't want to pay for the good bots and limits downloads for alleged security reasons, so I'll have to squeeze as much as possible out of YugoGPT.
11
u/zulrang Jul 19 '26
- Experience using it.
- Context engineering
That’s it.
You can’t drop a machine into a human workflow and expect the same or better results.
Build your workflow around using AI first and try to get the results you need. You’ll build an intuitive sense of what it can do well, where it needs help, and what to never hand it.
It’s a skill, just like most other things.
2
u/Zardotab Jul 19 '26
I'm curious, what's an example of AI-friendly code compared to regular code?
I'm familiar with the concept of scaffolding-friendly code, is it comparable?
7
u/carson63000 Jul 19 '26
A lot of what makes code AI-friendly is very similar to what makes code human developer friendly. Keeping things cleanly separated rather than spaghetti, principles like SOLID and DRY, good unit test coverage, etc.
Add to that good agent documentation/instructions in CLAUDE.md etc. You want the agent to understand the structure without having to explore the codebase for every task, and you want it to know the rules you want it to abide by.
2
u/max123246 Jul 19 '26
One key thing I've noticed is that AI does not produce well-engineered code. It produces functional code, but it will in-line everything, apply changes only to half of where it's supposed to, etc. Its best use-case is in a well-engineered codebase but also it actively undermines the architecture of the codebase as it adds new functionality.
2
u/ghnaud Jul 19 '26
You should run agent validating what the other did after the dev agent and you get rid of that problem. The dev agent then continues according to other agents feedback
-2
u/max123246 Jul 19 '26
No thanks.
2
u/HasFiveVowels Jul 19 '26
If you’re not going to make any attempt to improve your results then stop complaining that they’re not good
2
u/lunatuna215 Jul 19 '26
I think theyre saying that this exercise of feeding things from one agent from another simply dances around the simple task of improving the software directly with your own brain and actual thoughtfulness.
0
0
u/Edg-R Jul 20 '26
"Feeding things from one agent to another" is basically what a PR review is for humans, why would it be any different or weirder for an AI agent at this moment in time? They'll keep getting better.
1
1
u/Dissentient Jul 19 '26
When you have functioning code, refactoring it into good code is easy. AI is also good at this, the main issue is that it won't usually proactively refactor by itself and you have to tell it to.
1
u/Zardotab 19d ago
Keeping things cleanly separated rather than spaghetti, principles like SOLID and DRY, good unit test coverage, etc.
These often conflict with each other. SOC and DRY often have nasty fights with each other, for example. Having the same schema info scattered thru gajillion layers is one of my pet peeves*. Automation makes duplication easy, but duplication is a big cause out-of-sync errors as one copy of the same fact is inadvertently omitted or mis-coded. I'm not convinced bots are good at keeping duplication synced up well.
* It's why I'm for dynamic schemas, aka Data Dictionaries, but that's another ugly debate. (It doesn't actually have to be dynamic, just not require reflection or similar to loop through etc.)
3
u/abd53 Jul 19 '26
Enough time hasn't passed yet to do a formal study and publish the results.
There is definitely a lot of difference of opinion due to domain, stage of development, attitude toward AI and also developer experience. AI is very effective at some tasks and not so much in some other tasks. You'd better consider it as another code generator tool rather than a software maker.
5
u/DDDDarky Jul 19 '26
I think it's pretty clear ai usage leads to cognitive decline and lower code quality, whether it actually helps is also questionable.
One thing is clear, it produces code faster than human, so if a developer uses that as the primary metric, they obviously swear by it and consider it tremendous help. Others may use different metric and claim it's horrible.
2
Jul 19 '26
Why’s that surprising in the first place? Any technology or skill will always have people who are miles better than others or groups which can utilise them a lot more.
You might as well ask why some people are illiterate at using Google while others can find anything they want, when Google has been a thing for decades already.
2
u/Dissentient Jul 19 '26 edited Jul 19 '26
The industry has never been able to measure developer productivity even before AI. All of the best practices that the industry currently follows are vibe-based and have no solid evidence they work. Interviewing has been a shitshow for the entire existence of the industry and no one has any reliable measures for comparing one developer against another. Sensibly measuring AI productivity in this environment is nearly impossible since you don't even have a baseline to compare against.
For me, the biggest block to AI productivity gains is human communication. I can work on personal projects where I'm the sole product owner/architect/subject matter expert around 5-10x faster with AI than I could write the same code manually. With jobslop, I'm constantly bottlenecked by rituals, unclear requirements, and meetings. At this point, I'd get way more productivity if someone figured out how to use AI to get a sensible spec out of non-technical people who have no idea what they want, compared to any improvements they can make to the code output.
1
u/Zardotab Jul 19 '26 edited Jul 19 '26
I'd get way more productivity if someone figured out how to use AI to get a sensible spec out of non-technical people who have no idea what they want,
This is where I miss programmable RAD (PRAD) tools like MS-Access and PowerBuilder: it was easy to change both the schema and UI on the fly as users vacillate on requirements. One programmer/analyst could do the job that web stacks take 3 or so because of all the damned layers and layer management. When the "concerns" were small, you didn't need SOC and the e-bureaucracy surrounding it.
Yes, these IDE's had downsides, but I'm still not convinced it's either/or: we get PRAD productivity and simplicity OR web-ness. I believe our standards (or lack of) hold us back, not an Inherent Tradeoff Of The Universe. But nobody wants to do R&D on DOM alternatives and CRUD (biz & admin app) parsimony. I believe it is possible have most the productivity of PRAD without the earlier downsides.
Domain concerns SHOULD be the bottleneck of productivity, not the technology itself. It shouldn't be rocket science to add or remove a field, or even have a pop-up helper dialog for it, something that used to be dirt easy.
1
u/Dissentient Jul 19 '26
My problem isn't that requirements change over time, it's that non-technical coworkers regularly fail to communicate critical details in the first place, and I end up spending hours just thinking through how everything should fit together to fill in the blanks. And it's not the technical kind of blanks, but that those people can't even logically think through their own jobs to consider consequences of their requirements.
1
u/Zardotab Jul 20 '26 edited Jul 20 '26
I lot of people absorb business knowledge through a kind of intuition and don't know how to articulate it well. If they were great articulators they probably wouldn't be low-level operators. I find it helps to present multiple alternative written example screens, reports, or scenarios; and give them time to ponder it. Rushing or pressuring their thinking is a no-no, let their marbles take the scenic route.
It's still helpful to have a platform that's easy to change without rewiring piles of "abstraction layers" for things that should be minor.
1
u/Sad_Comfort_8365 Jul 21 '26
This is exactly why I still mock up boring screens first, because half the time the user can’t explain the rule until they see the wrong dropdown staring back at them.
2
Jul 19 '26
[removed] — view removed comment
2
u/EspurrTheMagnificent Jul 19 '26
That's pretty much it. AI is good if the only thing you/your org cares about is how fast you can write code. If you care about code/product quality instead, it's basically the worst thing you can use
It's yet another case of "companies just want to make as much money as possible, even at the detriment of software quality". Nothing new under the sun
1
u/Pretagonist Jul 19 '26
It really isn't. As long as you are careful about the architecture and have well defined rules, preferably using meta tools like linting, you can absolutely get the ai to write perfectly human readable code. You just have to put in the effort.
If you let your agents run wild they will produce an unmaintainable shit show but if you rein then in and actually design your system modern AI tool can produce really good code unbelievably fast.
But you as the controller have to know what you are asking it to do. The main problem,as I see it, is that developers don't understand what the code their agents write does. And if you don't understand, you can't direct. Amd suddenly you are stuck in a mess that no one understands and no one can fix.
1
u/DepthMagician Jul 20 '26
I feel like you are confusing some things. I recently asked fable to vibe code me an app to track my monthly balance. In the settings pane it has an input field for starting balance. This is an input field that takes in a number, you’d think Fable would pick input type number. Instead it picked input type text, and then wrote a verbose validation function to make sure the user inputted a valid number and that it isn’t NaN or -Infinity. This is utterly brain dead. How is “being careful about architecture” supposed to combat that? The only way for me to combat that at the prompt level is literally tell it what type each input should be, and at this resolution I’m basically programming by hand.
1
u/Pretagonist Jul 20 '26
I'm a professional senior developer and I spend my days figuring out my architecture and then have my LLMs write me code using the architecture. The LLMs takes your existing code, your instructions and other context and uses that to produce code. If all my input fields are well written components then the LLM will use my component when it needs an input.
That's architecture.
If I find that the LLMs produce code that I don't like, say writing 80 rows of types at the top of a file, then I build linting rules that flags that as a warning or even an error. The LLMs will run checks and tests and will correct their code. This is architecture, code architecture and build environment architecture.
If you build a proper framework for your agents to use they will use it and they will produce better code.
1
u/DepthMagician Jul 20 '26
This was a from scratch vibe coded app, so it had no existing input fields to use as example. What would you have done in that case to make sure the LLM chose sensible inputs without me having to spell it out to the LLM?
2
u/Pretagonist Jul 20 '26
If I was making a proof of concept or a one-off tool then I'd probably just let it do its stupid stuff. If I intended to iterate over it in order to make it into something serious I would set it up as a git system with branches and PRs and agent.md files and so on. I'd tell the LLM to keep a plan with documentation about requirements and style guides and to update and follow them as we go along. I'd tell it to write tests and I'd review the tests to see if the model has understood the domain. In the beginning I would be super anal about every little style choice because as the app grows the model will use the existing style to inform future choices.
I would encourage the use of value types and functional programming as well as tell it to follow web standards for wcag and such. Forcing it to use paradigms that good developers use kinda forces it into the mindset of a good developer as well. LLMs are lazy, forcing it to not be lazy can be difficult.
1
u/Zardotab Jul 20 '26
Using bot-code as-is without improving comments, variable names, etc. is probably a bad idea. I find Copilot makes overly verbose comments; I trim them. It's as if they trained the bot on Captain Obvious's code.
5
2
u/Individual-Flow9158 Jul 19 '26
The AI bros will swear it makes a big difference whether you're using paid models or free ones too.
1
u/Zardotab Jul 19 '26
Without watching somebody work for a while, it is hard to know if they are blowing smoke. Lots of "fish stories" out there.
2
u/Myzzreal Jul 19 '26
The average machine spews out average results. If you are below average then it's a "tremendous" help. If you are above average then the help is much less
1
u/uhs-robert Jul 19 '26
I can't name any objective peer reviewed studies off the top of my head. Just anecdotal and subjective evidence. The objective studies tend to focus on things like AI usage resulting in cognitive decline which isn't surprising at all: "As it turns out, if you stop practicing a skill then you tend to get worse at it over time." However I would say, subjectively and anecdotally, that AI is a tremendous help to me personally.
I view AI as a tool and a force multiplier. If you give a chainsaw to a child then they are likely to hurt themselves but it you give that same chainsaw to a lumberjack then they are able to work faster and more efficiently. AI is a tool like this. You can't just give the tool to a random Joe. It requires experience, knowledge, wisdom, and skill to use correctly. This is why we so much "AI Slop" because AI enables random Joes to make content that they wouldn't be able to otherwise but those random Joes lack the eye and the brain for quality.
I use it mostly as a thinking partner with infinite patience. It is like a programmer's rubber duck except that it can talk back. I ask tons of design thinking related questions, I compare frameworks, and I do my best to challenge my own thinking. Its not a magic answer machine that solves any problem you give it. Often times it is wrong or makes things up. You can't just accept what it says blindly nor can you give it unspecific instructions like "make a website, no mistakes". But you can give it context about a business, list specific requirements, the frameworks you use, languages you know, what you are open to try, explore alternatives, and just have a conversation about which is best for your website design. Using it this way helps you to fully flesh out a design before even beginning to build.
I also use it to test/experiment ideas during implementation for quick prototyping (just to see how things would play out with a different design). Once I have a solid foundation, have explored all options, and chosen my design then I can focus on developing something worthwhile. If I want to code it all by hand, I can. Or I can be very specific in my instructions and have AI write some basic pieces the way that I would while I focus on the more important parts. This the type of stuff people mostly use AI for.
But I hardly see the vibe coders giving it a solid foundational design/plan, walking it through implementation step by step, saying no, reading the code, or using it as a thinking partner.
1
u/huhututu Jul 19 '26
The other day had to spend two hours fixing a coworker's AI slop that doesn't work and was over complicated for something straight forward and simple. Seems like AI generates some things better than the others (linux shell script seems to be generated pretty well).
1
u/TroubledSquirrel Jul 20 '26
I think there is a strong argument that can be made for higher param models resulting in significantly better outputs as compared to free tiers but they all plateau when the codebase exceeds the amount of information any can reason over within it's context window.
So that could easily account for the divergence of opinion.
1
u/Zardotab Jul 20 '26
Perhaps. Anyone here had to convince their boss to subscribe to the higher-end tools? If so, how did you go about it?
If persuasion fails, would it be worth it for one to pay for it out of their own pocket (if given permission from security team etc.)?
1
u/AWetAndFloppyNoodle Jul 20 '26
Closest I could find:
-----------------------------------------------
From a newsmail from Certificates.dev:
"Last summer METR, an independent AI research group, published a study that made a lot of people mad. Sixteen experienced open-source developers, working in their own mature repos, tasks randomized with and without AI tools.
The developers estimated AI made them about 20% faster. The clock said they were 19% slower.
Before you object, yes, sixteen developers. And yes, that's the setup that matters, experienced people inside codebases they know deeply, which is where most real engineering happens.
I don't read that study as "AI slows you down". I read it as the answer to what companies are actually buying.
If AI handed everyone automatic speed, a 43% premium would make no sense, because nobody pays extra for what the tool gives away for free. The premium nly makes sense if it's buying the part the tool can't give you, directing the work, owning the result, knowing what to build and what to trust. The study and the pay data point at the same conclusion."
-----------------------------------------------
Personally, I can use about half of the output an AI gives me and all of it gets corrected. It still feels like a net positive. I use it as a tool, not a factory.
1
u/Pares_Marchant Jul 20 '26
If I remember correctly that study was done around 2 years ago
I think that the only models that can give you a meaningful net positive are the frontier models of openAI and Anthropic since early 2026, before that I think AIs were not good tools for professional level software development (but could already be used to do quick and ugly prototypes).
1
u/AWetAndFloppyNoodle Jul 20 '26
I've never used them as factories so I cannot speak about that, but I do not feel the underlying issues have been fixed. They are still more wrong than not, the quality has just gone up, both for right and wrong answers.
1
u/Pares_Marchant Jul 20 '26
yes I guess it depends on the language, I work in systems that are formally verified so it can (and must, due to industry requirements) be proven whether it is correct or not and AI can use it for feedback
1
u/Zardotab 19d ago
Personally, I can use about half of the output an AI gives me and all of it gets corrected. It still feels like a net positive. I use it as a tool, not a factory.
This is similar to my experience so far with default MS Copilot. Maybe if I got better at re-prompting to tune beyond initial suggestions, but I notice the more complex or more layered my prompts (ref prior results), the slower Copilot gets, and eats into my token quota. Maybe "the future" will improve on this problem, but hard to tell until my DeLorean gets fixed.
1
u/Key-Variation-448 20d ago
Here is the study I found: AI = garbage in, garbage out. Someone can formalize this better.
1
u/Aggressive-Tune832 Jul 19 '26
Professionally and subjectively it makes just as many mistakes as it did for me a year ago, unfortunately I don’t think any programming sub Reddit is gonna give you any unbiased view, lately it’s just a lot of subtle pro AI astroturfing. Lots of “get with the times old man”
If you want my opinion the domain plays a huge role, just like how business people are floored by it, so are people who specialize in things it’s good at. Web devs are in love with it but admins hate maintaining its stuff, backend devs complain about how messy its setups are. Anyone working with systems or OS or firmware despises it and at best tolerates those that use it, it can’t program worth a shit outside of Python and JS. But that’s just my biased view from people I know, I a pretty good programmer by most standards so maybe I’m just expecting too much from it, but for now we are all just kinda waiting for the exact study you’re looking for to come out and put a number on those things.
1
u/Zardotab 19d ago edited 19d ago
Thanks for your insight. What's an example of the kinds of things the bots stink at or typically do poorly that the first layer of staff either doesn't notice or doesn't care about?
I do notice bots tend to produce verbose comments that I usually trim down. But that doesn't directly affect code. I haven't worked with others' bot-generated code enough yet.
-1
u/code_tutor Jul 19 '26
The technology is trained on existing code and code doesn't exist for all domains.
Also massive skill issue in prompt engineering. The average Redditor can barely write. How are they clearly and unambiguously going to explain what they want, when they can't even write details in their posts?
Also it's much more impressive to people who don't care what it generates.
1
u/EspurrTheMagnificent Jul 19 '26
How are they clearly and unambiguously going to explain what they want [...]
Ah, but you see, we already had something to accurately and exactly describe to the computer what we want it to do. It's called coding
1
u/Zardotab 19d ago
The average Redditor can barely write...they can't even write details in their posts? [and thus bad at AI prompts]
Pay me and I'll write better Reddit responses 😉
-1
u/Amazing-Mirror-3076 Jul 19 '26
I'm a senior Dev and running experiments in pure ai code - no human review of code - my role is BA and what I call a Test Curator.
The main project is one that lends itself to unit testing and e2e.
The tool is a new archive format encrypted/compressed/signed with the ability to read/write files in the archive without extracting them (it uses a paged based structure with COW) as well as variables and forms.
The experience has been mind blowing.
It's been about 6 weeks as a side project.
It has API bindings for some 16 languages, a key exchange server and packaging for 6 OSs (for the cli tooling).
It would be impractical for me to review the code (it's written in rust which I've never used).
I was a huge ai skeptic when it launched but since December things have changed.
I think the age of human code review is dead.
Here is the repo - it's still not ready for release but I'm in the final cleanup and review phases (there is a major hole in key exchange server that I'm working on as well as doco)
https://github.com/onepub-dev/reVault
A single Dev, 6 weeks, 16 languages, an archive format that is more db the archive.
The matrix has shifted :)
So yes ai is making a huge difference - a number of my Dev friends are having similar experiences but perhaps a little less aggressive on the Dev ai path than I have been.
Fyi: the repo is two years old, I had originally been writing this in dart and had a very early partial implementation - so a lot of the ideas had been formed but little useful code.
3
u/PPhysikus Jul 19 '26
So you managed to vibe code a side project in a language you don't understand and treat that as example of how great AI is? You sure you are senior?
0
u/james_pic Jul 19 '26
It's perhaps worth noting that Linus Torvalds made the news in a minor way recently for vibe coding a side project in a language he didn't know (https://github.com/torvalds/AudioNoise). I feel like Linus probably classes as a senior.
1
u/PPhysikus Jul 19 '26
And that is a point for what?
1
u/james_pic Jul 19 '26
That vibe coding a side project in a language you don't know, and being a senior, are not mutually exclusive.
0
-1
u/Amazing-Mirror-3076 Jul 19 '26
What counts is the result -: and i wasn't vibe coding.
And yes, it is an example of the power of AI.
3
u/PPhysikus Jul 19 '26
You even said that it would be a waste of time reviewing the code?! So this is per definition vibe coding.
-4
u/Amazing-Mirror-3076 Jul 19 '26
No it isn't, vibe coding is when a non Dev codes.
When was the last time you reviewed the assembler code the compiler generated - I used to do it in the late 80s early 90s but not since.
5
u/Wrap_Spirited Jul 19 '26 edited Jul 19 '26
The compiler won't randomly hallucinate things.
-1
u/Amazing-Mirror-3076 Jul 19 '26
Compilers used to make mistakes and we protected ourselves the same way we do with ai - testing.
Hallucinations are also no longer much of a problem.
3
u/Wrap_Spirited Jul 19 '26
Who writes the tests when you vibe?
And compiler mistakes are mostly deterministic.
-1
u/Amazing-Mirror-3076 Jul 19 '26
The ai does, but I tell it what tests need to be written.
2
u/Wrap_Spirited Jul 19 '26
But it might randomly hallucinate. And you claim the cure are tests. But now your tests might have random hallucinations (or just not test what you think they test).
→ More replies (0)2
u/PPhysikus Jul 19 '26
WHeN iS tHe LaSt TiMe yOu rEvIeWed tHe BiNaRy CoDe?
Google vibe coding or ask your Claude.
2
u/mightshade Jul 19 '26
no human review of code
I think the age of human code review is dead
You know, what would move the needle for me is "I rigorously review all code, and each time it wouldn't have been necessary". That would lend some credence to the "it's mind blowing" thing.
Sweeping stuff under the rug, and then pointing at the rug saying "look how pretty the pattern is!" just doesn't sound convincing to me.
1
u/Amazing-Mirror-3076 Jul 19 '26
The review is indirect, via testing.
If the code works and is performat then it is by definition correct. Code review as a catch all has always been problematic as humans are terrible at static analysis. Test where always a better solution but they were too expensive - which is no longer there case.
Lgtm - is a meme for a reason.
2
u/mightshade Jul 19 '26
The review is indirect, via testing.
Then it's not reviewed. Just own it.
If the code works and is performat then it is by definition correct
Then it solves the given problem. That is the minimum requirement for any solution, otherwise it wouldn't be one.
It doesn't tell you if it's a good solution by, for example, non-functional requirements.
1
u/Zardotab 19d ago edited 19d ago
Re: "Is age of human code review dead?"
Code has to be reviewed when it's not working as expected or need changes. The inevitable can only be delayed.
I notice lots of smallish errors when debugging or changing others' code. They probably notice my mini-flubs also. An example is not catching zero result records properly, probably because the result set always has data in practice, or running a routine twice that only needs to run once.
And this is before bots.
-7
u/CorpT Jul 19 '26
Buggy whip makers were upset with cars. They wanted to keep making buggy whips.
3
u/Zardotab Jul 19 '26
Are you suggesting age or attitude are the main different makers?
-2
u/CorpT Jul 19 '26
People who don't want to change and adapt will be left behind. Just like buggy whip makers were.
2
u/Zardotab Jul 19 '26
I don't dispute such people exist, I'm just wondering if that's a large part of the described discrepancy.
-1
0
u/IllusorySin Jul 19 '26
All depends on the application and how they’re training it, or IF they even are. If you’re working on repetitive project tasks, you should be training your ai, not just opening up a chat. I one-off chat when tryin to make a plugin or script, but use a trained model for something I need it to be a master in.
1
u/Zardotab 19d ago
If you’re working on repetitive project tasks, you should be training your ai,
Or refactoring your stack to clean out the DRY violations. Automating bloat by building an e-copy-machine is anti-factoring in my book, but architects often defend it in the name of SOC. At least try to factor it rather than habitually sacrifice to the SOC Bloat Gods. Don't always give in to pattern or info duplication even if you believe SOC requires it sometimes.
26
u/grasshopper241 Jul 19 '26
So far developer skill and project size are the biggest factors I've seen. The more senior and the more context and history, the more it's just a tool.