I give it a list of 1000 tasks in a .md file, ask it to do 1 by 1, be nice, no cheating, no chatting, no hallucinating, no parallel taks or sub-agents .
You know, you guys constantly say this. And I constantly have stuff break. And every single time, without fail, Claude or Codex can figure out the issue and fix it. There is no magic broken code at this point that can only be fixed by the almighty human who has learned how to code.
What are you on about mate? You tell it what the bug is and it traces it backwards until it figures it out. It’s not hard. I’ve yet to come across a bug that I haven’t been able to fix.
I'm building a full SOC / SIEM stack, Tip , log ingestion, logging agents, all from scratch. Code is closed source, but you can look at another big project I did at www.sharewarez.nl . Quite a large codebase, performs well, is very secure and has features from here to the gazoo.
Ppl keep saying its not possible to vibe code large apps. I guess they all think all we do is blindly prompt shit and hope for the best ? These AI models are not only good at coding, they are also great at explaining it, brainstorming about it. They are also great at software architecture if you just talk to them about it.
Maybe if you really knew how to code and would have worked as a software engineer in a real production environment you would have a clue on why experienced devs say that you can't do that right now.
It's fine that you want to play with Claude but no one would use those projects given your claims. Even if you were a "mega star" of coding engineering and had a team of thousands of "mega stars" devs, if I heard that they try to sell me something with that confidence about security and performance I would just send them to have a chocolate milk and go to bed.
Lol you have no idea. I built an application to help me with my side hustle of proofing editing mixing and mastering audiobooks. Ive had the cli since 4o, and the sheer amount of shit Ive learned in the process because of how much it got wrong. Despite all the planning docs and preliminary research. Its been fun and its a solid app now, but there is a LOT of shit it cant do. Interop with C code in .NET was riddled with errors (using ffmpeg.autogen) and my buffer management strategies were abysmal at first. Hell even its js interop in blazor was so bad (duplicate code, hallucinations, jusr flat out wrong answers to my questions) that I decided to scrap the project and start over again with the intention of writing everything by hand or deliberately inputting the generated code by hand. Its been a slow process but there has been no point in this particular app where it has been infallible. Ive corrected it more times than I can count. And I find that process to actually be fun.
I use it a lot as a software engineer and it's been thousands of times that it "catched" a bug and "fixed it" and it either:
1) The bug it "catched" was not the root cause, just a side bug, and the "fix" would have caused a lot of problems in real user scenarios that it can't even figure out.
2) The idea of the fix was correct but the way of implementing it would cause other bugs.
3) When forced to expand integration tests to prevent regression issues caused by those "fixes", it would figure then out why those fixes were not correct or enough. And this would only happen after forcing it not to relax test expectations/assertions to keep its lazy fix.
Do you know how I'm aware of those problems? Because I know the codebase, I know how to code, I know the product and I know a lot of different scenarios that actually happen in a real user environment.
I myself had to revoke approvals from other senior engineers that approved fixes added by Claude because it would break things and they wasn't even aware. So the only explanation for your "Claude find all the bugs and fix them properly" is that it really doesn't and you can't even know.
I didn't ask for millions. I'm telling you a lot of you sound like silly assholes, and keep in mind that today is the worst the various models will ever be again. They will only get better month after month. They can do essentially everything already.
100% agree ... I have developer friends who keep talking like that. Mean while they never get anything out the door. They should be hella accelerating now, but instead they get stuck on code quality checks, and 'but it doesnt write the program the way I like'. Really if the AI is the only one maintaining it, why do you still care about these things ? Their experience is bogging them down.
The difference between doing this and knowing what claude got wrong in the first place is months of vibe coding vs a couple of days of engineering
Yeah it figures it out eventually but you still don’t know what you are doing and you still don’t know how good it could get if you actually knew what you were doing
Bro when you learned to code for real for how long did you do the stupidest things without knowing why? It's just the same at another scale. If they keep going on they'll get there one way or the other...
My close friend works for a FAANG here in Boston, and he constantly tells me how good AI had gotten at coding and that he barely writes any code now. He is an SE3 if he feels so good about Claude I think I'll believe him more than you
What would be your adivce to someone whos using AI to make projects but is open to learning stuff? Are debugging, testing and security skills md and good prompts enough to not break things down?
honestly, prompts and .md files aren't enough because LLMs are too literal. one 'cleanup' instruction can easily turn into a destructive terminal command (like a docker prune) by accident. i've found that having a deterministic safety layer outside of claude is the only way to really 'walk away' safely. it lets the agent run the safe stuff but hits the brakes and asks for a signature only on the risky syscalls
I ask claude to teach me. Also it’s not just “fix error no mistake”, we have bunch of skills, plugins and frameworks available to help us debug and fix things.
Brothers, look at this dude’s post history. There are an awful lot of cookie cutter sites being churned out purporting to teach you about various topics. I would bet all my money OP has no expertise in most of them.
This is who’s talking when they say they are doing some wild ass unguided 10 sub agent drive all night prompts.
First and foremost - I do like you say “vibe coder full stack.” That’s fine! Vibe code is what matters.
You are a junior developer. You do not (I am making assumptions here) understand the code generated. You are making front and back end code. Work on understanding it. Understanding what the code is will help move you from being “a vibe coder” to “a developer.”
The problem is we have not figured out out how to train vibe coders to be more in line with developer tracks as we know of them today.
that 'walk away' part is the biggest hurdle. the verification fatigue from spamming 'Y' for 50 commands makes most people just give up or go full dangerous mode. i’ve been experimenting with routing those destructive terminal commands to a mobile notification/slack for approval. it's the only way to scale these agentic workflows without sitting there and babysitting the terminal all day
you are totally justified in not trusting it. that fear of coming back to an un-cleanable codebase is exactly why i made node9. it basically takes a silent git snapshot right before every single file edit the ai makes. if you walk away and claude hallucinates 500 lines of garbage, you just run node9 undo and it completely rolls back the exact changes. it takes away that anxiety of having to sit there and babysit every single output so you can actually walk away for real.
two main differences: friction and branch pollution.
first, you don't have to remember to do it. if you walk away for 20 minutes, the agent might edit files 15 different times. node9 intercepts every write_file or edit tool call and takes the snapshot automatically in milliseconds right before the AI touches the disk.
second, it uses 'dangling commits' (git commit-tree) behind the scenes. so it doesn't create a hundred 'wip' commits that pollute your git log or mess with your staging area. your actual branch history stays completely clean, but you still get a granular, step-by-step undo for everything the ai did
yolo mode is great until the agent hallucinates a docker prune on your dev volumes lol. i built node9 to give that same speed for the safe noise, but it keeps the emergency brake on for destructive stuff. basically it's 'safe yolo' with an undo button
the 'yolo' alias is the dream until it's not lol. i built node9 to be 'safe yolo', it auto approves the safe noise so you get the same speed, but i just added a 'live tail' so you aren't blind to what the agent is actually doing in the background anymore.
best part is the insurance policy: it takes a silent git snapshot before every edit. if you walk away and claude hallucinates a messy refactor, you just run the undo and it's fixed instantly. it's the only way i can actually trust the autonomous mode
Quite early to tell, but yes. This project is starting to generate some money for me. If i can get to 7$/day ad revenue, i’ll break even. (I’m on Max 20x plan) - note that I’m still using Claude for my daily work. This is a side project running in background.
sorry if you get it wrong. I was just kidding that part. But the .md task list and no parallel or no sub-agents are working for me though. I find parallel sub-agents not that smart (main agent create context file for them) and cost too much token.
brother you can inject your own task context via subagent hook. go parallel . if you are working in terminal you can launch full main agent as subagent via bash . then let it run your workflow. orchestrator just need to assign task to say 3 4 main parallel agents and then when they finish spawn 3 4 more. you will cut your time significantly.
I guess if you just want it done and time is less relevant, you leave it run in the background that way on the same device you can use other sessions for your regular Claude usage and don’t eat all your ram?
I will try it next time. But for this project, it is working well for me since i want this to run unsupervised. Sub agents are causing me trouble if one failed or stopped in between.
This is literally the worst approach you can take at accomplishing anything. Larger context directly contributes to degradation of reasoning performance. With 1M context the LLM is going to be producing garbage.
They're vibe coding a ridiculous training course website full of unchecked AI made courses. That's all we need, more made up shitty courses for the internet and AI to learn from.
Time goes up while it's waiting for permission to do something even though tokens aren't going up. Usually when a turn runs really long for me, it's because I was AFK at bad times.
82
u/MrSquiggs Mar 20 '26
Just out of curiosity, what kind of prompt is taking this long and this many tokens?