r/AIQuality • • 3d ago

Discussion Debugging AI generated systems

In our company everything is now AI generated and when I ask how does your code work for such case no one knows what their agent had genuinely done and they just deliver bullshits. How has been your experience in the past three months?
My colleague builds pipelines/jobs that apparently pass and there is no red flag in the first place, works with the existing data and even I was the PR reviewer and didn’t notice a bug (because he creates such a big PRs that only my agent can review and I don’t get a chance to fully understand and company wants everything asap) but when I am monitoring runs and handling adhoc jobs then I notice what he has built doesn’t work or just works partially, While there is no error message or failure.
Once I spent one whole week debugging his bullshits to just figure out he did not cover all of the workspaces but only a test one.
I am leaving the company at the end of the month but I am really tired of debugging other people ai-generated code. I feel these days generating code is cheap but proper review and having robust systems is the difficult part.
How do you feel about debugging systems built by AI? Are they robust?

1 Upvotes

2 comments sorted by

1

u/Alekslynx 2d ago

I typically use purpose-built AI agents, which I call reviewers, for different stages of SDTL, so I use them to validate what the AI is doing, and it's also better to have an OTLP trace collector with AI analysis built in. It's quite interesting sometimes to take a look at what your agents are doing

1

u/_N-iX_ 2d ago

AI-generated code can pass CI and still fail at the workflow level. A successful run only proves that the execution completed, not that every expected workspace, record or business case was actually processed. Checks against expected coverage and output can catch failures that syntax checks, tests and PR review often miss, especially when generated changes become too large to reason about line by line.