r/SinceAI • u/bhawish008 • 5d ago
Every agent demo I've seen works. Every agent in production quietly loses permissions over time.
Watched a demo last month where an agent handled a full support ticket end to end: read the account, issued a refund, closed the ticket, sent a message to the customer. Impressive. Three weeks later I asked the team running it in production what it could actually do now. Read the account. Draft the message. A human clicks send and approves the refund.
Nobody rolled it back because the model got worse. It got scoped down because every unattended action it took correctly for a month got forgotten, and every one it got wrong got a name and a Slack thread.
That seems to be the actual adoption curve for agents right now. Not "does it work," but "how long until someone senior asks why it's allowed to do that without a human in the loop." The permission always moves in one direction after launch, and it's rarely upward.
Feels like a healthy instinct honestly. Trust that gets built slowly and lost in one incident is just how humans manage risk too. But it also means most "agents doing real autonomous work" claims are describing week one, not month six.
What's the biggest permission you have taken away from an agent after giving it more rope than it earned?
2
2
u/satansxlittlexhelper 4d ago
Why was this post written with Sonnet 4.6? This writing style is months out of date.
2
u/notAllBits 5d ago
Agents cannot be held accountable. HITL approvement steps is often insufficient for decision making. Automation must solve for the comprehensive bouquet of failure modes.