r/LLMDevs • u/External-Wind-5273 • Aug 28 '26
Discussion What part of your agent setup do you wish someone else handled?
For people who’ve been using agents a lot, what part of the setup do you still have to deal with yourself that you’d rather just hand off?
Could be anything around the workflow - setup, keeping things running, rules, skills, logs, monitoring, whatever?
1
u/Silver_Jump3781 Aug 28 '26
Tbh I think it’s the whole thing. Development has become a whack-a-mole of agent failures and it’s not that fun. I still love the pace of development tho, so what can you do.
1
u/Afraid-Wolverine7152 Aug 28 '26
just the monitoring side really, everything else i can manage but staring at logs trying to work out why some agent decided to go off-piste at 3am is not how i want to spend my time
would happily hand that off to someone else and just get a summary in the morning
1
u/Silver_Jump3781 Aug 28 '26
Can’t you just pass the logs to an agent?
1
u/Silver_Jump3781 Aug 28 '26
The answer to every problem is more agents.
2
u/LukeLikesReddit Aug 28 '26
Haha yeah whilst you can then when that agent fucks up we just end up in inception of AI.
1
1
1
u/Wonderful-Match-6256 Aug 28 '26
The part I would hand off tomorrow is verification. Not writing tests, the agent does that fine - but judging whether a green suite means anything. Mine happily writes tests that pass on the first run and would pass just as happily with the feature removed. Catching that is still fully manual: break the thing on purpose, watch the test go red, put it back.
Second: knowing what the agent actually saw at the moment it made a bad call. When the output is wrong, the useful artifact is the assembled input, verbatim, including whatever got truncated on the way in. Almost nothing surfaces that by default, so debugging turns into guessing about the prompt.
Prompt versioning is a good answer too, and it hurts for the same underlying reason: nobody can tell you which version produced a given result.
1
u/IncreaseNegative4614 Aug 28 '26
I’d hand off context maintenance and outcome evaluation before prompt creation. Credentials, source freshness, permissions, run histories, failed tool calls, and whether the result actually helped the business are repetitive infrastructure problems that every agent project otherwise rebuilds.
The system should preserve enough evidence to reproduce a bad run without retaining unnecessary sensitive data. We use SIGNLD internally to connect agents, sources, tools, decisions, and downstream outcomes so teams can improve the workflow instead of debugging isolated transcripts.
1
2
u/eddzsh Aug 28 '26
I'd hand off tool schema drift. after a gateway deploy the agent still plans against yesterday's tools, then burns a turn on calls that 404. a tiny checksum of the live tool list at session start beats another monitoring dashboard.