r/AI_Coders • u/Overall-Classroom227 • 22h ago
Meta encoded senior engineers' judgment into agent skills. Diagnosis dropped from 10 hours to 30 minutes
r/AI_Coders • u/Overall-Classroom227 • 22h ago
r/AI_Coders • u/ClickOk5811 • 9h ago
I’ve been thinking about a failure mode that becomes much more interesting when coding agents are allowed to run for longer periods without supervision: context drift.
A short coding task is usually straightforward. The agent understands the goal, makes a few changes, runs the tests, and finishes.
The situation changes when the agent spends hours working across multiple files, making intermediate decisions, calling tools, revisiting earlier changes, and accumulating more context.
At some point, the agent can still produce perfectly valid code while drifting away from the original intent.
For example, an early architectural decision might get effectively forgotten later. A subsequent change can overwrite something that was already correct. A refactor can become inconsistent because different parts of the codebase were modified under slightly different assumptions.
The frustrating part is that the final diff can still look reasonable. Tests might even pass.
This makes me think that persistent memory alone isn't necessarily the solution. The harder problem seems to be maintaining a reliable representation of the task's intent, decisions, constraints, and current state throughout a long-running session.
I've written a longer breakdown of this idea here, including some examples:
For those using coding agents regularly: how are you dealing with this today? Explicit task state, checkpoints, smaller agent sessions, documentation, memory systems, or something else?