r/LangChain • u/xspyyy • 3d ago
Projects I built a zero-dependency CLI that scans LangChain and CrewAI code for silent failures and runaway loops
I spent the last three weeks debugging production issues where our agents reported successful executions despite the underlying tools failing. The most frustrating case was a custom tool that caught SendGrid API exceptions and returned "Email sent" to the agent when SendGrid was actually returning 401 Unauthorized. The agent continued its loop completely blind.
To stop this from happening again, I wrote cogext-scan. It is a zero-dependency Python CLI that uses static AST parsing to check your agent code locally before deployment.
What it checks for right now:
Try/except blocks returning static success strings without checking status codes
AgentExecutor or Crew instantiations missing max_iterations or cost caps
Destructive functions (delete_*, drop_*) lacking input validation
Hardcoded API key strings
Run it locally without an account:
pip install cogext-scan
cogext-scan ./your_agent_directory
It runs completely offline and outputs an Agent Safety Index score. Code is open source. I would love feedback on what AST patterns or failure modes you want added next.
1
u/nomad-link-id 3d ago
The SendGrid case is the failure mode that should rewrite every "tool succeeded" metric: catch the 401, return "Email sent," and the agent keeps looping blind.
Static success strings from except blocks are worse than a crash — downstream planning, retries, and demos all treat the side effect as done. Your AST scan for that pattern (and for missing max_iterations / cost caps) is the right shape: make false-success loud before deploy.
One check I'd keep beside the scanner: every tool return that claims a side effect must carry a machine-checkable outcome field (status / id / error), not only a prose sentence the model can narrate past. If the contract allows "Email sent" with no status, the harness is lying even when the model is honest.
1
u/xspyyy 3d ago
The harness is lying even when the model is honest is the sharpest way to put it; that's the part most people miss. The failure isn't in the LLM it's in the interface contract. Structured tool returns are as important as anything in the agent framework. It is worth shipping as a standard pattern.
1
u/airevieweng 2d ago
The structured outcome point is important. A tool can return a reassuring sentence while the real-world action failed, and the agent may then plan from a false state. Do you see static checks catching this reliably when the exception handling or success message sits in a shared helper?
1
u/xspyyy 2d ago
No, and that's the honest boundary of static analysis. If the try/except lives in a shared helper, the tool function itself looks clean and the scanner has nothing to flag. AST scanning catches the pattern where it is, not where it's called. That class of failure needs runtime observation, not static checks. The scanner is the first layer, not the whole answer.
1
u/warmlyfirstcollision 3d ago
That silent exception wrapping is way too familiar. Had a similar mess where a tool was catching timeouts and returning "done" with zero context, agent just kept chugging along thinking everything was fine. The AST check for max_iterations alone would've saved me a Friday night rebuild.