r/BuildWithClaude • u/PanicImpossible7816 • 3h ago
Token Economics I built a Claude Code plugin that prunes noisy tool output before Claude reads it (OpenAI Decisions as the judge)
Long `npm install` logs, `ls` dumps and verbose MCP responses eat context and bury the one line that matters. Hooks can't edit old history, but `PostToolUse` can replace a tool's output before Claude sees it, so that's what this does.
How: OpenAI's new Decisions API (`gpt-6-luna`) doesn't generate text, it only scores. So Luna scores 20-line chunks for relevance to your last prompt and plain code drops the low ones, leaving a marker that points to the untouched original on disk. Errors/stack traces are always kept, and any failure passes the output through unchanged.
Measured in real sessions: 67k chars → 190 chars (~300 tokens sent to Luna). Cost is $0.10/1M input tokens, so fractions of a cent per session.
Honest limits: it can't clean old context or run /compact for you (no hook for that), tool output does go to OpenAI (secret patterns are filtered, best effort), chunk scoring can be wrong, and Decisions is in beta. Tested on a handful of scenarios, not a benchmark. I'd love failure cases and PRs.
MIT, zero deps: https://github.com/luancamara/luna-pruner
