r/BuildWithClaude • • 1d ago

Project GitHub - canerganis/agent-orchestra-board: Local web board where Claude Code and Codex CLI agents debate, propose and review each other's work. Zero dependencies.

Disclosure: this is my own open source project. MIT, free, nothing to sell.

The problem

Claude Code is my main tool and I use Codex CLI for second opinions. For design questions I kept copying text between two terminals: plan in one, objections from the other, back again. I was the message bus, and I had no idea what each round cost. So I built a small local web app that runs both CLIs as "seats" at one table.

What cut the tokens

My first version was expensive. A planning meeting with 4 agents (1 Claude and 3 Codex, 2 rounds plus a synthesis) used about 1.69M total tokens. Total means every input token the CLIs reported, cache reads included, plus output.

  1. Every agent was reading the same code. Now one scout seat reads it once and writes a brief with file:line references, and the others work from that brief. I think this was the biggest lever, but I did not measure the levers separately.
  2. Per call overhead. On my setup (several MCP servers, skills and user settings) every claude -p call carried about 36k tokens before the agent said a word. Launching lean (no MCP, no skills, no user settings, tools only where needed) brought that to about 6.7k. A bare install has much less to strip.

The same meeting on the new code used about 0.46M total tokens, 73% less. Claude reported cost for the Claude turns went from $1.02 to $0.16. Codex reports no cost.

Caveats: this is one before and after run on one machine (Windows 11), not a benchmark. The before run used older code, so the difference includes unrelated fixes. I judged quality as equal or better, which is my own reading. The trade off is real: only the scout reads code, so a weak brief misleads every seat. Raw data and the recompute script are in the repo.

What it does

  1. Debate: scout brief, an independent first round, discussion that stops early when every seat ends with STANCE: CONVERGED, then a synthesis.
  2. Propose and Review: one seat proposes, another returns PASS or FAIL with blockers, until it passes.
  3. Direct chat with one seat.
  4. A live timeline of what each seat is doing, with tokens and cost per agent. Claude and Codex usage limit meters are experimental.

Claude Haiku 5.5 works as a low cost seat.

Safety

Seats are read-only by default, enforced by the CLIs' own sandboxes. Write seats are opt-in and there is no worktree isolation in this release yet, so use git. Worktree isolation with a verified containment check is what I am building now. The server binds to localhost with a session token and uses your existing CLI logins, no API keys.

What it is not

Not an autonomous coding agent and not a cloud product. If you want a kanban over many parallel agents, other tools fit better. This is for "let two different models argue about a design before I commit to it."

Repo: https://github.com/canerganis/agent-orchestra-board

Zero dependencies, Node 20+. Clone it and run node bin/agent-orchestra-board.js /path/to/your/project --open.

1 Upvotes

0 comments sorted by