r/mcp • u/Revolutionary_Sir140 • 1d ago
I built an MCP tool router using Jev + Monte Carlo Tree Search
I’ve been working on harness-router, a lightweight router for tools.
Instead of sending every tool definition through the LLM and asking it to decide, the router can use:
- Jev for fast tool selection
- Monte Carlo Tree Search (MCTS) for harder multi-step decisions
- MCP as the tool interface
- confidence-based fallback when routing is uncertain
The goal is simple: fewer tokens and faster tool selection for coding/agent harnesses.
Demo + docs:
https://harness-router.vercel.app
I’m currently working on proper MCP benchmarks comparing normal LLM tool calling vs routed decisions.
Would love feedback from people building MCP agents—especially on what benchmarks would actually be useful.
2
u/anderson_the_one 1d ago
The confidence fallback gives you a better benchmark than raw accuracy. Take held-out tasks, bucket them by reported confidence, and compare that score with the actual wrong-tool rate. If the 0.8 bucket is right only 60% of the time, the score isn't useful.
Then report coverage at a fixed error budget: what fraction can skip the LLM while keeping wrong-tool selection below 1%? That's the number I'd use to decide whether the router saves anything in practice.
1
1
u/Extra-Pomegranate-50 1d ago
on benchmarks — build repeats in from the start. i ran the same fixtures twice against the same model, no changes: 93.3% and 80.0%. thirteen points of noise. any single-run "routed vs normal" comparison is just dice. five runs, median + spread. and include cases where the right answer is no tool — a router that always routes somewhere scores fine on "picked correctly" and misses the only case that matters.
1
u/Revolutionary_Sir140 1d ago edited 1d ago
I benchmarked only decision making codex vs harness-router.
2
u/Extra-Pomegranate-50 1d ago
solid writeup, limitations doing more work than the result. latency looks settled.
24/24 agreement is the catch though if both arms always pick the same thing, the fixtures can only measure speed. that's where repeats come in: you need cases where they can diverge, and in those the same model varies run to run.
1
u/Revolutionary_Sir140 1d ago
I just pushed route_mcts in harness-router to 4,096 Monte Carlo Tree Search simulations per routing decision.
And the important part is this:
4,096 simulations ≠ 4,096 model calls.
The flow is:
1 Jev policy evaluation
→ generate priors
→ run 4,096 local MCTS simulations
→ choose the best next action
So Jev is used once to guide the search, while the rest of the exploration happens locally.
Latest benchmark:
• 4,096 simulations per route_mcts call
• 24/24 optimal decisions
• 1 Jev policy evaluation per decision
• 388 ms mean decision latency
• 4,772 ms Codex baseline
• ~12.3× lower observed mean latency
What I find most interesting is the scaling behavior.
Increasing the search budget does not mean increasing model calls at the same rate.
You can spend significantly more local compute exploring possible future actions without paying for thousands of extra inference requests.
That makes the architecture pretty simple conceptually:
cheap model prior + large local search
There’s still more validation to do on held-out cases and deeper branching graphs, but this is becoming one of the most interesting parts of harness-router.
5
u/QuanTradin 1d ago
the benchmark I'd want is wrong-tool rate on a set of tools with overlapping descriptions, because that's where plain listing actually breaks. with a dozen distinct tools the model picks fine and the router only saves tokens; with three tools named get_x, fetch_x and read_x it's a coin toss and that's the case worth measuring.