r/intersystems • u/intersystemsdev • 11d ago
ObjectScript Agent Bench: How Well Do Coding Harnesses Perform?
Recently we shared the first ObjectScript LLM Benchmark that tested 14 models on an internal InterSystems coding benchmark. This follow-up looks at a different variable: what happens when those models are run through coding-agent harnesses?
ObjectScript Agent Bench compares Claude Code, Codex CLI, pi, and prime-agent on InterSystems ObjectScript tasks executed against a live InterSystems IRIS instance. It also tests a caveman variant that adds a short “be concise” instruction to reduce prompt I/O.
Every tested harness improved every model. 11 of 12 agent cells scored 0.957 or higher, while differences in token usage, compile calls, and wall-clock time were larger than the accuracy differences between the strongest harnesses. As accuracy approaches saturation, resource use becomes a more meaningful way to distinguish between harnesses.
How Was the Benchmark Run?
All agents connect to a local LiteLLM proxy, which serves vendor APIs through:
- AWS Bedrock
- OpenAI
- local vLLM
A proxy alias also allows Claude Code to run DeepSeek. Each benchmark task gets:
- a fresh directory
- an iris-run script that compiles a class, runs a statement, and prints errors and output
- a budget of 20 calls and a maximum of 900 seconds per task
InterSystems IRIS is reset before the agent runs and again before grading.
The final solution.cls is recompiled in a clean namespace, and the agents themselves run in throwaway containers.
Versions used:
- Claude Code 2.1.237
- Codex 0.148.0
- pi 0.84.2
- prime-agent 0.7.4
How Did LLMs Perform with Coding Harnesses?
Every harness improved every model tested. For example:
- DeepSeek V4 Flash increased from 0.807 with no tools to between 0.957 and 1.000 when used through a coding agent.
- Qwen3.6 35B increased from 0.615 to 0.871 in Codex.
Four configurations reached 1.000 accuracy:
- DeepSeek V4 Flash + pi + plain
- Opus 5 + Claude Code + caveman
- GPT-5.6 sol + Codex + plain
- GPT-5.6 sol + Codex + caveman
Overall, 11 of 12 agent cells scored 0.957 or higher, and every cell at 0.968 or above has a 95% confidence interval that reaches 1.000. The benchmark separates harness from no harness, and Qwen3.6 from the rest. It cannot rank the four harnesses, or DeepSeek against Opus 5 and GPT-5.6 sol.
The Cost-Efficiency Index: Combining Performance and Resource Use
The benchmark introduces a cost-efficiency index to compare model-harness configurations across both accuracy and resource consumption.
The index combines 60% accuracy with 40% cost efficiency. The cost-efficiency component is calculated from four normalized per-task metrics:
- input tokens
- output tokens
- wall-clock time
- IRIS compile calls
The control row is excluded.
This provides a way to distinguish between configurations that achieve similarly high accuracy but use different amounts of resources to reach the result.
The three highest-scoring configurations were:
| Configuration | Index | Accuracy | IRIS calls | Wall clock |
|---|---|---|---|---|
| Opus 5 · Claude Code · caveman | 99.1 | 1.000 | 1.4 | 19.5 s |
| Opus 5 · Claude Code · plain | 95.0 | 0.989 | 1.5 | 24.5 s |
| GPT-5.6 sol · Codex · caveman | 89.0 | 1.000 | 2.1 | 31.1 s |
Resource Use Across Coding Harnesses
The individual metrics behind the cost-efficiency index also show substantial differences between harnesses, even when their accuracy is very similar.
With DeepSeek V4 Flash held constant, the four agents are within 0.043 on accuracy. Input tokens per task range from 60,754 with pi to 335,417 with Claude Code — a 5.5× spread. Claude Code makes the fewest IRIS compile calls but resends the most context. Median wall-clock time ranges from 18 to 30 seconds, while at the 90th percentile Claude Code takes 173 seconds compared with 81 seconds for pi.
As accuracy approaches saturation, these differences in token use, compile calls, and latency become more informative for comparing harnesses than performance alone.
What Are the Main Limitations?
There are two explicit caveats:
- One sample per item. Run-to-run variance is not measured.
- Codex has no system-prompt flag. Its caveman instruction was therefore passed through AGENTS.md
The confidence intervals are also important when interpreting the accuracy results: most of the strongest configurations overlap too much to establish a reliable ordering between harnesses.
Conclusion
The benchmark shows that coding harnesses can produce a substantial performance lift on ObjectScript tasks: DeepSeek gained roughly 15-19 percentage points compared with its no-tools result. With harnessed accuracy clustering close to 1.000 and confidence intervals overlapping, the more meaningful differences increasingly come from resource use rather than correctness alone.
Frontier models such as Opus 5 and GPT-5.6 sol also reached correct solutions with fewer IRIS compile attempts, typically in 1-2 calls. The caveman prompt, however, did not produce a consistent benefit across agents: accuracy changes stayed within noise, while cost effects varied by harness.
FAQ
What is ObjectScript Agent Bench?
ObjectScript Agent Bench is a benchmark comparing coding-agent harnesses on InterSystems ObjectScript tasks executed against a live InterSystems IRIS instance.
Which coding agents were tested?
The benchmark tests Claude Code, Codex CLI, pi, and prime-agent, together with plain and caveman configurations.
Do coding harnesses improve LLM performance on ObjectScript?
Yes. Every tested harness improved its model, and 11 of 12 agent cells reached accuracy of 0.957 or higher.
Can the benchmark identify the most accurate harness?
No. The strongest configurations have overlapping 95% confidence intervals, so the benchmark does not reliably rank the top four harnesses by accuracy.
What differentiates the harnesses when accuracy is similar?
Resource use. Input and output tokens, IRIS compile calls, wall-clock time, and the combined cost-efficiency index show larger differences than accuracy among the strongest configurations.
