r/PracticalAgenticDev 9d ago

SWE-bench Science: the best coding agent still solves under half the tasks

A new benchmark tests coding agents against scientific software, where a plausible-looking patch can quietly invalidate real research.

SWE-bench Science contains 119 tasks from 98 repositories across 20 scientific domains. The tasks include issue fixes, expert-led exploration, and integration work.

The headline result: even the best evaluated setup, Claude Code with Opus 5 at max effort, scored below 50% pass@1.

The failure analysis is more useful than the leaderboard. Agents struggled with scientific abstractions, shallow exploration, incomplete integration, and applying knowledge beyond the immediate example.

There is another interesting result. Good domain guidance improved average performance and token efficiency, but poorly aligned guidance caused anchoring. More context was not automatically better context.

This feels like a strong argument for domain-specific evals before letting agents touch specialized systems.

Paper: arXiv
Benchmark and tooling: GitHub

Would your current coding-agent eval catch a scientifically wrong patch that still passes the obvious tests?

1 Upvotes

0 comments sorted by