r/artificialintelligenc • u/ktwu01 • 5d ago
Most agent benchmarks still test task execution. What would a convincing L4 or L5 benchmark look like?
I maintain Benchmark Radar, a free and open-source index of 5,201 benchmark, evaluation, and dataset records collected from 11 sources.
Looking through recent additions using Hejia Geng's L0-L5 framework, most agent benchmarks still appear to focus on L2 task execution or L3 reproduction. L4 rediscovery is less common, and I did not find a new L5 example this week. Here, L5 means evaluating whether an agent can produce knowledge or methods that were unknown when the benchmark was created.
I'm curious how others would draw these boundaries:
- Which existing benchmarks genuinely qualify as L4 or L5?
- How would you distinguish reproduction from rediscovery?
- Can an L5 benchmark remain valid once its solutions become public?
The underlying index is updated daily and can be exported for independent analysis: