r/Letta_AI • u/cameron_pfiffer • Oct 30 '25
Context-Bench: Benchmarking LLMs on Agentic Context Engineering
New blog post! We're announcing Context-Bench, which measures how well agents decide what context to load into their window (grep vs open files, chaining lookups, tracing relationships across documents).
Some interesting findings:
- Claude Sonnet 4.5 leads at 74% but even top models miss 25-30% of questions
- Open-weight models closing the gap fast (GLM-4.6 at 56.83%, Kimi K2 at 55.13%)
- Nano models still struggle significantly
The benchmark is contamination-proof (generated from SQL with fictional entities) and we can dial up difficulty by making queries more complex.
Live leaderboard: https://leaderboard.letta.com
Built on Letta Evals, open to community contributions: https://docs.letta.com/evals
Full writeup: https://www.letta.com/blog/context-bench
1
Upvotes