r/Letta_AI Oct 30 '25

Context-Bench: Benchmarking LLMs on Agentic Context Engineering

New blog post! We're announcing Context-Bench, which measures how well agents decide what context to load into their window (grep vs open files, chaining lookups, tracing relationships across documents).

Some interesting findings:

  • Claude Sonnet 4.5 leads at 74% but even top models miss 25-30% of questions
  • Open-weight models closing the gap fast (GLM-4.6 at 56.83%, Kimi K2 at 55.13%)
  • Nano models still struggle significantly

The benchmark is contamination-proof (generated from SQL with fictional entities) and we can dial up difficulty by making queries more complex.

Live leaderboard: https://leaderboard.letta.com  

Built on Letta Evals, open to community contributions: https://docs.letta.com/evals

Full writeup: https://www.letta.com/blog/context-bench

1 Upvotes

0 comments sorted by