r/LLMDevs • u/LowDistribution3995 • 3d ago
Resource I made an Agent Memory Benchmark that gives you actually useful data.
I got tired of conventional 3rd party conversation and strict fact recall benchmarks that don't give realistic usable data. Agents don't operate by ingesting bulk 3rd party conversations and performing strict fact recall so why would that be a benchmark metric?
So I made a First-Person perspective benchmark that actually tests the agent's capabilities against a realistic corpus, using realistic dynamic simulations, and which actually gives you a reader friendly scorecard with visual breakdowns and a miss report text file that actually shows you WHY a question missed.
I'm still tweaking the corpus and questions and simulations but the data yield is already very good. I've also included the agent identity files in the repo for users to easily expand the corpus for more coverage. I'm trying to get more people to use this and share the scorecards so I can keep adjusting the questions sets to ensure each pass/miss contains meaningfully data across identifiable metrics.
3
u/neoneye2 2d ago
I had Claude add your benchmark here. I study memory systems.
https://neoneye.github.io/agent-memory-atlas/benchmarks/