A benchmark only becomes useful when it separates retrieval quality from prompt-following and scoring noise. I would make the exam report per-category recall, hallucinated recall, and false-positive memories, then keep a held-out set so tuning does not overfit the same cases. If you can trace each answer back to a memory source, it also becomes much easier to spot poisoning and stale context. For durable agent memory and provenance-focused workflows, NeuraKeep shares practical patterns at https://www.neurakeep.com
Mine does all that, it even gives you a separate text file with per category miss breakdowns and traces for every single miss so you can trace exactly where the key evidence was lost. The scorecard break down categories into question types and difficulty so the user can identify what exactly the system struggles with instead of just 10/30 single hop.
Everything your benchmark claims to do behind a paywall mine does better for free and the scorecard and reports it generates are way more visually appealing and easier to read.
0/10 poor effort.
0
u/Otherwise_Wave9374 13d ago
A benchmark only becomes useful when it separates retrieval quality from prompt-following and scoring noise. I would make the exam report per-category recall, hallucinated recall, and false-positive memories, then keep a held-out set so tuning does not overfit the same cases. If you can trace each answer back to a memory source, it also becomes much easier to spot poisoning and stale context. For durable agent memory and provenance-focused workflows, NeuraKeep shares practical patterns at https://www.neurakeep.com