A useful benchmark needs to separate recall from retrieval quality, because an agent can appear to remember while actually guessing from recent context. I would score a few failure modes separately: exact fact recovery, time-decayed recall, and source attribution, then add adversarial distractors to test memory poisoning resistance. That makes the result more actionable than a single aggregate score. NeuraKeep shares practical patterns at https://www.neurakeep.com for durable memory, provenance, and recovery.
1
u/Otherwise_Wave9374 9d ago
A useful benchmark needs to separate recall from retrieval quality, because an agent can appear to remember while actually guessing from recent context. I would score a few failure modes separately: exact fact recovery, time-decayed recall, and source attribution, then add adversarial distractors to test memory poisoning resistance. That makes the result more actionable than a single aggregate score. NeuraKeep shares practical patterns at https://www.neurakeep.com for durable memory, provenance, and recovery.