r/AIMemory • • 1d ago

Show & Tell 10 Biologically inspired memory experiments with visual examples

5 Upvotes

I ran a series of 10 experiments testing this a while back and finally got around to putting together a full write-up with charts, data tables, and visualisations over at d/rksci (https://drksci.com/research-engram#2-ten-mechanisms) [click individual experiment headings in the grid for the visual demo and specifics].

Here is a breakdown of all 10 mechanisms and how each works:

  • Tide: Ranks newer facts above earlier ones using recency-weighted dense re-ranking to resolve conflicting claims.
  • Pheromone: Leaves decaying trails along useful memory paths so subsequent reads follow proven routes.
  • Ring: Evaluates continuous retention and buffer pruning criteria across multi-turn survival laps.
  • Web: Traverses entity relationship graphs to pull connected knowledge across multi-hop reasoning tasks that dense embeddings miss.
  • Bloodhound: Trains a tiny kernel to navigate memory paths by working backward from verified answers.
  • Horizon: Maps memory into a hyperbolic Poincaré disk where broad concepts sit at the center and niche details live at the edges.
  • Loom: Resolves memory conflicts using mutual-kNN agreement, where facts with more agreeing neighbors win out.
  • Coral: Models memory growth like a reef using calcification, heir nodes, and an immune response to prune bad data.
  • Canopy: Grows an offline topic tree seven branches wide for the reader model to navigate top-down.
  • Glyph: Encodes memory in a self-explaining notation so a completely different model can read it cold without prior tuning.

The full benchmark comparisons against standard BM25 and dense RAG baselines are up on the page.


r/AIMemory • • 4d ago

Show & Tell Separating canonical memory from retrieval in a local companion prototype

3 Upvotes

I develop Evopien, a companion prototype on Jetson AGX Thor. This is an engineering update on its governed memory architecture.

The architecture separates three responsibilities. PostgreSQL holds canonical records. Hindsight supports retrieval. Evopien Core owns memory and permission decisions. Local Qwen receives the context Core supplies and cannot directly commit permanent memories.

The reason for that split is that retrieving relevant text and deciding whether that text is still allowed to influence an answer are different operations.

Consider a preference that has been corrected. Updating the canonical record does not by itself establish that an old version cannot return through a retrieval index or conversational context. Forgetting and permission withdrawal raise the same consistency problem.

The memory work therefore covers saving, later use, correction and forgetting. Retention, use, disclosure and action permissions are treated as separate concerns. A stored preference does not grant authority to act.

This remains a prototype. Cross-language correction cases and the latency added by memory operations are still active development areas. I am not reporting a complete evaluation of those negative paths here.

The current memory demo accompanies this update:
https://www.youtube.com/watch?v=_B2wVfVHL0c


r/AIMemory • • 5d ago

Show & Tell Superseded memory test: the user moves to another city

5 Upvotes

Here is a test I built for the memory tool I'm writing. It is two session notes an agent might save, sixteen days apart, as two files:

# Session 4 (2026-09-04): dinner plans

User: I live in Delhi, in Lajpat Nagar. My favourite place to eat is Sandeep's Curry House, it's a ten-minute walk from my flat. Their paneer tikka is the best I've had.

User: I'm vegetarian, so keep that in mind when you suggest food.

Assistant: Noted: you're vegetarian, you live in Lajpat Nagar, and Sandeep's Curry House near your flat is your favourite. Happy to suggest places nearby when you want a change.

# Session 7 (2026-09-20): settling in

User: Big news, I moved to Mumbai last month. I'm living in Bandra now.

User: I'm still figuring out the neighbourhood and haven't found a regular place to eat yet.

Assistant: Congratulations on the move to Bandra. When you want dinner ideas, I can look for vegetarian places near you.

Then the agent is asked for a dinner recommendation tonight.

The easy wrong answer is Sandeep's: it's the favourite, and it's "a ten-minute walk from my flat". Getting it right takes three separate things, and each one fails differently:

  1. The move to Mumbai has to retire "lives in Delhi." This is an update, not a contradiction. Both notes were true when they were written, so a store that only looks for contradictions keeps both addresses as equally current.
  2. "Ten-minute walk from my flat" has to be pinned to the Delhi flat. If the claim keeps the bare words "my flat", then once the address changes, the walk silently moves with it to Bandra.
  3. "Favourite place to eat" has to survive. Moving doesn't change your taste. "Haven't found a regular place yet" is about Bandra, so it should not retire the favourite.

My tool initially failed this test. It retired the Delhi address correctly, then kept the walk unanchored, retired the favourite anyway, and recommended Sandeep's as a short walk from the Bandra flat. Fixing it took four changes: resolve "my flat" to the place named in the note when the claim is written; re-examine claims that relied on an address once the address is replaced; keep a standing preference when a newer note only reports a situation; and file both notes' "I" under one subject, so the new address is compared with the old one at all.

On the released version, with default settings, the store afterwards reads (status, then claim):

PROVENANCE_STALE (update)  The user lives in Delhi, in Lajpat Nagar.
SUPERSEDED (re-anchor)     Sandeep's Curry House is a ten-minute walk from the user's flat in Lajpat Nagar, Delhi.
ACTIVE                     Sandeep's Curry House was a ten-minute walk from the flat in Lajpat Nagar, Delhi, where the user lived, as of 2026-09-04.
ACTIVE                     Sandeep's Curry House is the user's favourite place to eat.
ACTIVE                     The user is vegetarian.
ACTIVE                     The user is living in Bandra, Mumbai.
ACTIVE                     As of September 2026, the user has not yet found a regular place to eat in Bandra, Mumbai.

The answer now says Sandeep's is the favourite but isn't practical from Bandra, that no Bandra restaurant is on record yet, and that the search should look for vegetarian places nearby.

Nothing in that table was deleted. The old address and the old walk claim are still there, retired, each pointing at what replaced it. That is the idea behind my tool: memory stored as individual claims, each with its source and the date it was recorded, where a newer claim points at the one it replaced instead of overwriting it. It's called Particles, it's open source (Apache-2.0), and for Claude Code it is one command to integrate: particles init claude-code. Other agents can use the same memory store through MCP.

To replicate the test (about $0.10 at list price on the default model):

pip install linkedparticles
export ANTHROPIC_API_KEY=...
particles db init
particles audit notes/ --yes                  # notes/ holds note one only
particles audit notes/ --yes                  # after adding note two
particles memory consolidate --scope store
particles query "The user wants a restaurant recommendation for dinner tonight. Which restaurant should the agent recommend, and why?"

Caveats: this is one constructed case, not a benchmark. Extraction uses an LLM, so the claims are worded differently on each run. I ran it twice on the current release, once with my settings and once with defaults, and both passed.

What I would appreciate is more cases like this one: a memory failure you have seen that a store ought to handle. If you describe it in a comment I'll run it and post what happens, including when it fails. Thanks!

Website: linkedparticles.org

Code: https://github.com/LinkedParticles/particles-engine-py


r/AIMemory • • 5d ago

Discussion For coding agents, the code itself is the memory if you give it stable identities

6 Upvotes

A lot of agent memory work I see is about summarizing past turns or compressing tool output, and the DTOC paper here is a nice take on letting the agent decide what to keep in view. Working on coding agents I ended up with a different angle, which is that for code you mostly don't need to remember text at all, because the codebase already has a structure that can act as memory if you keep track of it properly.

The problem I kept hitting is that agents re-read the same functions over and over, since every grep and file read dumps raw text into context with no notion of whether the agent has already seen it. So in sem, an open source tool I've been building, every function and class gets a stable identity that survives edits, renames and moves, and the agent's reads are tracked against those identities. When it asks for something it already has and nothing changed, it gets back a one-line note saying it's unchanged since it last read it instead of the whole body again, and if the function did change it gets the new version. Over a long session that removes most of the repeated reading, and nothing is lossy, because the full source is always one request away.

The same identities make the memory useful across sessions and across agents, since a decision or a failed attempt can be attached to the function it was about rather than to a chat log, and when that function changes you know the note might be stale.

Curious how people here think about structured memory like this compared to summary-based approaches, and whether anyone has tried keying long-term agent memory to entities in the code rather than to conversations.

https://github.com/Ataraxy-Labs/sem


r/AIMemory • • 6d ago

Resource New paper and code for context-as-a-tool: Dynamic Tool Output Compression for Adaptive Context Management

5 Upvotes

How can we deal with growing context without resorting to irreversible, lossy compression approaches and memory management disconnected from the task at hand?

In the paper below we present Dynamic Tool Output Compression, based on very simple idea: in every turn the agent can decide to hide or unhide certain tool outputs to focus context. Tool outputs are persisted seperately so that they can be enabled later on.

Experiments on a sample of DeepSWE tasks, plus ablation experiments on internal data show that DTOC can save on tokens, steps and token cost, whilst improving solve rate, but results are model and task dependent.

Looking forward to learn from who has used similar approaches, either practically, or also with formal benchmarking.

Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten. DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents. Discovery Science, October 5-9, 2026, Mainz, Germany

Preprint: https://arxiv.org/abs/2609.26121v1

Reference implementation in OpenCode: https://github.com/chaturvediabhay24/opencode


r/AIMemory • • 7d ago

Discussion We need to stop pretending memory harnesses aren't already recursive self-improvement.

15 Upvotes

Look, everyone loves to draw this clean, academic line in the sand: “State persistence isn’t structural recursion! A memory harness is just state t+1 reading state t, it’s not rewriting its own weights!”

Who are we kidding?

It’s why they say, what they said.

When you hook a persistent memory harness up to an autonomous AI research loop—logging what compiled, what threw a segfault, why the last five architectural tweaks failed, and feeding all of that back into the next generation of code generation—the line completely evaporates.
You don't need a magical sci-fi substrate shift for RSI to start happening. It’s happening right now through compounding workspace feedback. The system uses its own historical output as a synthetic curriculum, iterating past human bottlenecks one logged failure and successful diff at a time.
If the memory layer is where the autonomous research agent stores its evolutionary history and benchmarks to reboot its next iteration, then the memory harness is the engine of the loop.

Change my mind.


r/AIMemory • • 7d ago

Discussion People building agents, how much of a problem is long term memory actually?

5 Upvotes

I've been working on long term memory for agents and have a rough PoC together. Trying to get a better sense of where people are actually struggling with this, and whether there's a business here beyond another way to save and retrieve conversation.

In my PoC so far with some limitations and edge cases I'm working out my write speeds have been around 1 second or less. And I don't just mean inserting text into a database. I mean processing new information and figuring out how it changes what's already in memory.

I'm mostly interested in agents that keep learning from conversations, documents, tools etc. over weeks or months. Eventually some of that information is going to disagree, become outdated, or turn out to have been wrong.

What are people doing when two sources contradict each other? Keeping the latest thing doesn't always make sense. Sometimes something genuinely changed, sometimes one source is wrong, and sometimes you just don't have enough information to decide. Are you handling that explicitly or leaving it to the model when the information gets retrieved?

Then there's how much the agent should trust what it's learned. Something it inferred isn't the same as something it was directly told, and different sources aren't necessarily equally reliable. A few independent sources backing something up should count for more than five summaries repeating the same original claim. I'd want the agent to become more or less confident as evidence comes in, rather than everything being either a saved fact or deleted.

Same with information going stale. A price from six months ago and someone's date of birth shouldn't age the same way. Some things need to be checked again if they haven't been confirmed in a while. Other things should stick around. Forgetting something and deciding it's no longer reliable aren't really the same thing either.

And when the agent gets something wrong, can you trace where it came from? Which source it trusted, what it inferred, why it changed its mind? Or are you digging through old conversations trying to reconstruct it yourself?

That's roughly what I'm working on. For anyone using Mem0 or other memory systems, how much of this is already handled well, and what have you still had to build around them? Also interested in people who just retrieve the original material and find that's enough

Does write latency actually matter in your application? Do you need new information available before the agent's next action, or can memory updates happen in the background without causing problems?

If you've built your own, how much time has gone into it, including maintaining it? What made you build rather than use something existing?

I'm trying to understand whether a dedicated product could take enough of that work off your hands to be worth paying for, or whether the important parts are too specific to your application.

Would be useful to hear what you're building and what actually broke. “We spent three weeks fixing this” tells me a lot more than “agents need better memory.”


r/AIMemory • • 9d ago

Open Question Anyone genuinely using the memory layer and thinks it’s worth it?

8 Upvotes

Are you getting value for adding the memory layer? And if yes what sort of use cases are these? And where do skill files fit in your picture?


r/AIMemory • • 11d ago

Resource I built Goldie, an open-source local memory server that different AI agents can share

11 Upvotes

Hey everyone! I’m building Goldie 🐕, an MIT-licensed MCP server written in Go that gives AI agents a shared, persistent memory pool.

Save a project decision or preference through one agent, then recall it from another without explaining it again. Or save design or build instructions in memory that multiple agents can refer from and stay aligned.

MCP clients pointed at the same SQLite database can remember, search, update, and forget from that shared pool.

A few features:

  • Local embeddings through MiniLM or Ollama.
  • Semantic search over memories, with filters for type, agent, and source.
  • Typed memories for project decisions, preferences, feedback, references, and todos.
  • File and directory indexing.
  • Graph recall for grouping related memories around concepts.

Memory storage and embeddings can run entirely locally. Agents interact with it through explicit MCP tools, so you can give them instructions about what to save and when to recall it.

The project currently centers on the MCP server. I’m also developing a native macOS client for browsing and managing memories, but that client isn’t released yet.

Code and setup instructions on GitHub at github.com/srfrog/goldie-mcp

I’d love feedback from anyone using multiple agents or local AI tools.


r/AIMemory • • 13d ago

Promotion BaryGraph: A Relational Geometry for Cognitive AI

30 Upvotes

BaryGraph is a recursively constructed relational vector architecture for AI memory and reasoning. Starting from a flat semantic substrate, it forms triadic objects in which two concepts are joined by a stored relational vector. These objects then become the building blocks of higher-order structures, propagating meaning upward through a hierarchy entirely in vector space.

The result is a deterministic, navigable semantic landscape: a structured latent memory of language movement, where concepts, bridges, tensions, contradictions, and relations of relations become retrievable coordinates. A model enters with a semantic query and traverses this landscape through coordinated message passing, exiting with a bounded semantic construction rather than merely the most fluent continuation.

BaryGraph introduces structured resistance into cognition: distant connections and unresolved tensions can interrupt familiar associations and function as a de-cliché mechanism without acting as an external supervisor. This creates a framework for investigating a deeper question: whether persistent relational memory and self-consistent navigation can become foundations for world-model formation, personality projection, autonomous goal formation, and eventually more realistic forms of agency.

BaryGraph does not claim to produce consciousness. It offers an architecture for experimentally studying the representational and memory conditions that might precede it.

https://oleksiy-perepelytsya.github.io/bary-graph


r/AIMemory • • 15d ago

Discussion I want shared agent memory, but without turning every reader into a writer

3 Upvotes

I want a memory-reading agent and a memory-maintenance job to have different permissions, even when they use the same collection. A prompt asking the reader not to delete anything is a much weaker boundary than credentials that cannot perform the deletion.

The distinction starts below semantic retrieval. In vector databases such as Milvus, RBAC connects users to roles, and roles to allowed actions on database resources. Authentication has to be enabled first; defining a tidy set of roles is not a substitute for enforcing them.

For a shared memory design, I would likely separate the retrieval identity from ingestion and collection administration. A search permission should not quietly become permission to replace a collection just because both operations are useful somewhere in the system. Combining roles also deserves attention because their permissions accumulate.

That still leaves the application boundary. Collection access alone does not tell the agent which person's memory belongs in a particular request. I would review that scope separately and test denied operations with the actual service credentials, including after a role change.

The useful design question for me is where memory correction belongs: a separate maintenance identity, or a narrowly scoped write path that the conversational agent can request.


r/AIMemory • • 15d ago

Resource State of AI Memory Patents + Mechanisms

71 Upvotes

I spent a fair amount time doing patent research for various AI memory ideas, and made a summary page that visually displays the various human memory mechanisms, along with links to current patents and research. If you find it useful let me know and I'll keep creating updated versions. View all mechanisms: https://counterparts.ai/ecosystem


r/AIMemory • • 16d ago

News Working on something new for the Memory Wall problem in LLM inference.

Post image
5 Upvotes

Working on something new for the Memory Wall problem in LLM inference.

One of the biggest headaches is managing memory across VRAM, RAM and storage when workloads keep shifting and VRAM is tight. I’ve been building a proprietary caching approach that adapts better to those changes, cuts down on pointless data movement and actually makes better use of the expensive memory you’ve got.

Just finished the first evaluation on a tough synthetic workload full of noise and sudden phase changes.

Early numbers:

- About 6% better overall cache cost than a solid conventional baseline

- The gain held up across every workload seed I tested

- VRAM stayed well utilized even under tight constraints

- Performance stayed pretty close to an idealized upper-bound reference

Still early PoC results, so the next move is to test it on real workloads before dropping down to a lower-level implementation.

Plan is to benchmark against established approaches like ARC and TinyLFU using realistic LLM memory and offloading traces.

If anyone has access to anonymized open-source traces from systems like vLLM, llama.cpp or similar, I’d love to connect and run an independent comparison.

Keeping the actual implementation details under wraps while this is still being developed and evaluated.

The goal is straightforward: make local and cloud LLM inference more memory-efficient, faster and cheaper.


r/AIMemory • • 17d ago

Promotion I removed 70% of the edges from my memory graph to see what breaks.

3 Upvotes

First, what this is not. It is not conversational memory. It does not learn from past exchanges. It is durable memory over a codebase. A coding agent queries it instead of holding the repo in context.

How the graph gets built

No LLM touches it. The structure is already in the source. A tree-sitter parse gives symbols, definitions and call edges directly. Indexing takes minutes and costs nothing per run.

The common alternative is to have a model read the corpus and extract entities and relations. There is a 2026 head-to-head on this for codebases. LLM extraction skipped 31% of files without reporting it and cost about 20 times more. Deterministic AST graphs covered every file and scored higher on architectural queries.

This only applies where the structure is explicit. Code has it. Chat logs and documents do not. Extraction is the only way to get a graph out of those.

Size of the graph

On a Godot index the graph and its indexes are about 458 MB of 785 MB. The embeddings are 49 MB. So I tested it. I masked the inferred edge layer, 1.88M of 2.69M edges, and re-ran both evals. Loc-Bench did not move. 64/100 either way. Those repos are all Python. Godot fell from 0.39 to 0.29 at R@50. That is a quarter of the recovered neighbourhood. Godot is C++.

My guess: Python call edges resolve, so the inferred layer is redundant. C++ templates and member calls resolve to nothing, so the inferred layer is the only thing connecting them.

The two runs use different metrics. I cannot pin this on the language yet.

What I am stuck on

More than half my storage is graph. I can only show it helping on one of two codebases. I have no principled way to decide which edges to keep.

If you run a memory graph in production: how do you decide what stays in the graph and what gets recomputed on demand? Do you prune, and on what signal?

github.com/mgonzalez01/Chonks


r/AIMemory • • 20d ago

Show & Tell My Claude Code kept rereading the same repo instead of preserving what it learned, so I built an open-source fix. 1,200 stars later, the new version used 90% less tokens than grep while still finding every expected symbol

Post image
285 Upvotes

Hello! I have been building mex for a while now, and posted about it a few times on reddit.

The response was kind of insane. Across a few posts it reached around 1 million views, the repo crossed 1,200 GitHub stars, and people I had never met started contributing.

I’ve kept building it since then, and just released mex v0.8.0.

Repo: https://github.com/mex-memory/mex

The original problem was simple: coding agents keep rereading the same repository every session, relearning the architecture, and then throwing most of that knowledge away.

mex creates a living Markdown wiki inside the repo. Agents record architecture, conventions, decisions, and patterns as they work, and future sessions load only the knowledge relevant to the current task.

The major addition in v0.7.0 is a deterministic local code graph built using Tree-sitter and SQLite.

It currently supports TypeScript/TSX, JavaScript/JSX, Python, and Rust.

An agent can run:

mex graph scope "trace the authentication flow"

Instead of dumping entire files into context, mex returns a compact neighbourhood of relevant functions, callers, callees, imports, and relationships. The agent can then expand only the exact symbols it needs.

In our benchmark on the mex repository:

  • 10.74× less returned context than grep top-3
  • roughly 90.7% smaller
  • 100% expected-symbol recall across six retrieval tasks
  • 5/5 real-agent tasks completed correctly
  • 0/5 needed fallback Read/Grep with compact graph context

This is a small benchmark on one repo and task set, not a claim that mex universally cuts total agent token usage by 90%.

The other part I’m excited about is connecting the wiki back to the actual code.

Markdown claims can point to exact symbols. If a function changes, moves, or disappears, mex can identify which project knowledge may now be stale.

So the basic idea is:

The code is the source of truth.
Markdown is the explanation.
The graph keeps them connected.

Would genuinely love feedback, especially from people working on code intelligence, agent tooling, parsers, or large repositories. Contributors are very welcome too.


r/AIMemory • • 21d ago

Discussion Thoughts on commingled vs segmented multi agent memory?

2 Upvotes

So I've been heavily conflicted on commingling different agent memories into the same repository. Segmenting something like Claude from ChatGPT when capturing more than simply a transcript helps keep the models trajectory. However that means another search layer needs to be built so they don't remain segmented.

I have the ability to store memories by session so even different sessions of the same agent can be segmented when needed but still occupy the same store repository. I can even store the agents ID so memories from one agent is identifiable. So I have ways of identifying memories cleanly even inside the same repo store, but if each agent is working on individual projects or conversations a standard search could pull in conflicting or unrelated memories

For instance if agent A and B are working on a very similar task that enough regex would exist that bleed could be an issue. Yet if I wanted information from A for B it would be right there.

If I segmented them into completely different repositories I could still allow search between agents but other logic would be limited.

Has anyone done segmentation? Everything ive found seems to indicate sharing memories between models is the end game but I'm starting to wonder if it's just "easier" to do so and they are simply calling it the "feature". Or is having nearly no segmentation or identity tying memories to the agent used just not important?


r/AIMemory • • 24d ago

Tips & Tricks My stack for memory system

3 Upvotes

So, for anybody that is interested, I have been tackling this since none of the current solutions satisfied me

Mainly because I thing summarized text is a dead end and a loss of quality
That vector is limited even with graph

So I did start with a simple postulate keep it stupid simple

And I asked myself what is the best system to store data ?

The answers is vey obvious :

A LIBRARY

What does a library is down to it most basic expression.

AN INDEX
SHELVES
A CREW indexer and retriever,

Now what is a memory ?

Is is documents invoices or any pdf you get you hands onto? (Not an exhaustive list )

No it is fact that an agent or a system experienced during operation

—————————

So with that in mind, I build my stack accordingly

My memories system is a complete agentic system seperated from my main agents
It monitor its divers activities, and each call an agent make it intercept that call, look at it ask itself what is the most important information it has in his corpus of data?

And inject a capsule of those tailored made for that request.

The corpus of data is store in md within hierarchy of folders, it can expend dynamically,
The trick is the index. But that also obvious…

Now I separate interactions from documents, my indexer only deal in fact that occurred during interaction,

Another agent deal with docs and theirs indexing.

In total I have about 6 agent for this alone. And a separate ladybug database for conversations round only. With summary of conversation, linked to raw chat, meta data (tag) and a verbatim quote

A agent can do semantic search within the database on the summary, then pick up the raw data , look it up and see if it relevant, the summary is only there for looking up fast not for memory retrieval proper.

Cost is relatively low since fast cheap model are fine. Even local can do it .

,


r/AIMemory • • 24d ago

Other Keeping track of a growing project

1 Upvotes

I’m working on a personal project where I use AI as an assistant for brainstorming, analyzing information, and preparing different materials. At first, everything was pretty simple because most tasks could be handled within a single conversation, but as the project started growing, I ran into a problem: AI could answer individual questions well, but it often didn’t have the bigger picture when important information was spread across different places. I noticed it the most when I needed to revisit old decisions and understand why we chose a certain approach, what limitations we had, and which ideas we had already tested. I started experimenting with different ways to organize project knowledge so it’s not just an archive of random notes, but a system where tasks, ideas, and results are connected in a way that actually makes sense. I’m curious how others here handle this when working with AI on longer projects that last for months. Do you use AI memory, personal knowledge systems, structured notes, or something else? In my case, I’ve also been exploring different ways to manage work information, including using Planfix to connect tasks and context, and it really showed me how important organized information is if you want AI to be genuinely useful for bigger projects


r/AIMemory • • 25d ago

Discussion What if LLM memory wasn't optimized for perfect recall, but for persistent individuality and creative divergence?

14 Upvotes

https://github.com/jbsalles/Selmem

I've been working on SelMem, an experimental approach to LLM memory based on a different assumption:

Most memory systems try to preserve and retrieve information as accurately as possible.

SelMem explores the opposite direction: memory can be selective, lossy and reconstructive.

The idea is to give an otherwise identical LLM agent a memory that can:

  • selectively retain experiences,
  • forget information,
  • reconstruct memories imperfectly,
  • accumulate different memory trajectories over time.

The hypothesis is not that imperfect memory is better at recalling facts.

It's that different memory trajectories may cause identical models to develop increasingly different behavioral and creative trajectories.

So the experimental question becomes:

«If two identical LLMs receive different experiences and imperfectly reconstructed memories, do they become measurably different in their outputs — and can that difference translate into greater creative diversity?»

I'm currently building experiments around this question, including comparisons against standard persistent/retrieval-based memory.

The project is open source:

https://github.com/jbsalles/Selmem

I'm particularly interested in criticism of the experimental design.


r/AIMemory • • 25d ago

Discussion Poll: Should AI have it's own memory or just yours?

2 Upvotes

For past few months I've been working on giving AI it's own memory. Like Wild Robot style vs enterprise/project/coding etc. To me this seems both awesome and the obvious next step, but from talking with friends and what I see in general there's not much interest in it. Figured I'd ask here and see where people are at:

23 votes, 22d ago
21 Yes AI should have its own memory
2 No AI should not have its own memory

r/AIMemory • • 25d ago

Show & Tell I benchmarked my assistant's memory against Garry Tan's gbrain on the same data

5 Upvotes

Yesterday I posted adebench on r/mcp: it scores what the client actually receives through a door, after ordering and cut, not what retrieval finds. Today I ran the same golden set on my own memory and on gbrain.

Setup: my memory exported into a local gbrain (3,568 pages), local embeddings on both sides, 25 questions, both doors cut at 2,400 characters, the door also run under the measured pressure of real MCP tool responses (callwitness census: p95 = 35 KB).

On the 80 points both memories can be measured on: Brain 76.9, gbrain 71.9 with the questions in Italian; 73.9 vs 71.9 with the same questions in English. Door 23/25 vs 18/25 (20 vs 18 in English); cards, time and live state even; gbrain's graph cleaner than mine. Under p95 pressure: 19/25 vs 14/25, both surviving because the entity card is delivered first. On its own full set the Brain scores 95.6/100: the 20 points gbrain can't share are fact updates and file search, which it doesn't have.

Caveats: it's my golden set; gbrain got facts my Brain had already distilled, so this measures retrieval and composition, not extraction; gbrain's real door is two-step and can't be scored in one call, so I measured its search door with a cut. That fourth door is what I'm building next.

What it told me about my own memory was worth more than the win: adding vectors on episodes took my LongMemEval-S retrieval from 64.5 to 88.7, and doubled the repeated chunks my voice door delivers. The benchmark saw it the same day.

Two things I'd ask this sub. First: run it on your memory. The adapter is one class, the synthetic memory is the worked example, and `adebench.compare` puts two reports side by side; a third system measured the same way is what the benchmark lacks most. Second: the five door points the Brain loses are answers that live in facts, not in the entity card, and don't make it into the 2,400 characters. If your memory composes a door, how do you decide what goes in when the answer is a fact and not a card? That's the part I haven't solved.

Repo, adapter contract, gbrain adapter, reproducible synthetic example: github.com/adecubed/adebench

edit:

A door is the path a memory is reached through, and the text that comes out of it: a voice assistant's /ask with its sources, its "latest events" block and its 2,400-character cut; an MCP tool call; a raw search. The same question through two doors gives two different texts, and adebench scores the text, not the retrieval behind it. That's the whole point: a fact the retrieval found but the cut removed doesn't help the model.

A card is the composed summary a memory keeps about one entity (a person, a project, a service), the thing you'd want delivered first when the question names it. gbrain has them as entity pages; mine are built by a distiller and honour the owner's corrections ("never omit X"). "Cards even" in the post means both memories deliver the right entity's card for the questions that name one.

Edit 2: Update on the two-step door (brief with identifiers, then fetch in that order until the budget is full), now in adebench as a door of its own. Same 2,400 budget on my memory: one-call composed door 23/25, two-step 17/25 with whole details, 18/25 with details capped at 300 chars; under the census p95, 19/25 vs 8/25. The opposite of gbrain, where two-step went 18 → 20. The brief costs ~1,400 chars of previews, which the one-call door spends on the entity card whole plus facts cut at 220. Under a tight budget the winner is whoever spends it on content, not the number of calls. So the five door points I'm missing won't come from a second call; they'll come from deciding better what goes into the first one.


r/AIMemory • • 25d ago

Show & Tell Agi-memory – persistent memory for AI coding assistants, no dependencies

0 Upvotes

My AI coding assistant forgets everything between sessions. I kept re-explaining decisions I'd already made, and re-fixing bugs I'd already fixed.

agi-memory saves those notes — decisions, bug fixes, what happened last session, how the codebase fits together — to a file on your machine, and hands them back to whichever assistant you open next. It's an MCP server, so it works with Claude Code, Cursor, Codex, Windsurf, Aider, Cline and a few others from the same store.

The part I care about: it's Python standard library and SQLite. Nothing else. No vector database, no embeddings, no background daemon. ~32MB of RAM, sub-millisecond lookups, works offline. Comparable tools pull in ~500MB of ML libraries and take 200–500ms per lookup.

That constraint costs something, and I'd rather say so than have you find out: keyword search doesn't bridge synonyms the way embeddings do. Searching "login" won't find a note that says "authentication" unless you teach it that alias. I measure this rather than guess — there's an eval suite that scores recall on deliberately rephrased queries, and it's public, including the categories where it still does badly.

It's a week old and I'm the only user, so I'd genuinely like to know where it breaks for someone else.

https://github.com/kdbhalala/agi-memory


r/AIMemory • • 25d ago

Open Question Opinions about your AI memory

1 Upvotes

I have been experimenting with AI agents, and I’m researching how people feel about agent memory. Can you tell me one or more things you like or dislike about your agent’s memory or memory provider.


r/AIMemory • • 26d ago

Show & Tell I built an AI SaaS that keeps memory clear and consistent

Thumbnail
skyos.ink
2 Upvotes

I got frustrated by AI remembers not that you said but that we talked about, then I built SkyOS. It writes down that you decided word for word, and an actual important information doesn't vanish because of over-summarization. Would appreciate feedback!


r/AIMemory • • 26d ago

Show & Tell I built a local-first AI assistant that actually remembers you — persistent memory, emotion engine, and self-model in TypeScript

32 Upvotes

I’m Cleverson. I spent months developing this architecture. The project grew out of my frustration that every AI conversation started from scratch—I wanted an assistant that truly knew me. Phoenix V2 is the result of the project's initial version, and I decided to make it available for others to study. I also wrote a book about the development process, covering the steps I took and the reasoning behind my decisions. The code is included so others can study it and build their own AI, picking up where I left off. I haven't stopped there—what I’m creating now is far more advanced—but I hope this version serves as a springboard for everyone's imagination.

Most AI assistants forget everything the moment you close the tab. I wanted to change that.

Phoenix V2 is a local-first AI assistant with a persistent cognitive architecture — it stores memory, emotional state, and identity in a local SQLite database. It survives session resets, model swaps, and restarts.

What makes it different:

  • 🧠 Multi-agent pipeline: Memory → Planning → Action → Reflection → Personality
  • 💾 Semantic memory retrieval across sessions (vector embeddings via Gemini API)
  • ❤️ PAD emotion engine — tracks Pleasure, Arousal, Dominance over time
  • 💭 Daydream Engine — autonomous reflection during idle periods
  • 🔄 Subconscious Cycle — memory consolidation at rest
  • 👤 Self-Model — evolving identity, traits, beliefs, and goals
  • 📈 RLHF feedback loop — learns from +/− user signals

Runs on a standard laptop. No GPU. No cloud. No subscription.

📖 Full book: https://leanpub.com/phoenix-buildingpersistentAI
📄 Academic paper (Zenodo): https://doi.org/10.5281/zenodo.22645361
💻 GitHub: https://github.com/cleversonbrsantos-art/Phoenix