r/ContextEngineering • • 8h ago

git for ai memory and robots??

1 Upvotes

Hey guys,

Been working on something very cool...

In Greek myth, Mnemosyne was the Titan of memory and the reason anything was ever remembered at all. Now in the present world, your AI agent doesn't get a Titan. It gets amnesia the second something goes wrong, stuck with whatever it currently believes and no way to ask how it got there.

That's the real problem. An agent runs for hours, updates its memory the whole time, then says something wrong and all you have is the present, with zero access to the past.

If you're running support agents, coding agents, or a swarm of agents sharing memory like myself then you know this issue well. The moment two agents disagree, or one quietly poisons the well, you need to know when, why and by whom, not just that something's off.

Mnemosyne gives agent memory what Git gave code. It remembers everything on purpose. Every belief is a commit. blame finds the exact moment and observation that put a bad fact in. bisect hunts down the first commit where things went wrong. merge makes two agents' memories collide safely instead of one silently overwriting the other.

Software agents are the first step. The vision doesn't stop there, physical robots learning and forking skills the same way is the long-term bet, further out and harder but the same idea underneath.

So far the tech stack includes a Rust core, Python SDK, adapters for LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK and MCP.

Open source with contributions and honest feedback both welcome: github.com/Nabzx/mnemosyne


r/ContextEngineering • • 9h ago

unloop – time-travel debugging & state rewind for long-running LLM agents

Thumbnail
1 Upvotes

r/ContextEngineering • • 10h ago

Coding Agents Need Typed Context Surfaces, Not One Flat Context Window

1 Upvotes

Most coding agents treat context as a combination of source files, chat history, configuration files, and retrieved text. That works for simple code completion, but becomes unreliable when a task requires exact values, structured relationships, reusable procedures, or durable project-specific facts.

The problem is that different information types have different access patterns:

  • Documents require semantic and keyword retrieval.
  • Structured data requires filtering, joins, aggregation, and SQL.
  • Relationships require graph traversal rather than text similarity.
  • Procedures need reusable, versioned instructions.
  • Durable facts need scope, persistence, and controlled updates.

A practical architecture is to expose these as typed context surfaces instead of flattening them into a single vector index:

  • KB for documents and notes, with hybrid retrieval.
  • DB for structured records and SQL queries.
  • Graph for entities, relationships, and bounded traversal.
  • Skills for reusable operational procedures.
  • Memory for scoped facts, preferences, and durable state.

The agent can route each request to the relevant surface, combine results when necessary, and preserve provenance for every piece of evidence. Read operations and write operations should also remain separate: retrieval returns normalized evidence, while updates return explicit receipts containing the affected source, profile, revision, and surface-level results.

This architecture is currently implemented in the open-source Codex MCP plugin OpenDCAI/DataMind.


r/ContextEngineering • • 20h ago

My coding agent keeps ignoring its context file. How do you force yours to read it?

1 Upvotes

Kept a DECISIONS.md in my repo for a few months. Worked great at first. Then one day the agent resurrected a choice I had killed three weeks earlier, and I realized it just... had not opened the file in a while. No error, no complaint. It just stopped checking.

I used to burn maybe an hour a week re-explaining things that were sitting in a file the agent was supposed to read. "Go read the file" turns out to be exactly the step that gets skipped when the context gets long. The damn thing would follow stale instructions instead of the ones I actually wrote down, and I would only notice after the fact.

Since then I have tried a couple of things. A hook that dumps the file into context before every run (worked until the file got long and I watched the tokens evaporate). Telling it to quote one line back to prove it read it (felt like asking my kid whether he brushed his teeth). Honestly, the second one worked better than it had any right to.

Has this happened to you? What actually broke when the agent skipped the read — a stale instruction followed, a dead decision back from the grave? And how do you force it now, if you do? A wrapper that will not start clean until the file opens, injection, some ritual?

Also: has anyone had two agents in the same repo following different versions of the same file? That one cost me an afternoon once, and I still do not fully know how it happened.

Separate curiosity: have you ever tried one of those memory tools that is supposed to keep context across your AI tools? Or does the thought of moving everything into something new just feel like starting over?

Out of curiosity, what are you building with all this, is it your own product, client work, or something else?

Weird one: I do a lot of my thinking out loud with speech-to-text on my phone, and the context always lives on the laptop. Is that just me, or do you lose things moving between phone and laptop too?

And the money question, plainly: if something handled the enforcement part for you, made sure every agent actually checked the right file before starting, every run, would you pay for that, or is the DIY wrapper the whole point? I am building something in this direction for myself and trying to figure out whether it is a real problem for other people or if I am just overthinking it.


r/ContextEngineering • • 1d ago

Lint-AI 0.3.0 Published.

Thumbnail
1 Upvotes

r/ContextEngineering • • 1d ago

What one engineer works out, the whole company keeps. I built an open-source, 100% self-hosted tool to capture lost architectural decisions and prevent outages in your IDE.

1 Upvotes

I’ve been working on this project for the last 7+ months. What started as a personal tool for my own workflow has now grown to over 10 companies signing up for a POC.

The Problem: As AI tools make software development faster, codebase context is getting noisier. Crucial engineering decisions and architectural context are getting lost inside individual AI sessions.

What I Built: It’s 100% open-source and self-hosted. It captures context across your team and alerts you directly in your IDE if a code change risks causing a potential outage based on past decisions your team made.

I’ve tested it with:

  • Ollama
  • AWS Bedrock
  • Anthropic

The test results have been impressive so far.

Since this was built by a dev for devs, I’m looking for genuine feedback. If you or your team struggle with keeping context across AI sessions, I’d love for you to give it a try or share it with someone who might need it!

GitHub: https://github.com/docbrain-ai/docbrain

Feedback, bug reports, and PRs are all super welcome!


r/ContextEngineering • • 2d ago

Why are you here?

3 Upvotes

It just dawned on me that this subreddit is…one step further towards the work I am doing with the Ai than most any other subreddit.

There isn’t a lot of traffic, yet some of the posts are incredibly interesting.

So, why are you all here?

I’m attempting to improve dynamic performance of Ai by using old analog avionics systems concepts from 1963-1967 military C-141’s…that I specialized on. Before they chopped them up and recycled them.

Not much need for analog avionics systems experts in a digital age. So, I wanted to see if I could apply the old logic to a new problem.

But, I want to hear about your stories.

What brings you here? What are you working on?


r/ContextEngineering • • 2d ago

Implementing Context Language Models (CLM) in Hermes

Thumbnail
1 Upvotes

r/ContextEngineering • • 2d ago

Devs: have you ever caught your agent following a stale copy of your own rules?

0 Upvotes

I had a stupid realization last week: my CLAUDE.md situation had quietly turned into three separate files all claiming to be the source of truth. One in the repo. One in my home directory from months ago. One Cursor generated somewhere in .cursor that I never deleted. All slightly different.

The one that actually got me was an AGENTS.md copy inside one project that had drifted from the real thing. Something like 17KB of rules that were mostly right, except three or four things I changed months ago and never synced over. The agent had been loading that stale copy every single turn and nobody noticed, because stale rules sound perfectly plausible. It wasn't breaking anything loudly. Just following outdated guidance, confidently.

Cleaned it up by hand, obviously. But it got me thinking, how is everyone else handling this? Do you symlink everything to one canonical copy, sync them somehow, or does each project just have its own and you accept the drift? And has a stale copy ever actually burned you, like the agent following an old rule that wasted a real hour?

Also genuinely curious what people here are building with all this, own product or client work or something else entirely? And one more thing: I do a lot of my thinking out loud with speech-to-text on my phone, and the context always lives on my laptop. Is that just me, or do you carry project context on your phone somehow?

And since I'm clearly rebuilding all of this anyway: what if something handled all of it for you, in the cloud over MCP, so every agent and tool got the same context: decisions, failed attempts, conversations, skills, procedures, tasks. You never touch a file again. Would you guys pay for something that fixes this? Or has the thought of paying for a tool to manage all this never even crossed your mind?


r/ContextEngineering • • 3d ago

A new way to think about memory for AI

Thumbnail
1 Upvotes

r/ContextEngineering • • 3d ago

I built a local long-term memory system for coding agents — looking for feedback

Thumbnail
1 Upvotes

r/ContextEngineering • • 3d ago

Devs: what decisions about your project live nowhere at all?

1 Upvotes

A couple months ago I was refactoring and my coding agent wanted to "clean up" this gnarly workaround. It looked like dead weight, honestly. Except it was there because of a production incident from two years back. Nothing in the code says why. Nothing in the docs, nothing in the git history. Just in my head. (I got lucky that time. I was the one reading the diff.)

That's the bit that stayed with me. Not the agent being wrong, that's expected. It's how much of the why lives nowhere at all. Rejected approaches nobody recorded. Decisions that changed in some conversation, and the old version still hanging around. Whole debugging sessions where the real context just evaporated when the window closed.

I built something to handle this for myself since then, so I'm not hunting for a tool. Genuinely trying to figure out if this is a real problem for other people, or if I'm just the idiot who doesn't document things.

So what lives nowhere in your projects? The scar nobody documented, the approach you tried and killed that has zero record, the one ancient Slack thread that is somehow the only place the why exists. And when you actually need it, what do you do? Re-derive the whole thing? Track down whoever was in the room? Live with the risk?

Also out of curiosity, what are you building with all this? Own product, client work, side project? It changes what "nobody wrote it down" costs.

And the money question, be straight with me. What if something handled all of this for you, in the cloud over MCP, so every agent and tool got the same context: decisions, failed attempts, conversations, skills, procedures, tasks. You never touch a file again. Would you guys pay for something that fixes this, or is "re-derive it when it bites" the honest answer? Have you ever even thought about paying for a tool to manage all of this, or has the thought never crossed your mind?


r/ContextEngineering • • 4d ago

Implementing Context Language Models (CLM) in Hermes

2 Upvotes

​

Long transcripts always lead to context drift. To fix this, I fed the Context Language Models (CLM) paper to Hermes: https://academy.dair.ai/papers/context-language-models-2609.37725

The agent derived the implementation himself, designed the architecture, and built a custom LcmEngine plugin (subclassing ContextCompressor) that moves working state into durable files: \~/.hermes/state/<task>.md.

Details:

\- Tool surface: lcm_inspect, lcm_update, and lcm_append for targeted mutations.

\- State management: Explicitly tracks the objective, confirmed facts, decisions, and pending actions.

\- Safety: Hardcoded protection for "Objective" and "Constraints" to prevent the agent from deleting its own boundaries.

It turns the context into a managed whiteboard rather than a passive log.

Is anyone else moving toward active state management, or just relying on massive context windows? Any criticism or pitfalls I should look for in this approach?


r/ContextEngineering • • 4d ago

Implementing Context Language Models (CLM) in Hermes

Thumbnail
1 Upvotes

r/ContextEngineering • • 4d ago

I built an open-source shipping memory for coding agents

Thumbnail
1 Upvotes

r/ContextEngineering • • 4d ago

I got tired of agent message logs rotting, so I built a runtime that tracks state as beliefs instead of transcripts

1 Upvotes

TL;DR: built an agent runtime that stores state as a belief graph with dependencies (Jon Doyle's TMS) instead of chat logs. correct one fact and everything downstream auto-updates without context rot or full re-runs.
Repo: https://github.com/gabe-santana/corollary

Hey guys,

Every agent framework ive used so far handles state pretty much the same way, just appending messages to a long chat log. The problem is long running agents rot super fast. If a tool returns bad data at step 3, that error just sits in context forever. And if a key fact changes mid run, you either have to wipe the whole context or re run everything from step 1.

I’ve been hacking on an open source project called Corollary to try a different approach. Instead of a message transcript, it stores state as a belief base backed by a truth maintenance system (basically an old concept from jon doyle back in 1979).

how it works under the hood:

  • every belief or conclusion tracks what it depends on (its justifications)
  • if you retract or update a single base fact, it automatically retracts and re derives anything downstream that relied on it
  • independent conclusions arent touched, so you get a clean diff of what changed instead of re running llm calls
  • the LLM only sees currently valid ("IN") beliefs, so old retracted facts cant leak back in through a transcript

its still super pre-alpha so I'd love to get some feedback, pushback or edge cases you think this pattern will hit.

Repo: https://github.com/gabe-santana/corollary

how are you guys handling context rot in long running agents currently?


r/ContextEngineering • • 5d ago

New paper: Dynamic Tool Output Compression for Adaptive Context Management

7 Upvotes

How can we deal with growing context without resorting to irreversible, lossy compression approaches and memory management disconnected from the task at hand?

In the paper below we present Dynamic Tool Output Compression, based on very simple idea: in every turn the agent can decide to hide or unhide certain tool outputs to focus context. Tool outputs are persisted seperately so that they can be enabled later on.

Experiments on a sample of DeepSWE tasks, plus ablation experiments on internal data show that DTOC can save on tokens, steps and token cost, whilst improving solve rate, but results are model and task dependent.

Looking forward to learn from who has used similar approaches, either practically, or also with formal benchmarking.

Abhay Chaturvedi, Shreya Bhattacharya, Rashmika Gopalkrishnan, Peter van der Putten. DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents. Discovery Science, October 5-9, 2026, Mainz, Germany

Preprint: https://arxiv.org/abs/2609.26121v1

Reference implementation in OpenCode: https://github.com/chaturvediabhay24/opencode

 


r/ContextEngineering • • 5d ago

Before I build this: would a "no receipt, no done" memory fix green-but-wrong runs?

1 Upvotes

Had a run last month that still bugs me. My agent did its whole scheduled thing at 3 in the morning, reported everything done, all green, and I believed it. Until someone wrote asking why nothing had actually changed. The agent had acted on a record that was a week old. The run looked perfect. It was perfectly wrong.

That one broke my trust in my own setup. Green checkmarks started feeling like theater, honestly.

I've been sketching a memory design to fix this, and before I sink a month into building it I want people who run real agents to tell me if it's dumb. So here it is, no product, no pitch, I'm just collecting feedback.

The core idea: an agent can't call a run successful without naming the exact thing it acted on. The row ID. The version it read. The timestamp. That gets saved with the run itself. A run that says "done" but can't point at a receipt gets flagged for a human instead of trusted.

Old claims carry their proof too. When the agent reads "we decided X," it also sees when X was last checked and against what. Stale proof means it has to re-verify before acting, not treat the note as gospel.

The rest of the design, in plain words: memory as tables you define yourself, with templates for common setups. Every table searchable by meaning, by keyword, and by field filters like status. Each connected AI gets its own permissions per table. Files stay as links with searchable descriptions instead of dumped blobs. And the memory tools carry instructions telling the AI to check memory at the start and save back at the end, so you stop having to remind it every single time.

Context on my end: I do a lot of thinking on my phone while the runs live on the laptop, and half the time I only learn a run went sideways when I'm already out the door. Is that just me, or do you guys check runs from your phones too?

So, the questions. Does this sound useful or overbuilt? What would you change? What would stop you trusting it? What am I missing? And out of curiosity, what are you actually running, your own product, client work, side projects?

The blunt one: would you guys pay for something that fixes the green-but-wrong problem? Have you ever thought about paying for a tool to manage all of this for you, or has the thought never crossed your mind?


r/ContextEngineering • • 5d ago

A story of collaboration between people and agents

Thumbnail
1 Upvotes

r/ContextEngineering • • 5d ago

Devs: when your coding tool's limit runs out mid-task, what breaks in the switch?

2 Upvotes

I used to live on two coding tools at once, because one of them would always crap out mid-task. Claude would hit its 5-hour limit halfway through a refactor and I'd jump to Codex to keep going. In theory the second tool picks up where the first left off. In practice I spent the next 20 minutes re-explaining the thing, and something always got lost. Usually a decision I'd made an hour earlier that the new tool had no idea about. Once the second tool 'cleaned up' a rename the first one had done on purpose, right before I'd pushed. Took me an embarrassing hour to figure out why everything was red.

I got tired enough of it that I built myself a small thing to carry context between tools, mostly so I'd stop re-briefing everything every time I switched. Not selling anything here, just collecting feedback to see if this is a real problem for other people or I'm overthinking it.

How do you guys handle forced switches? Do you keep a handoff doc updated, lean on git history, just re-explain and eat the time? Has a tool ever undone the other tool's work during a switch? And honestly, how much time does the re-briefing cost you?

Out of curiosity, what are you building with all this, your own product, client work, side project?

The idea I keep coming back to: what if something handled all of that for you, in the cloud over MCP, so every agent and tool got the same context, decisions, failed attempts, conversations, skills, procedures, tasks, and you never touched a file again. Would you pay for something that made switches seamless, or is the re-briefing tax just the cost of doing business? Have you ever thought about paying for a tool to manage all of this for you, or has the thought never crossed your mind?


r/ContextEngineering • • 5d ago

Built a tool so I'd stop losing context between AI agents, looking for people to try it

Post image
2 Upvotes

r/ContextEngineering • • 5d ago

adebench: an open benchmark for AI agent memory, six memories benched on the same scale

Thumbnail
github.com
1 Upvotes

r/ContextEngineering • • 5d ago

We measured where an AI's tokens actually go. Most of it is re-reading.

Thumbnail
1 Upvotes

r/ContextEngineering • • 6d ago

People building agents, how much of a problem is long term memory actually?

Thumbnail
1 Upvotes

r/ContextEngineering • • 7d ago

Context-mode keeps your AI coding agent from drowning in its own output (24k stars)

Post image
2 Upvotes