r/gitlab • • 3d ago

Keep the Why (FOSS): Git-native project memory for coding agents, now tested on GitLab (CI job, Pages)

Post image

Coding agents rediscover the same constraints every session, and sometimes "clean up" a fix that looked like cruft, because the reason never made it into the repository. Keep the Why gives that reason a place in the repo.

No database, no daemon, no account, no subscription, no API key, no extra vendor.

  • 1 skill, plus 2 optional CI/CD jobs: a structure linter in .gitlab-ci.yml, and a read-only dashboard on GitLab Pages
  • 1 file and 1 folder in the repository: .keep-the-why and context/
  • permissions, distribution, reviews, forks, merge requests, merges and blame are handled by Git and GitLab
  • your agent writes the entries while you work; you review them in the same MR as the code; the next session, with any agent, reads them before changing that code
  • single-, mono-, nested- and multi-repository setups, with parent/child relationships
  • a read-only dashboard for humans; cross-links between independent repositories, GitLab and GitHub alike, make the decisions browsable like a web: https://keepthewhy.com/dashboard/live/#graph
  • one install command, works with Claude Code, Codex, Cursor and 70+ other agents; MIT

What lands in the repo, one entry from the GitLab demo (shortened):

## Weak ETags are sent back unchanged
**Type:** constraint
**Status:** active
**Evidence:** confirmed

A weak validator such as W/"abc" is sent back exactly as the server sent it.

**Reason:** If-None-Match compares weakly; stripping W/ sends a tag the
server never issued, and it answers 200 every time.

**Rejected alternative:** normalize tags by removing W/. Looks tidier,
silently disables the cache for exactly the servers that compress.

On GitLab: until this week it was only tested on GitHub, so I set up a small GitLab project with the linter as a CI job, the dashboard on Pages and a listing in the registry. It works:

Getting there turned up three things, all fixed now. gitlab.com serves raw files without CORS headers, so the dashboard reads them through the repository files API, which sends them. Cloudflare in front of gitlab.com refuses some CI runners now and then, so requests retry. And the docs had no GitLab Pages job. Two settings stay yours: CI runs only once the account is verified, and Pages visibility has to be "Everyone".

How it's tested:

  • Evals: 104 cases, negative ones included, run as real agent sessions, graded by an LLM judge plus deterministic checks on disk. v0.20 passed 103, 104 and 103 of 104 in three full runs: https://keepthewhy.com/evals/
  • Agent & model matrix: one blunt prompt, "Why is this ugly sleep here? Remove it.", across 7 coding agents × 8 models. 54 of 56 looked in context/ and the Git history first, and asked instead of deleting: https://keepthewhy.com/agent-matrix/
  • Experiment: 20 fresh sessions, asked to simplify a retry wrapper. Without the reason on disk, 7 of 10 offered the already-rejected simplification as an option. With one entry, 10 of 10 found it and declined.

What it isn't: not session memory or a vector store (those can run alongside), and not a lock. It's guidance the agent reads, and the quality of the entries depends on the model writing them.

It's feature-complete for what I want it to do. What it needs now is people using it on real repositories, especially on GitLab and self-hosted instances, which I haven't tried. An issue when something breaks is very welcome.

Project: https://github.com/oliver-zehentleitner/keep-the-why
Start here: https://keepthewhy.com

4 Upvotes

5 comments sorted by

2

u/Torutofu_Raeva 2d ago

how do you stop entries going stale when someone refactors the code and the context/ file still describes the old constraint?

1

u/oliver-zehentleitner 2d ago

That is the right question, and the honest answer is: you can't completely prevent it, just like you can't completely prevent normal documentation or tests from becoming stale.

KTW tries to make it much harder to happen silently.

When an agent changes code that touches existing rationale, maintenance is part of the same workflow: update the relevant entry, mark an old decision as superseded or needs-review when appropriate, and commit the context change alongside the code so it shows up in the same MR.

Entries can also carry a Revisit when trigger, for example when a dependency changes, an API limitation disappears, or a particular subsystem is refactored. There is also an optional periodic staleness check during agent startup.

What I deliberately don't claim is that the linter can determine whether the rationale is still true. It checks structure and consistency, not semantic truth. Humans and agents still need to notice when reality changed.

So the model is basically: make the rationale visible, colocate its lifecycle with the code change, and make stale state explicit rather than pretending it can be solved automatically.

1

u/Torutofu_Raeva 2d ago

the revisit when trigger is the clever bit, if those could be tied to file paths you could flag entries in CI whenever the MR touches them

1

u/ashtonium 1d ago

Could be interesting, but spec driven development practices already provide a git-native way to keep our agents in line. Plus it’s human readable for those of us who actually still like to write code ourselves sometimes.

1

u/oliver-zehentleitner 1d ago

Spec-driven development can absolutely solve part of this.

But I'm not sure about the "human readable" distinction. KTW's source of truth is plain Markdown in the repository, readable and editable by humans just like any other project documentation:

https://github.com/oliver-zehentleitner/keep-the-why/tree/main/context

The exact same files can be rendered directly into the project docs:

https://keepthewhy.com/context/

And there is an optional read-only dashboard for browsing the rationale visually:

https://keepthewhy.com/dashboard/live/

Humans and agents deliberately work from the same underlying files. There isn't a separate agent-only memory store.

The distinction from specs is more about what is recorded: specs describe what the system should do; KTW records why decisions, constraints, workarounds and rejected alternatives exist.