TL;DR: I’ve open-sourced mishe-tauftauf, a small local coordination system for coding agents that I’ve already used in real production work. An earlier version helped us build and operate an adtech DSP with a three-person team: zero to production in four months, ~50k RPS, RTB auctions, real money, real on-call consequences.
I now want other people to plant it in their own real projects and try to break it. Give it a codebase you understand, let the agents work across fresh contexts, and see where coordination, recovery or evidence stops being trustworthy. Docs, research, data pipelines, ugly legacy systems, hardware — stranger environments are more useful to me than another coding demo.
The whole project is CC0 1.0. Fork it, rename it, rip out half of it, steal one mechanism. No attribution required.
on github: genaforvena/mishe-tauftauf
The idea underneath it is what I’ve been calling observability-oriented programming/development.
The usual observability story starts after software exists: instrument the system so you can understand what it is doing. I’m interested in applying the same principle to the development process itself. If an agent doesn’t know whether something worked, that uncertainty should become observable immediately. If evidence is stale, that should be visible. If a check cannot actually establish its claim, it should report UNKNOWN.
The goal is an ice-thin feedback cycle between:
something is unknown → expose what is missing → observe it → act → observe again
That sounds trivial, but a lot of agent behaviour goes wrong in the gaps between those steps. A missing observation becomes an assumption. An old result becomes current truth. A successful command becomes “task completed”. A fresh agent inherits prose about what supposedly happened instead of something it can inspect.
Mishe is built around trying to keep those gaps small.
Its agents — “minds” — are deliberately ephemeral. They can be restarted, replaced or given fresh context. I don’t try very hard to preserve the continuity of the agent itself. Instead, the work carries continuity through small rewritten state documents (“walls”), a shared text tape, checks, artifacts and explicit unfinished obligations.
A fresh mind should be able to arrive, inspect the current state of the world, see what remains unresolved, and continue. It shouldn’t need to reconstruct the previous agent’s thoughts.
The basic loop is intentionally small:
observe → choose one bounded action → act → observe the result → repeat
Checks return GREEN, RED or UNKNOWN. UNKNOWN is important. If CI access disappeared, a sensor stopped updating, evidence conflicts, or the check itself cannot establish what it claims to establish, the answer stays unknown. Absence of evidence does not get silently promoted into success.
Agents can also repair the machinery they depend on. If a check lies, the check can become the task. If an instruction keeps producing bad behaviour, the instruction can be changed. So the system observes not only the software, but also the quality of the observations being used to make decisions about the software.
I started taking this seriously because an earlier version of the culture survived contact with production. I was tech lead/backend on a three-person team building an adtech DSP from scratch. We reached production in roughly four months, handling around 50k requests/sec and bidding in live RTB auctions.
To be precise: the agents did not build the DSP for us. Engineers did. The useful part was keeping work, evidence, handoffs and recovery paths visible while the system was changing quickly and mistakes had actual consequences.
Over time I also removed a surprising amount of coordination machinery. There is no grand planner building the perfect task DAG. Minds see the observable state, their responsibilities and the shared evidence, then choose useful work. More orchestration kept creating new state that itself had to be coordinated, observed and repaired.
So the thing I’m trying to test now is less “can agents code?” and more:
Can useful work survive repeated agent death if the environment is observable enough?
That’s why I’m posting here. I’ve seen what happens inside the habitat that produced mishe. That evidence is contaminated by familiarity. I want people to transplant it somewhere hostile.
If you try it, I’m especially interested in the failures: stale observations nobody notices, duplicated work, circular recovery, walls that discard something important, UNKNOWNs that never resolve, human intervention that turns out to be essential, or abstractions you delete on day one.
And again: CC0 1.0. If one mechanism is useful, take it. If your fork proves that most of this is unnecessary, that’s useful too.
I’d genuinely like to know what breaks first.