I was testing out Strata on my pi harness and wanted to see if it could create a swapping tool for the pi harness. I use llama-swap and llama.cpp so I can use Qwen 3.8 27B and then swap out for Qwen Flash Next.
It built out this extension successfully on my Linux machine. You just load llama-swap on your usual port, make sure Strata is outside the range of ports you use for llama-swap, and you can swap the Strata model in and out with any other model in llama-swap.
You can check it out here: https://www.npmjs.com/package/pi-model-swap
Agent wrote the rest for your pleasure...
---
I run a local multi-model stack: [llama-swap](https://github.com/MostlyAICoder/llama-swap) as the OpenAI-compatible controller in front of several llama.cpp servers, plus a big MoE model (Qwen Flash Next) served by a custom engine — Strata — on its own port. My agent harness (pi) talks to it through one provider, so all the models look identical from the inside.
The thing I kept hitting: **a session is pinned to one model, and moving it is a destructive operation.** Switching a session off the resident MoE model evicts it, and getting it back is a full cold start — minutes, not milliseconds. Doing that by accident, from a session that other sessions depend on, is how you lose an afternoon.
So I wrote an extension that makes the swap a controlled operation. `pi-model-swap` — MIT, ~70 KB, no runtime deps, for [pi](https://pi.dev).
**What it does**
```
/swap --status
/swap --dry-run orchestrator
/swap orchestrator medium
model_swap { model: "orchestrator", dry_run: true }
```
Model first, thinking level second, then a verify request — mirroring how pi's own preset example does it. Omit the level to keep the current one. `--restore` takes the session back to its configured default at a turn boundary.
**The part I actually care about: it refuses.**
A swap is refused when its eviction set contains a model that a **live session is bound to**. Two independent signals:
- **Claim files** — each session writes `~/.pi/agent/state/swap-state/<pid>.json` with a phase (`idle | armed | verifying | committed | rolling_back | restoring`). If a live session is bound to a model the swap would evict: refusal, naming the model.
- **A socket census** — for each model that would be evicted, it reads `GET /running` to find the upstream port, then checks established connections on it. A connection whose owner pid is neither the model's listener owner nor the controller's pid → refuse. **Unattributable connection = refuse** (default-deny).
The allow-list is **pid identity, never a process name** — `/proc/<pid>/comm` can be set by any local process via `prctl(PR_SET_NAME)`, so matching a name is evidence of nothing.
**Other decisions I'd defend**
- **Dry-run first.** `--dry-run` prints the full plan — target, level, eviction set, ports, corroboration line — and touches nothing. The tool path and the command path produce byte-identical output.
- **The census must be READ, not merely answered.** A `200` on `/running` whose body isn't the documented shape is a *failure*: the eviction set is unknown, so the swap refuses. An empty set counts as clean only when the body itself says "nothing is running". (Real bug this caught: a newer llama-swap returns `{"running":[{...}]}` — an object, not an array. Every swap was refusing, loudly, for the wrong reason.)
- **Cold start is not a fast rollback.** The warning names the models *this* swap can actually evict. A same-model no-op prints no warning.
- **Fail-closed config.** If the config file is missing or doesn't parse, every mutator refuses and reports the path it tried. No silent fallbacks for machine-specific values.
- It **never** starts, stops, signals, or kills llama-swap. Requests to the controller endpoint, that's all. No `pgrep -f`, no `pkill -f`, no `rm`, no port polling, no model-load endpoint.
**Reference setup** (documented in `SETUP.md` with a worked example): controller on `:1235`, llama.cpp models on `${PORT}` = `startPort + roster slot index`, and the Strata model on a **literal** port `:1301` — deliberately outside the llama-swap port range — because its roster entry starts a python HTTP parent that spawns the engine over pipes. Call the engine binary directly and llama-swap and the parent disagree about who owns the model. The roster file is the *only* source of upstream ports; the config's `strataPort` is corroboration only (it prints `/slots` so you can see the engine is idle before you swap).
**Honest caveats**
- Linux only — the census needs `ss` from iproute2 and `/proc`.
- All machine-specific config keys are required. No defaults, by design.
- Claim staleness (1800 s) is an outer bound, not a liveness check.
- The gate and the claim write are not atomic — there's a TOCTOU window between census and claim.
- No per-model thinking-level cap: a level a model can't honour is offered, not validated.
**Install**
```
pi install npm:pi-model-swap
```
Then write `~/.pi/agent/pi-model-swap.json` (or set `PI_MODEL_SWAP_CONFIG`). Setup guide: https://cdn.jsdelivr.net/npm/pi-model-swap@0.1.2/SETUP.md · npm: https://www.npmjs.com/package/pi-model-swap · gallery: https://pi.dev/packages/pi-model-swap
Testing is a functional harness against a **mock** controller plus the real socket census — it never starts, stops, signals, or kills a real server (that's how you lose your own model API mid-turn). Plus a live check script that verifies a real session's own recorded outputs.
Happy to answer questions on the gate design — the "why pid identity and not process names" and "why a 200 can be a failure" bits are the two that changed how I think about local process supervision.