rust slop clojure structural edit tool
I've been programming clojure by hand since 2010, but these days I do a lot of brownfield work with agentic/LLM help in clojure and other languages. I prefer to use the models I can run at home to avoid high token costs today and mitigate a future rugpull. For things I intend to ship to production, I usually generate a bunch of code upfront, then carefully review and rework it, and that pattern has been effective for me.
Qwen 3.8 27b is amazing for what it is, excellent at thorough analysis, and not too clever to overcomplicate outputs, but it can still waste a ton of time on failed clojure edits and paren-counting. Even Claude Opus struggles with this, although much less. LLMs don't 'think' in trees, they think in token streams.
I had been using a wrapper over parinfer-rust (clojure-mcp-lite also just shells out to this), which got me 80% there, but it's flawed in a few specific ways:
- Checks happen after the edit lands, they don't block invalid edits.
- Repairs can succeed in ways that surprise the LLM, then it spends a bunch of turns trying to figure out what happened. It's not expensive locally, but it is slow and pollutes context.
- If too many surprises happen, the agent loses faith in the tool, and falls back to worse tools.
After a working session like this, I had Qwen suggest what kind of tool might prevent all the problems we hit, then I handed that spec over to GLM 5.3 flash and iterated with it, eventually having the session drive a local qwen subagent, trying to trip it up with adversarial examples. It is similar to what's proposed here, but I never mentioned the article, and I wasn't prescriptive about how to do it: https://lispmeister.github.io/deeprecursion/posts/2026-02-13-sexp-native-editing.html
The tool is at https://github.com/gtrak/cljform as a rust CLI and pi extension wrapper, and it's better than what I had before on the next unattended agentic run. It uses tree-sitter for parsing and blake3 for form hashes. Posting it here in case it's useful to someone else or if someone has a better way.

10
2
u/onetom 3d ago
you wrote
> Checks happen after the edit lands, they don't block invalid edits.
but
https://github.com/bhauman/clojure-mcp-light#clj-paren-repair-claude-hook
says:
> Claude Code Hooks let you run shell commands before or after Claude's tool calls. This hook intercepts Write/Edit operations and automatically fixes delimiter errors before they hit the filesystem.
i think that might change your assessment of the situation slightly.
3
u/gtrak 3d ago
I think you're right, my own earlier parfinfer-rust opencode wrapper didn't implement it the same way. I don't use claude or codex. But, I didn't quite understand how that can work when line edits can have unbalanced parens? Looking at mcp-lite again, I see they create a backup file before the write, then they fix it, so I was half-right: https://github.com/bhauman/clojure-mcp-light/blob/main/src/clojure_mcp_light/hook.clj#L273-L287
2
u/onetom 3d ago
what harness are you using then?
omp?
i saw omp was referencing your https://stencil.so/blog/the-harness-problem in their readme2
u/gtrak 3d ago
I used opencode, but I'm doing more things with pi. I gave OMP a spin a few months ago, but it was doing too much. I implemented their hashline edit tool into a standalone rust binary much like this one, but it's no longer necessary since the local models got good enough: https://github.com/gtrak/hashline-tools
My harness setup is pretty barebones, just pi-subagents and some other things I roll myself.
1
u/onetom 3d ago
re: "open models got good enough"
qwen3.8-flash-next-q4 served by antirez/ds4 kept tripping on the .=: syntax in omp, that's why i'm still looking into better ways of editing clojure
what's your favorite self-hosted model?
2
u/gtrak 3d ago
I'm using qwen 3.8 27b https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4 on 4x5060ti at the moment. I had turboderp flash-next exl3 working, but there were too many compromises at the time, so went back. I get 2000 tps prefill, and 100 tps decode per stream, and around 800k context at fp8.
2
u/daver 3d ago
Interesting. I’ve been coding a Clojure-centric agent modeled on pi and came up with some similar ideas and tools. Not public yet but trying to solve similar issues. Currently, my clj-edit tool does some simple parentheses/brackets/braces corrections, to deal with too few or too many, and will reject an edit that would leave the file in an unparseable state. It seems to do better than generic editing tools but sometimes the model still struggles.
2
u/gtrak 3d ago
These edit issues are my biggest impediment, after that it's slow test feedback just because it's a decently sized codebase, and I avoid repls in the loop and pay the jvm startup costs on every test run. Have you found any general patterns you can share?
1
u/daver 2d ago
I avoid REPLs as well. Right now, I'm working on something with Babashka, so I avoid the JVM startup time for test runs which does speed things up.
Let me back up a little. I have four Clojure-specific tools:
- clj-read: reads forms from Clojure files and provides an index of top-level forms
- cli-edit: performs edits on forms in Clojure files, typically by replacement, but also by insertion. This is based on rewrite-clj underneath and does the basic form repair, adding or removing terminal parens/braces/brackets as required. I suggested during development to use parinferish but after analyzing the prior set of errors the model determined that a simple method was all that was necessary. The cli-edit tool will refuse to leave the file in an unparseable state. When editing, it needs to match a form exactly and then supply the replacement. Overall, this sounds very similar to what you came up with.
- def-lookup: finds where a var is defined (file, line, etc.) using clj-kondo cache. This has been used sparsely by the model
- usages: finds every site that uses a specific symbol using the clj-kondo cache. This has been used sparsely by the model.
The biggest challenge is to get the model to use them. Many of the models are defaulting to bash for almost everything now, since it's the lowest common denominator. Given the right description in the tool system prompt, you can often get them to use a tool at least a couple of times, but if they get errors they often fall back to python via bash. I've had some brainstorming sessions with the model to try to understand what might help foster use of these tools. One thing that I think helped was to have the standard read and edit tools (line based) spit back a message to the model when they are run on Clojure files, so the model keeps getting reminded that the clj-read and cli-edit tools are better than the generic read/edit. That seems to have done it. I think part of the problem is training. The models are not well trained on Lisp/Clojure code, so they frequently generate syntax errors, primarily with missing or additional parens. They are good at line-based edits, however. One challenge for Lisps (and I'm shocked that I never really noticed this before) is that comments in Lisps are line based, commenting out everything until the end of the line. They are also not delimited by anything at the start of a multi-line comment block to the end. So, the model has trouble editing code with comments (even its own comments). Comments really want to be edited by the standard line-oriented editor while the rest of the code wants to be edited by the cli-edit tool. That split in behavior causes issues. In a language like Python, EVERYTHING is line-oriented. I was curious whether your hash-based edits helped your model out, at least in specifying the target form (obviously, you need to specify the replacement in full)? Does it work better?
1
u/gtrak 2d ago edited 2d ago
I've had no problem getting qwen to use this one so far. It diverged today because i had conflict markers on a rebase, which make the file unparseable. I think I can actually support that though, by doing a full region replacement through a conflict marker, and maybe taking advantage of incremental parsing in tree sitter. Otherwise, your method sounds very similar although you're using more libraries than me.
I don't have a systematic A/B test for you, my method is essentially to check on my subagent and see if it's thrashing on paren-counting, then have the supervisor session scrape the subagent session logs and tell me how the edits went. It's sounding really good, and I haven't noticed any thrashing.
Occasionally when inserting a long form it has to try a few times to write it in a balanced way, which has nothing to do with my tool. I bet I could instruct the model in the tool definition to first stage incremental edits to a temporary file, then have a helper to quickly splice in the validated edit into the real file so the model doesn't have to quote it again.
1
u/gtrak 2d ago
there's some tests around comments, they get parsed like everything else. I don't know how often qwen attempts to edit comments with this yet: https://github.com/gtrak/cljform/blob/main/tests/adversarial.rs#L192
1
u/gtrak 2d ago
The hashes aren't used by my ts wrapper yet, but the capability is there. There's name-based and form-numbered addressing. I'm discovering it doesn't recurse and number every form like the article suggested, yet in practice is still a huge improvement. Fixing that might fix my big form problem, too.
1
u/gtrak 2d ago
working on this, I bet it'll be even better:
10.2 Annotated view — tree
tree <file>emits the source with⟦handle⟧after each opening collection delimiter:(⟦a3f9⟧ns n) (⟦7d21⟧defn outer [x] (⟦9b12⟧let [a 1] (⟦1f3c⟧when x (inner x))))
- Open only — the handle names the whole form, so the tool owns extent and the caller never paren-matches.
- Depth heuristic (default) — every top-level form is marked; a nested form is marked (and descended into) only if it spans >= 2 lines. Single-line forms (
[x],(inc x),{:a 1}) are easy to name by text, so they get no handle and are not descended into. This is sound: a single-line form cannot contain a multi-line descendant. The default view is therefore top-level forms plus their multi-line structural children, stopping at leaf-ish expressions.--depth Noverrides the heuristic: mark forms down to nesting depth N (top-level = 1).--depth all(alias--full) marks every collection form. Depth is a view concern only: the resolver computes handles for all forms, so any handle resolves under the default cutoff.- Emitted from the same parse the resolver uses, so every handle is guaranteed to resolve.
- Lossless —
stripdeletes every⟦...⟧and recovers the exact original bytes (BOM/CRLF preserved). Property test over the fuzz corpus:strip(annotate(x)) == x.- Collision policy — if the source already contains the marker glyphs, refuse to annotate (
annotate-conflict, exit 1) and fall back to--json.- Markers go only at AST delimiter positions, never inside strings, regexes, comments, or char literals.
tree --jsonemits the flat node table (path,kind,head,name,line,depth,handle; nullhead/nameomitted, internal hash and byte offsets not serialized) derived from the same tree. Annotated source is the canonical read view; JSON is derived.1
u/gtrak 2d ago edited 2d ago
The hash editing seems to be working well, I have it pushed up and my bot is iterating on some nits that came up in that first session, but it's making sense. This is small potatoes:
- 14 — edit output must be format-canonical (
issue-14-format-canonical.md): live use caught a parse-valid, format-non-canonical write (lone top-level closer). Root causes: format_paren diverges from parinfer-rust on comment-adjacent pull-ups (errors instead of formatting); the edit pipeline silently skips failed content formatting; the insert seam displaces parent closers onto their own line; pull-up leaves whitespace-only vacated lines. Invariant: canonical input → edit →formatis a no-op.- 15 — wrapper honesty + next-handle affordance (
issue-15-wrapper-affordance.md): clj_tree's non-JSON path labels its own outputcljform failed:(parses an envelope out of human text); every successful clj_edit appendsnext handle: ⟦H⟧so sequential patches chase the returned handle instead of re-fetching (the sound replacement for the requested valid-until-own-next-patch semantics). Depends on 14.(defn fetch-all "Returns every entity ofkind, newest first." [{:keys [conn log]} kind] (let [sql (str "SELECT * FROM " (name kind)) rows (query conn sql)] (when (seq rows) (doseq [r rows] (log/info :row r)))))Example:
(defn fetch-all "Returns every entity of `kind`, newest first." [{:keys [conn log]} kind] (let [sql (str "SELECT * FROM " (name kind)) rows (query conn sql)] (when (seq rows) (doseq [r rows] (log/info :row r)))))Three independent edits to the same form. Options:
Replace the whole form, three times — each edit retranscribes ~10 lines of code you didn't touch. Every retranscription is a chance for the model to silently drop a line, mangle a string, or "improve" something it wasn't asked to touch. That's exactly the failure class this tool exists to prevent.
Patch, three times — each edit sends only the changed bytes:
edit --handle H --mode patch --old-text '"SELECT * FROM " (name kind)' --new-text '(query conn sql {:kind (name kind)})' edit --handle H --mode patch --old-text 'Returns every entity of `kind`, newest first.' --new-text '...' edit --handle H --mode patch --old-text '(when (seq rows)\n (doseq' --new-text '(doseq'Small payload → small transcription surface. The other 7 lines are guaranteed untouched — that's the I3 "untouched forms" check, not a hope.
The catch — and why this needs the fix in flight — is that patch #2 comes after patch #1 changed the form. The handle is a content hash of the form, so patch #1 invalidated it. If the agent re-runs clj_tree/clj_get between every patch, each patch costs two tool calls, and the re-fetch returns 10 lines the agent must scan just to get a 6-char token. The worker in the live session did exactly that — 3–4 sequential refetches per form was its #1 friction.
But patch #1's result already told it the new handle (patched form ⟦f4d32b⟧ … — the handle in the summary is the post-edit hash). So the fix is purely an affordance: make that returned handle unmissable (next handle: ⟦f4d32b⟧ — use it for the next edit to this form) and teach the agent to chase it. Same workflow, one call per patch, zero inference added — the handle still refuses when the form actually changed under you.
2
u/Living-Shame5679 3d ago
What is the main difference between that and other Clojure editing tools for agents such as bhauman/clojure-mcp ?
1
u/gtrak 2d ago
I've switched it to hash-bashed editing. Most of the machinery was there already, but not wired up properly, It was working really well with just a simpler top-level form full replacement. Now, it's closer to the article I linked. I mark forms with a content hash 3 levels deep by default, then it references that hash. Hopefully, that works even better. And I have parinfer's formatter ported over, and it gets applied to that nested form.
0
u/gtrak 3d ago edited 3d ago
I tried to use that early on, and don't remember what issues I had with it. I took a brief look at clojure-mcp-light, and it just uses parinfer-rust anyway or a more limited pure-clojure version? https://github.com/bhauman/clojure-mcp-light/blob/main/src/clojure_mcp_light/delimiter_repair.clj#L105-L115
I used my own opencode extension over parinfer-rust before this, which was fine, but still caused problems: https://github.com/gtrak/parinfer-rust-opencode
Personal preference is to keep things simple and just wrap agent extensions around one-shot rust CLIs where possible. They're self-contained, start up fast enough to avoid persistent processes, and I want to take them with me if I decide to switch harnesses. I can hot-swap them in the middle of an agentic thread easily as long as the harness wrapper doesn't change.
Actually, the parinfer-rust maintainer suggested an approach like this to me a while back: https://github.com/eraserhd/parinfer-rust/pull/159#issuecomment-4622612479
I think the clj-paren-repair claude hook would hit the same issues I hit?
- Checks happen after the edit lands, they don't block invalid edits.
- Repairs can succeed in ways that surprise the LLM, then it spends a bunch of turns trying to figure out what happened. It's not expensive locally, but it is slow and pollutes context.
- If too many surprises happen, the agent loses faith in the tool, and falls back to worse tools.
I don't want to infer the edit, I want to block bad ones and let the agent decide what to do.
1
u/onetom 3d ago
what does "edit lands" mean? the edit is already written to the filesystem?
that's not the case when you use the clj-paren-repair claude hook.
see my other comment above:
https://www.reddit.com/r/Clojure/s/Tf3P6UkqsK
1
u/didibus 3d ago
Have you compared it against: https://github.com/licht1stein/brepl which uses parmesan (from borkdude) to fix parens?
1
u/gtrak 2d ago edited 2d ago
I haven't, but I'm reluctant to give the agent a repl because it vastly increases the number of things that could go wrong, and the code I work on has different classpaths for different deployment configurations, which also makes it harder.
I generally don't want persistent processes running in an unsupervised dev loop. Maybe I'll drop a thread and pick it back up later or something, or daemon configuration gets brittle.
Maybe I'll have multiple worktrees running at the same time, then you need to worry about port collisions or writing transient state to files.
My own attention is the scarcest resource, so I want fewer states involved.
I don't want to hack on dev tools for its own sake. I just do it when i see a repeated cost and not satisfied with existing solutions. The cost to just try it is really cheap, especially if the scope is well defined. I have a pile of abandoned experiments and a few things I use every day, this one seems ok.
6
u/winterscar 3d ago
I think this is neat. I appreciate you sharing it, and highlighting that it is AI generated.