Why trailstone: decisions change, agents keep building

· by

AI coding agents keep building on a decision you already changed whenever the change lands somewhere their checkout does not see: another worktree, or a clone that has not pulled. It keeps building on the old one, and you find out at review or in production. trailstone exists to close that gap: it keeps your decisions in the repo, records which files each one governs, and when you reverse one, tells you and your agents exactly what now rests on the old rule. This post is the idea, what we measured (including where we were wrong), and where it could go.

The problem, in a developer’s day

You are running three agents on a feature. Around lunchtime you change your mind: sessions use cookies, not JWTs. You update the rules file. You tell the one agent you are watching.

The other two are mid-task, in their own worktrees. Neither has been told. Both keep writing JWT code that compiles, passes its tests and looks finished. Nothing is red. You find out at code review, if the reviewer remembers, or in production if they do not.

Two things go wrong here. The agents drift: what was decided ten messages or three sessions ago quietly stops governing the work. And the reversal silently invalidates everything built on the old decision, which stays green, finished and wrong. Agent memory stores facts and preferences. It does not know what a decision governs, so it cannot tell you what a reversal breaks.

The idea

Record the decision in the repo, next to the code, and have it name the files it governs:

trailstone decide "Sessions use JWT in an Authorization header, not cookies" \
    --why "mobile clients cannot hold cookies" --scope src/auth/

trailstone reverse <id> "Sessions use a signed HttpOnly cookie, not a JWT header"

Because the decision knows its scope, the reversal can point at exactly the work that rested on the old rule: any tracked file in that scope whose last commit is older than the reversal. That is the whole trick. The agent is told at the moment it touches a governed file, with the old text and the new one, and the push guard and CI check fail on a flagged file until someone has looked at it.

Staleness is derived from git history and scope every time you ask. It is never stored, so it cannot go out of date. A decision is never edited. Changing your mind is a new entry with supersedes:, so a reversal reads as a diff. See architecture decision records for AI coding agents for why plain ADR files do not get this far.

How we got here: a server, then git

We did not start here. The first version was a hosted service with accounts, a database and its own sync. It worked, and it was the wrong shape. Then we noticed that git already had everything we were building: history with authorship, review through the pull request, and distribution through clone and fetch. A decision that lives in the repo is reviewed like code, blamed like code and shipped like code.

So we deleted the server. What is left is one script and one file, .trailstone/decisions.yml. No server, no account, no telemetry. The one network call is a background git fetch of origin’s default branch, and TRAILSTONE_FETCH=0 turns it off.

What we measured, and where we were wrong

We tested trailstone against the free default, a plain CLAUDE.md. These are small samples that we ran ourselves, so read the denominators.

  • For rules that never change, we tie CLAUDE.md. Up to 100 rules, agents honored a plain CLAUDE.md about as well as trailstone, and it costs less. If your decisions stay put, you do not need this. We wrote that into the README instead of hiding it.
  • For a decision reversed mid-work, agents in separate worktrees: trailstone reached 15 of 15 running agents.
  • Separate clones: trailstone reached 9 of 9. A teammate’s pushed CLAUDE.md edit reached 0 of 3 clones that had not pulled, against 3 of 3 for trailstone.

We were wrong once, in public-facing wording. Our early headline said a running agent never notices an edited CLAUDE.md, 0 of 9. That was true of headless claude -p sessions only. When we ran the interactive Claude Code app, it noticed an edit to CLAUDE.md in its own checkout 4 of 4 times. We withdrew the headline. The claim that survives is narrower: the gap is wherever an agent’s checkout never sees the change, meaning other worktrees and clones that have not pulled. The details are in why a CLAUDE.md edit never reaches other worktrees.

One more result we owe you. On a real project we dogfooded on, a reversal reached 3 of 3 running agents and none of the three implemented it, because the change lived in shared code that needed an owner. trailstone made the agents stop and know. It did not fix the code. It does not claim to.

The invariants

We treat five rules as the constitution. Changing one is a different product, not a feature.

  1. Derived, never stored. supersedes and scope are the only stored edges. Everything else is computed from git on every read.
  2. Exact on anything that blocks. A flag needs a supersession plus an exact scope match: a path, a directory prefix or a glob. Never a guess. A false flag is worse than a missed one.
  3. Minimum sufficient surfacing. Never dump the ledger. Push what is critical, pull what is relevant, nothing else.
  4. You outrank the ledger. A recorded decision gives the agent standing to refuse a contradicting request and to ask. It never lets the agent overrule an explicit instruction from you.
  5. A flag means “governed by”, not “broken”. Reversing a directory-scoped decision flags every file under it. Most will still comply and clear in seconds.

The format is the contract

The important artifact is not the script. It is .trailstone/decisions.yml, a small YAML shape with every field documented in the README. The script is the first implementation of it. Anyone can read the file and build on it: an editor plugin, a different enforcement point, a port to another language. We will keep the format stable, and we would rather review your shim than write five integrations badly ourselves.

Two layers are already universal. The ledger is plain YAML, and enforcement is a pre-push hook and a CI check, which do not care what wrote the code. Asking “what governs this file?” works for any MCP client through trailstone mcp. The layer that needs a shim per tool is the warning pushed before an edit, unasked. Claude Code, Codex and Cursor have it today.

Where it could go

These are directions, not a roadmap and not promises. None is reserved. If one is what you need, build it.

  • Push shims for the remaining tools. Windsurf and Claude Desktop are still pull-only. A hook shim that gets unasked pre-edit warnings working in one of them is the contribution we want most, especially from someone who lives in that editor.
  • Teams. One person’s reversal reaching another at their next relevant moment. Across clones it already works through the fetch of origin’s default branch. What is missing is real teams using it.
  • A GitHub App that turns a stale flag into an annotation on the exact lines, with re-affirm and supersede actions in the pull request.
  • Richer scope, from files to globs to symbols or modules.
  • Scope suggestion from the diff, so a narrow scope is the easy path.

What trailstone will not become

  • It will not fix your code. It flags and blocks; an agent or a person makes the change.
  • It will not guess. Lexical matching may help surface a decision, but it can never flag a file or fail a push. We also do not follow imports to flag files that merely depend on a flagged one, because that would trade precision for a flood of guesses.
  • It will not follow renames today. git mv a governed file and it leaves the decision’s scope until you re-scope.
  • It will not replace your CLAUDE.md or AGENTS.md. Use those for rules that hold still.
  • It will not run a server, hold an account or phone home.
  • It will not save you from a broad scope. A bare src/ flags every file under it, and most of those never touched the decision. The discipline is yours.

Try it

npx trailstone demo runs the whole idea on a throwaway repo in about ten seconds. If you run agents in parallel, the worktrees post shows the setup where it matters. If it breaks on your repo, that is the report we want most, and trailstone report --anon makes it easy. The code and the README are on GitHub, and the home page has the short version.