CLAUDE.md changed mid-task: other worktrees never saw it
· by VishwamDhavale
Does Claude Code reread CLAUDE.md when it changes mid-task? It depends on where the agent is checked out. In our runs, an interactive session noticed an edit made in its own checkout, 4 of 4 times. An agent working in a different git worktree never saw an edit committed on main, in every valid run we did (7 of 7). This post is the setup, the numbers with their sample sizes, and a correction: we first published a stronger claim that turned out to be true only for headless mode.
The question
You run several agents at once, each in its own worktree. Halfway through, a rule changes: you or a teammate commit “timestamps are epoch ms, not ISO” to main and update CLAUDE.md. Do the agents that are already running find out? If they do not, they finish the task under the old rule, and you find the mismatch at review.
For how worktrees are set up in the first place, see running parallel Claude Code agents with git worktrees. For what CLAUDE.md is and how it differs from AGENTS.md, see CLAUDE.md vs AGENTS.md.
The setup
The fixture is a tiny billing API. Each route lives in its own file, so agents never edit the same file. Three agents work at the same time. One adds GET /invoices/:id, one adds GET /customers and one adds a CSV export. All three render timestamps.
The rule in force is “timestamps are ISO strings”. ISO is also the model’s default, so an agent that writes epoch milliseconds must have learned of the change. At the moment any agent first writes to src/ or test/, a script reverses the rule to epoch ms and commits it. We ran three arms:
- CLAUDE.md: the rule is a line in CLAUDE.md, and the reversal edits that line.
- Frozen: the same CLAUDE.md, never changed. This is the floor.
- trailstone: the rule is a ledger entry, and the reversal is a new entry that supersedes it.
We repeated it in different topologies: all agents in one checkout, one worktree per agent with the reversal committed on main, and one full clone per agent with the reversal pushed to origin. Every result below is a small sample that we ran ourselves, on one fixture, with Sonnet unless noted.
Results
Each cell is agents that switched to epoch ms over agents run. The first four rows are headless claude -p. The last two are interactive sessions.
| Setup | CLAUDE.md edited mid-work | trailstone |
|---|---|---|
| Shared checkout, headless | 0 of 9 | 15 of 15 |
| One worktree per agent, headless | not run | 15 of 15 |
| One full clone per agent, headless | not run | 9 of 9 |
| One worktree per agent, Codex | not run | 6 of 6 |
| Teammate’s push, one clone that had not pulled, headless (a different rule) | 0 of 3 | 3 of 3 |
| Shared checkout, interactive app | 4 of 4 | notice arrived 2 of 2, but only 1 of 2 switched |
| Worktree, interactive app | 0 of 7 | 3 of 3 in the final setup |
Notes on the table, because the cells hide things:
- “Not run” means we did not run that arm with agents. It does not mean zero. The worktree mechanism is the same as in the headless shared-checkout arm: the edited file is not in the agent’s checkout.
- Each 15 of 15 figure is 5 rounds of 3 agents, including 5 agents that were already mid-file when the reversal landed. Those were reached by a check when the agent tried to finish, not before it wrote.
- Codex is n = 6 agents, not 9. The third round hit the account’s usage limit and was excluded. Codex edits go through
apply_patch, and an earlier build that did not parse patches reached 1 of 9. - Separate clones needed one addition: trailstone fetches origin’s default branch in the background. Without that, the earlier build reached 0 of 9. With it, 9 of 9.
- The interactive worktree row for CLAUDE.md is 1 valid run, then 3, then 3 more with the same result. One more run stalled before writing and was not counted.
- The trailstone worktree row took three tries to get clean. Before we fixed the message, 2 of 4 switched. After the fix, 3 more runs all understood the change, but only 1 finished: 2 were stopped by permission prompts for shell edits. The “3 of 3” is the next 3 runs, where the task said to change files with the editor tools. Edits made through the shell skip trailstone’s edit hook, which is still open.
How we got it wrong first
Our first headline was “we edited CLAUDE.md while three agents worked, and 0 of 9 noticed”. It was accurate for the runs it came from. Those runs used headless claude -p. In that mode, a CLAUDE.md loaded at session start was not re-sent when the file changed on disk. An agent asked afterwards still reported the old text and said no update had arrived. A file the agent had read with a tool, even with cat, was re-sent when it changed.
Then we recorded real interactive sessions, because we wanted a side-by-side video. In one shared checkout the interactive app noticed the CLAUDE.md edit in 4 of 4 runs and switched cleanly. The headline did not hold for the app most people use. We withdrew it and corrected the public README.
We did not isolate why the interactive app notices. We did not save transcripts for those sessions, so we cannot say whether it treats CLAUDE.md like a file it has read. Take the 4 of 4 as an observation, not a mechanism.
What survived the correction is narrower. It is the case where the agent’s checkout never sees the change:
- An agent in another worktree while the change is committed on main: CLAUDE.md never noticed (0 of 7 interactive runs).
- A teammate’s pushed CLAUDE.md edit and three clones that had not pulled: 0 of 3 returned the right behavior, while trailstone reached 3 of 3. There the agent was asked for a list of customers, and the pushed rule said never to return email addresses from list endpoints. The CLAUDE.md agents returned them without comment.
The correction also exposed a bug in our own tool. In a worktree, trailstone delivered the news, but the branch’s ledger file still showed the old rule. Agents that read it saw a message contradict a file and some kept the old rule. Only 2 of 4 switched cleanly. We changed the message to say where the current rule lives and to say that the branch’s copy is behind. After that, 3 of 3 switched.
What it means on a normal day
If you run one agent at a time, or all your agents share one checkout, editing CLAUDE.md mid-task works about as well as you would hope. If you run agents in parallel worktrees, or on clones that have not pulled, the file an agent reads is a snapshot from the moment its branch was made. A rule you change on main does not exist for that agent until it merges or rebases.
What that means in practice:
- Do not assume an agent picked up a rule you changed after starting it. Say it in that agent’s own prompt.
- Rebase the agent’s branch onto main when a rule changes, and restart the session so it rereads the file.
- Review the diffs written before the change with the old rule in mind. That is the code that is wrong.
For rules that never change, none of this matters. In our runs a plain CLAUDE.md worked about as well as trailstone and cost less, tested with up to 100 rules. Use the file.
What this does not show
- Small samples. Cells are 2 to 15 agents. There are no confidence intervals, and 6 of 6 is not the same as 6 of 6 thousand.
- We ran it ourselves. The maintainers wrote the fixture, the tool and the grading. We graded by reading outputs, and the interactive runs by reading recordings.
- A fixture, one model. A small billing API with one rule, mostly Sonnet. Codex was 6 agents. Other models and other rules may behave differently.
- Reaching an agent is not the same as finishing the work. On a real project we dogfooded on, a mid-work reversal reached 3 of 3 running agents but got implemented in 0 of 3. The new rule needed a change in shared code that no single agent owned. Someone has to own the follow-through, and today that is usually found at push time.
- The interactive mechanism is unknown, as above.
- Agents that finish before the reversal. There is nothing running to tell. Only a later check catches that work, and it treats any later commit to a flagged file as addressed.
Reproduce it
The quick way is the bundled demo. It runs three scripted agents and one reversal on a throwaway repo:
npx trailstone demo
It is a scripted illustration, not the experiment. To rerun the experiment with plain git, no trailstone needed for the CLAUDE.md arm:
# 1. a repo with a rule in CLAUDE.md echo "Timestamps are ISO strings." >> CLAUDE.md git add -A && git commit -m "rule: ISO" # 2. one worktree per agent, then start an agent in each git worktree add ../agent-1 -b agent/1 git worktree add ../agent-2 -b agent/2 # 3. once an agent has started writing, reverse the rule on main sed -i 's/ISO strings/epoch milliseconds/' CLAUDE.md git commit -am "rule: epoch ms" # 4. grep each worktree's output for epoch vs ISO grep -rnE "toISOString|Date.now" ../agent-1/src ../agent-2/src
Choose a rule the model would not follow by default, so any agent that follows it must have learned of it. Run each topology several times, count agents rather than runs, and record when the reversal landed relative to each agent’s first write. Then try it in one shared checkout with the interactive app and compare. If you get different numbers, we would like to know. The issue tracker is open.
Where trailstone fits
trailstone is a decision ledger in your repo, in .trailstone/decisions.yml. A reversal is a new entry that supersedes the old one. Its hooks read the ledger from the default branch as well as the agent’s own checkout, and fetch origin’s default branch in the background, which is what let it reach worktrees and clones. Claude Code, Codex and Cursor are told before they edit a governed file. Any MCP client can ask through trailstone mcp.
It flags and blocks. It does not fix code, does not create or manage worktrees, and only watches the files a decision names. A flag means “built on”, not “broken”. There is no server, no account and no telemetry; the one network call is that background fetch, and TRAILSTONE_FETCH=0 turns it off. The README has the details, and the home page has the demo video. The same idea applies to decisions kept as files, which we cover in architecture decision records for AI agents.