Skip to content

review: harness-mocks changes since the 2026-10-03 demo (do not merge) - #240

Draft
n-sviridenko wants to merge 40 commits into
review/demo-baselinefrom
main
Draft

n-sviridenko wants to merge 40 commits into
review/demo-baselinefrom
main

Conversation

@n-sviridenko

Copy link
Copy Markdown
Member

Comment-only PR: the diff from the last harness-mocks commit of the 2026-10-03 capabilities-matrix demo (9a26dad, #138; sloprail-community demos/20261003-capabilities-matrix) to current main. Base branch review/demo-baseline is pinned at 9a26dad. Do not merge.

Compare: 9a26dad...main

🤖 Generated with Claude Code

https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

n-sviridenko and others added 30 commits October 4, 2026 10:26
…l pinned binaries (#139)

* snapshots: pin harness versions; a binary bump no longer re-judges every capability

The harness binary a recording was made with is per run (run.yaml); the
MANIFEST's top-level `version` (and each page's `version`) are gone. A doc
page is frozen by its own sha256 and `fetched` date, and MANIFEST `pin` only
names which binary capture.sh runs next. touched.sh treats only a page whose
sha256 changed as in question, so bumping the binary or re-freezing a page
with an unchanged hash invalidates nothing; verdict keys carry page hashes
only. The three MANIFESTs are migrated (hashes untouched, fetched = the date
each hash was first committed).

tools/harness-bin installs an exact claude (npm), codex (npm) or cursor-agent
(versioned download) into a cache of its own, never the global install, and
`path` refuses a binary that does not report exactly the version. capture.sh
runs that binary with auto-update off, links the user's login as before, and
refuses to add a sample to a run recorded at another version unless
--rerecord. `capture.sh pin <version>` installs and sets the pin.

Sloprail-Cites-User: this ownership of different versions of harnesses probably also has to be a part of the infrastructure of mocks
Sloprail-Cites-User: if the version is just updated, it doesn't mean that we need to rescan all the features. So there has to be a mechanism on how to kind of freeze it.
Sloprail-Cites-User: we can kind of pin the things

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* harness-bin: refuse an escaping symlink in a tarball; capture.sh all checks the pinned binary before dropping samples, pin never overwrites a MANIFEST

Review fixes: untar accepted absolute and ../ symlink targets; `all` removed every run's samples before discovering the pinned binary was missing; `pin` fell back to overwriting an existing MANIFEST (and its docs) when yq failed.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* adr(pinned-harness-versions): keep to the decisions the user stated and the rules check

Sloprail-Cites-User: if the version is just updated, it doesn't mean that we need to rescan all the features. So there has to be a mechanism on how to kind of freeze it.
Sloprail-Cites-User: we can kind of pin the things
Sloprail-Cites-User: this ownership of different versions of harnesses probably also has to be a part of the infrastructure of mocks

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* adr(pinned-harness-versions): link only the rules that decide it; drop the unverifiable cache sentence

Sloprail-Cites-User: if the version is just updated, it doesn't mean that we need to rescan all the features. So there has to be a mechanism on how to kind of freeze it.
Sloprail-Cites-User: we can kind of pin the things

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
The marker was written as "# sr-mark: ci-verify"; the gate
sloprail/gate/ci-verify-required looks for "sr:ci verify" (written by
`sr-mark apply ci --verify=<path>:<line>`), so it kept reporting that no
committed file shows CI verifying the file-guards. Same step, same comment
position; only the marker text changes.

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…on (#149)

* ci(claude-mock): release darwin binaries, checksums.txt and --version

The release now builds linux and darwin (amd64, arm64) with the tag embedded
as main.version, and uploads checksums.txt. a10n-claude-mock --version prints
the embedded version (dev when unstamped). The mock never answered --version
for scenarios, so nothing existing is shadowed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* docs(adr): release-version — a mock's --version is its release, stamped at build time

Sloprail-Cites-User: confirm the 3 release-version bullets

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(sloprail): adr-conformance also judges the claude-mock release workflow

adr/release-version's first decision (the release stamps the tag into the binary) lives in .github/workflows/publish-claude-mock.yml, which the rule's match never covered.

Sloprail-Cites-User: Make the rule also watch `publish-claude-mock.yml`

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
#155)

* fix(ci): the e2e git identity is test-process environment, not the machine's global git config

A workflow step wrote the global git config; copied onto a developer's machine it rewrote
their real ~/.gitconfig (commits of PR #142 were authored ci-local@sloprail.invalid). The
test-e2e target now exports GIT_AUTHOR_*/GIT_COMMITTER_*, the step is gone, and
tools/repoguard fails on any workflow/Makefile/script/Go source that writes the global or
system git config.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(repoguard): see commands split over lines, --file on a home gitconfig, more file kinds

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(repoguard): comments never join a command, real line numbers, XDG config via --file, deleted tracked files

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(repoguard): a Go split joins only inside an open call; every line is also judged alone

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(repoguard): judge each shell command of a line alone; join a multi-line slice literal

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(repoguard): a read inside a substitution or comment no longer exempts a write; one report per command

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… (pilot) (#140)

* codex: replay every recording through the mock's own `replay` command (pilot)

The mock replays a recording itself: `a10n-codex-mock replay [--print-script]
<run-dir>` turns the recorded run's model turns into a scenario script in the
format any scenario has (in memory, nothing stored), runs the mock on the
run's own setup, and has the shared core (internal/replay) compare the mock's
whole event stream and hook payloads with the recording's. Exit 0 green, 1
differs, 2 not replayable yet. The table-driven test only enumerates
snapshots/runs/* and calls the same function once per recording folder.

Recordings that do not replay green are on an explicit allow-list with a
reason; an entry whose run is gone, or that now replays green, fails the test
(the list only shrinks).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* replay: the core speaks one unified format, every heuristic is the harness adapter's

The core (internal/replay) takes a recording as normalised turns (agents, calls,
a unified tool vocabulary) and compares normalised output; it holds no heuristic.
Reading a rollout into turns, the mock's script format, the dropped/rewritten
fields and the ordering of hook payloads move to the codex adapter
(codex-mock/internal/replay), behind core.Adapter so claude and cursor can add theirs.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* replay: the adapter starts children through procexec, takes its environment from the entrypoint

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* replay: an unparsable recording line fails the replay; the codex mock refuses feature switches

The adapter reads the recording's stream, hook log and rollouts strictly: a line that is not
a JSON object is an *Unbuildable naming the file and line, not a line left out of the comparison.
The codex mock refuses --enable / --disable (it implements no feature switch), so a run asked
for `--enable multi_agent_v2` fails loudly instead of passing as the default mode.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* codex-mock: the feature-switch refusal is a file of its own (exec.go stays under the size limit)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* scenario module: codex-mock/feature_switches.go is in its home beside exec.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* replay adapter: run.yaml and shell command lines are parsed, not scanned; the allowlist only shrinks

- load.go reads run.yaml with a YAML parser into a typed struct, and a rollout belongs to a
  thread by the thread id its file name carries (exact), not by a substring
- rules.go loses its shell regexes: a command line is parsed with mvdan.cc/sh/v3/syntax
  (shell.go) to unwrap `/bin/<shell> -c[l] <cmd>` and to recognise the replayed `sleep N` job;
  the ps listing parser stays (ps.go) and has unit tests
- TestAllowlistOnlyShrinks (CI job "replay allowlist only shrinks", ALLOWLIST_BASE=origin/main)
  fails when notReplaying gains an entry compared with the base branch

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* adr fail-fast-unimplemented: a mock refuses what it does not implement; the codex mock does

The codex mock refuses every flag it implements nothing of (--enable/--disable, -o, --output-schema,
--ephemeral, --sandbox, --profile, --color, --thread-source, --ignore-*, --strict-config,
--approve-for-me) and any -c key but agents.max_depth, with an error that names it; the
noninteractive-run deviations say so; claude-mock/run_flags.go is the ADR's one legacy exception
(the claude and cursor mocks follow).

Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* adr fail-fast-unimplemented: keep only what the user decided, and name the code it governs

Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* adr replay-exceptions-only-shrink: the replay exception list may only shrink

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* replay-exceptions-only-shrink: a file-guard enforces the ADR, replacing the CI test

The file-guard compares the notReplaying map's keys at the base and head and refuses an added one.
The ADR names the map and the guard; the Go test and its CI job go (one mechanism).

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* replay-exceptions-only-shrink: sr-test cases, and an emptied list is a legitimate shrink

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* tests: the refusal cases assert their reason; every refused flag is covered, --full-auto too

Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical
Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* replay-exceptions-only-shrink tests: each case proves both the permit and the refusal, with the reason

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* tests assert the rule's events and reason directly; noninteractive-run names stdin-as-prompt and resume --last as not modeled

Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical
Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* replay-exceptions-only-shrink tests: show the recovery after the refusal

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* codex-mock: a test that `exec resume --last` is refused and nothing runs

Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* modules: one home per file: the command line and the output frames leave internal/scenario

internal/cli owns each mock's entry points and flag refusal, internal/frames the output event
frames; internal/scenario keeps driving the script and routing records.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* Revert "modules: one home per file: the command line and the output frames leave internal/scenario"

The split made module-leaks refuse (the A10N_MOCK_* protocol read outside the scenario home) and
module-distinct still named other files; the module layout is fixed in its own PR (issue filed).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ngs must be clean (#142)

* Judges list every failure; deviations need user words; recordings must be clean

- judge templates: a fail verdict lists EVERY failing item, one numbered line each
- capability-grounded: user words for a new deviation or a newly unsupported cell;
  undisclosed doc/recording conflict fails; a deviation negating the cell's defining
  clause means supported: false (ADR capability-grounding updated)
- snapshots-current + capture.sh: unparsed hook-log lines, harness error lines/frames and a
  mode the mock does not imitate refuse a recording (_lib/recording.sh, expected.yaml declares
  legitimate ones)
- tests/rules.sh covers all of it; run in CI

Sloprail-Cites-User: #142 (review) <-- fix

Sloprail-Cites-User: (c) "When a doc and a recording disagree, the cell must say so as a deviation."
Sloprail-Cites-User: (d) "A deviation that contradicts what the cell claims means the cell is not supported."

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #142: drop the every-failure paragraph and rules.sh, quote only for what the mock does not model

- the "Every failure, in one verdict" section leaves the judge templates: the engine
  keeps only the first line of a multi-line fail reasoning (judge.go reasonFromVerifierOutput),
  so asking for a list would not reach the agent
- tests/rules.sh and its CI step go; the tests move to sr-test cases in a later PR
- added-or-removed.sh: a user quote is required only for a new modeled-surface deviation
  (the mock lacks what the harness does); a deviation or supported:false grounded by a
  recording that shows the harness lacking it needs none (ADR capability-grounding)

Sloprail-Cites-User: #142 (review) <-- fix

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #142: drop the recording checks (_lib/recording.sh, expected.yaml, capture.sh hooks)

Recordings are trusted as recorded: snapshots-current and the three capture.sh go back to
what main has.

Sloprail-Cites-User: #142 (review) <-- fix

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #142: deviations carry a kind; user words only for what the mock does not model

- capability.cue: a deviation has an optional `kind`: mock-not-modeled | harness-lacks
- capability-grounded/added-or-removed.sh (deterministic): words are required when a
  mock-not-modeled deviation (or one with no kind) is added or changed, or a cell turns
  supported: false without a cited recording; a harness-lacks deviation and a recorded
  unsupported cell need none
- ADR capability-grounding and the rule's comments describe it

Sloprail-Cites-User: #142 (review) <-- fix

Sloprail-Cites-User: (b) "A deviation where the mock doesn't model something the harness does needs my quote; one backed by a recording of the harness doesn't."

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* chore(spec): give each deviation its kind (migration)

adr modeled-surface -> kind: mock-not-modeled (97 less 3 unsure), adr capability-grounding ->
kind: harness-lacks (47 less 2 unsure). Five entries are left without a kind for the user to
decide (the mock differs from the harness in part, or what the entry names is unclear):
hooks-all-matching-run cursor[0], session-transcript-file codex[1], subagent-transcripts
codex[0], codex[1], cursor[1]. Content otherwise unchanged.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #142: a kind: harness-lacks deviation must cite a doc or a run

harness-lacks-cited.sh (a script check ahead of the judge in capability-grounded) refuses a
cell with a harness-lacks deviation that cites neither docs nor runs: that deviation is waived
from the user's words because a cited doc or recording shows it (ADR capability-grounding).

Sloprail-Cites-User: #142 (review) <-- fix
Sloprail-Cites-User: (b) "A deviation where the mock doesn't model something the harness does needs my quote; one backed by a recording of the harness doesn't."

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* chore(spec): classify the five deviations left without a kind

hooks-all-matching-run cursor[0], subagent-transcripts codex[0] and codex[1]: mock-not-modeled;
subagent-transcripts cursor[1]: harness-lacks; session-transcript-file codex[1] is dropped (the
mock does the same as the harness, so it is not a deviation). The user decided each.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #142: kind is required on every deviation

capability.cue: `kind` is required (every deviation is classified), ADR capability-grounding says so.

Sloprail-Cites-User: Every deviation says whether the mock doesn't model it or the harness lacks it
Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): cells whose deviation says the harness lacks the defining clause are unsupported

Rule (d): codex agent-input-validation, bash-tool-result, compaction-transcript-continuity,
task-stream-frames, transcript-record-envelope; cursor plugin-hooks, task-stream-frames become
supported: false, each keeping the doc and recorded runs that show the harness lacking it. The
harness-lacks deviation becomes the reason; the mock-not-modeled ones go (nothing is mocked).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): the remaining cells #142's rules flag

- (d) cursor agent-input-validation and foreground-subagent-bash-ends-with-response are
  supported: false (their docs and runs kept; the harness-lacks text is the reason)
- (c) cursor noninteractive-run and print-waits-for-background-agents, codex
  session-transcript-file disclose the doc-versus-recording conflict
- codex transcript-record-envelope: the reason names turn_context's cwd and the item events' thread id
- cursor hooks-all-matching-run[0]: kind harness-lacks (Cursor itself runs the hook once per source)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): codex session-transcript-file is unsupported (rule (d))

Both of its deviations negate a clause of the statement (keyed by working directory; the file
does not exist when the start hook runs), so the cell is supported: false, with its doc and
recorded run.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* codex-mock, cursor-mock: markers of cells that became unsupported go to the supported cell the code serves

- cursor background.go notificationFrame and its test: task-notifications/cursor
- cursor agent_task.go keeps foreground-subagent-result/cursor; subagent_shell_test: task-notifications/cursor
- codex collab.go and its test: foreground-subagent-result/codex

Comments only; no behaviour change.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #142: a cell may carry notes, each harness's own way of providing the capability

capability.cue: optional `notes` on a providers cell. A statement stays abstract; where a harness
does the feature in its own way (a location, a timing), that is a note, not a deviation.

Sloprail-Cites-User: maybe statement should be more abstract and deviated per harness?
Sloprail-Cites-User: each hartness just has its own transcript resolution - otherwise how we'd strucutre transcript frature?

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): abstract statements, each harness's own way as notes; ten cells supported again

agent-input-validation, bash-tool-result, compaction-transcript-continuity, task-stream-frames,
transcript-record-envelope, plugin-hooks, foreground-subagent-bash-ends-with-response and
session-transcript-file name no single harness's mechanism now; where a harness does the feature
its own way (codex agent-input-validation, bash-tool-result, compaction-transcript-continuity,
task-stream-frames, transcript-record-envelope, session-transcript-file; cursor agent-input-
validation, task-stream-frames, plugin-hooks, foreground-subagent-bash-ends-with-response) the
cell says so in `notes`, and the mock-not-modeled deviations dropped earlier are back. Stay
unsupported: cursor compaction-transcript-continuity and transcript-record-envelope, codex
foreground-subagent-bash-ends-with-response (they lack the feature).

Sloprail-Cites-User: maybe statement should be more abstract and deviated per harness?
Sloprail-Cites-User: each hartness just has its own transcript resolution - otherwise how we'd strucutre transcript frature?

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…#176)

Sloprail-Cites-User: and why it introduced notes btw? this is an extra level of complexity and N cases to jduge etc


Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ported again (#186)

#142 moved these markers off task-stream-frames (codex, cursor) and
foreground-subagent-bash-ends-with-response (cursor) while the cells were
unsupported; the cells are supported again, so capability-covered flagged them
as missing markers. Comment lines only, no behaviour change.


Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ng and refuses nothing (#167)

* feat(snapshots): docs follow recordings; a doc change re-judges nothing and refuses nothing

No check reads a live doc page; capture.sh re-freezes the docs its cells cite with a recording; capability-grounded/rigor keys carry recordings, never doc hashes; doc_copy_var fixes DOC_ERROR: unbound variable; sr-test cases for the four rules.

Sloprail-Cites-User: only recording changes matter - doc changes don't - bc recording changes when changed then docs are pulled

Sloprail-Cites-User: but why the fuck we need this drift job? we only pull docs when needed ad-hoc or

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(snapshots-read-only): assert the one refusal, no refused event for the recovery and boundary

Sloprail-Cites-User: only recording changes matter - doc changes don't - bc recording changes when changed then docs are pulled

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…158) (#184)

* claude: replay every recording through the mock's own `replay` command (adapter + generated test)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: split adapter and load by responsibility (file-size)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: review fixes (explicit assistant drops, ids found by the adapter not the core, refuse two sayings, specific reasons)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* ci: run the mock tests three times to surface flaky ones (#151)

Unit packages run with -race -count=3 (COUNT in each mock's Makefile). The e2e suites take
minutes, so two extra replicas per mock run in parallel next to the required e2e job instead
of -count=3 in one job: three runs under CI load, no extra wall time.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* ci: keep the repeat in the workflow, not the mock Makefiles

A change under *-mock/ re-runs the whole-tree capability-covered guard, which is red on main
for unrelated missing provides/proves markers. The unit repeat now calls go test directly and
the e2e replicas run the Makefile's commands themselves.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…_use_result, wire_tool_inputs (#158) (#192)

* claude-mock: stream frames carry session_id, parent_tool_use_id, tool_use_result and wire_tool_inputs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: stream frame stamping in its own file (file-size)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: stream frame stamping lives with the records (module home)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…m (part of #161) (#195)

* turnloop: a turn may have several tool calls; cursor-mock starts all before completing any

scenario.Turn reads the tool_use lines of one turn (Tools; Tool stays the first);
turnloop starts them all before completing any for a host that opts in through
turnloop.Interleaver. claude-mock and codex-mock do not opt in and carry out the
first call as before. cursor-mock opts in: shell started, Task started, shell
completed, Task completed (runs/task-stream-frames).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* agent_task: reflow the finishSubagent comment

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
#177)

The mock took 0 as no yield and ran the command to its end; Codex returns
the receipt at once, whatever the command does (recorded: task-notifications-bg).
The replay found it; task-notifications-bg leaves the exception list.


Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…161) (#196)

* cursor-mock: the thinking-frames test is not vacuous; a positive control shows the mock prints a script's assistant text (part of #161)

The test set assistant frames aside from the replay, whose scripts carry no text,
so "no assistant frames" held whatever the mock did. It now asserts the
recording has assistant frames, the replay has no thinking frames, and a new
test shows a script's text is an assistant frame in the stream.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: a test pins the tool_call frame bodies the force-write recording shows (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… wording, drifted-page cache, escaped path, trailing-slash lookup, tests) (#191)

Header and judge prompts no longer say a doc is verified against its freeze; doc_section/slug removed; a drifted page is cached once; trailing-slash URLs found; all-wiring and slash tests; the capability-grounded add-e case asserts the exact reason and recovers.

Sloprail-Cites-User: Apply the #167 reviewer's follow-ups: fix the stale snapshots header, dead helpers, judge wording about live docs, re-fetching drifted pages, the unescaped regex, and the trailing-slash doc lookup, with tests.


Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ing (#203)

* cursor noninteractive-run: the mock-not-modeled deviation claimed the mock emits no assistant frames; it prints them for script text, and emits no thinking frames (part of #161)

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor noninteractive-run: the mock stream lacks the recorded assistant frames too, not only thinking frames (part of #161)

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…egexes (#170)

* codex replay: read the model's JavaScript with a parser, not regexes

The adapter parses each exec call's script (goja's parser) and evaluates what decides
the tool calls: literals, constants, objects, arrays, templates, store/load and the
receipts of earlier calls. A call that depends on a result, or JS the adapter does not
read, makes the recording unbuildable instead of being guessed. user-prompt-submit-hook-bg
now replays green and leaves the exception list.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #170: scope names, refuse tool calls inside functions and aliases, read bracket keys

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* review of #170 (2): refuse var, tagged templates, defaulted params, mutating methods, exponent numbers, dropped arguments

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #170 (3): receiver before arguments, optional chains, dead zone, lone surrogates, whole-number yields, a spawn with no arguments

agent-input-validation replays green and leaves the exception list.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #170 (4): refuse reads of null, a spawn argument that is not an object

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #170 (5): a binary operator's right side runs unless it is ||, && or ??; the base of a ?. always runs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #170 (6): a ?. on a known non-null value does not short-circuit; tests for the chain's end and operands

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #170 (CI judge): the harness's other tool options go to the mock as given, not dropped by the adapter

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #170 (8): refuse non-finite options and another working directory; tests for the pass-through

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #170 (9): sub-agent scripts get the run's directory, a workdir that is not text is refused

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…tch both ways (#153) (#189)

* feat(rules): capability-reconciled, cells and sr:provides adapters match both ways (#153)

A deterministic file-guard scoped to the (capability, harness) pairs a change
touches: a supported cell needs its sr:provides code and a cited recording that
exists; a sr:provides marker needs an existing, supported cell. Replaces the
sr:provides halves of capability-covered (whole tree, forward and back).

Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): rigorous cases for capability-reconciled; cases for capability-covered

Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): capability-reconciled covers a supported cell citing no run

Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(capcells): the sr:provides half moved to capability-reconciled; capability-covered tests assert sr:proves

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-reconciled sees a cited recording deleted or replaced

Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* docs(adr): capability-once names the cell's deviations field and links capability-reconciled

Sloprail-Cites-User: The capability-once ADR names the deviations field and links the capability-reconciled rule.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* docs(adr): capability-once says each deviation names an ADR that exists

Sloprail-Cites-User: each deviation names its ADR, and that ADR must exist
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): capability-reconciled cases assert the refusal's own reason text

Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-covered re-runs when an ADR is deleted or renamed, with a case

Sloprail-Cites-User: each deviation names its ADR, and that ADR must exist
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-reconciled fails closed when the touched pairs cannot be worked out; the catalog goes through a file, not argv

Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… limited in width (session-end flake cause) (#201)

* codex replay: the scratch repository, the machine's text, the scratch root, at most a few replays at once

The recording's scratch repository has one empty init commit; its commit id, the harness's npm package root and the scratch directory (<TMP> is the root the repository, CODEX_HOME and TMPDIR sit in) are not behaviour; sub-agents' scripts sit beside the repository, as the real run's has none. Hooks run under wall-clock limits (SessionEnd: one second by default), so replays are limited in how many run at once: sixteen at once killed a recorded half-second hook, one alone never did.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* replay width is not decided here: the session-end entries stay, as flaky with the cause found

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* tests: sub-agent scripts are Scripts after the merge with main

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adapter.go under the size limit: the placeholders in a file of their own

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…son cites the recording (#171)

* spec(compaction-transcript-continuity): cursor stays unsupported, reason cites the recording

Part of #161.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(compaction-transcript-continuity): reason names the turn_ended line and where dynamic_tools sits

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ed (#169)

* test(cursor): plugin-hooks tests stop pinning an order the recording does not establish (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor): plugin-hooks comments say why no order is pinned (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor): a loaded plugin's hooks run before the project's, as recorded; tests pin it (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor): a manifest-named hooks file replaces hooks/hooks.json (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor): label the manifest-replaces-default assertion as doc-based (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor): plugin-hooks wording says what is recorded, what is configured order, what is doc (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…nt's next call (#174)

* fix(cursor-mock): the aborted notice of a sub-agent's command follows the parent's next tool call, as recorded

Part of #161. The test now asserts the whole recorded frame order unfiltered.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* refactor(core): the deferred report of what ends at a sub-agent's response is the core's, cursor sets when

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(core): one sr:capability marker for foreground-subagent-bash-ends-with-response

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor): foreground-subagent-result asserts the recorded args, hook and PINEAPPLE-7 result; the stream says unspecified for a generalPurpose Task

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* refactor(core): the deferred-report helper lives beside, not in, the sr:capability file

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…on-end flaky entries leave the list (#206)

* replays run with limited concurrency: an ADR and the core's gate, both mocks' replay tests use it

The two session-end entries (flaky: load) replay green under the limit and leave the exception list.

Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* replay concurrency: the core's Run holds a slot; the ADR decides only what the user said

Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* ADR replay-concurrency: names where the limit is held, no history

Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* ADR replay-concurrency: the user's words, with the slot count and where it is held

Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load: the replay core's Run (internal/replay) holds one of at most min(max(NumCPU/2, 1), 4) process-wide slots for each replay, for every mock.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…213)

* grounded judges: a short user reply grounds exactly what it answered

The adr-grounded, invariant-grounded and capability-grounded judges are handed each cited quote with its
transcript path:line and may Read. Their rubric now says how to read a short reply ("lgtm", "yes",
"1. lgtm"): the user's quote is the only authority, the surrounding messages are context to interpret it,
a short approval grounds exactly what it answered and nothing beyond, the approved proposal must itself state
the change, an earlier answered proposal does not carry over, an assistant message alone grounds nothing, and
what is read from the record is data (same rubric as sloprail's grounded-rule-changes, sloprail #264).
Two sr-test cases per rule: "lgtm" grounds the change proposed a few messages earlier; the same "lgtm" does not
ground an unrelated change (refused with the judge's reason).

Sloprail-Cites-User: They just get the line number of JSON error plus my quote, and then they can just go and read the JSON error to figure things out.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* short-approval cases: the refused unrelated change recovers with the proposed one

The unrelated-change cases now continue after the refusal: the agent drops the unrelated commit, makes the
change that was proposed, cites the same "lgtm", and the rule permits it (refused, then passed).

Sloprail-Cites-User: They just get the line number of JSON error plus my quote, and then they can just go and read the JSON error to figure things out.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* short-approval cases: strict is_error check, mock-judge disclaimer, transcript noted beside allowed_tools

Sloprail-Cites-User: They just get the line number of JSON error plus my quote, and then they can just go and read the JSON error to figure things out.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… make (#231)

* toolspec: one validator for the tool calls a scenario script asks a mock to make (ADR tool-calls-validated)

The core's schema and check: an unknown tool, an unknown parameter, a missing required one, a wrong type or an option value the mock does not implement is refused, naming the tool and the parameter; a kind the real harness answers itself is returned for the mock to answer as recorded. Ungrounded checks a schema against recorded calls.

Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* toolspec: the claude, codex and cursor runners validate every tool call a script asks for, against a schema each declares

The shared turn loop (codex, cursor) and claude's stream scan check each call before it is played; a refusal in a sub-agent's script fails the whole run. Schemas are grounded in the recordings (a test per harness), the kinds a harness answers itself are listed with their recorded run, and parameters the mock implements that no run shows are grounded in the docs. ToolSearch, recorded, is implemented.

Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* toolspec: the shared turn loop validates through the host (turnloop.Validator), a refused call is kept for the run; the ADR says only what the user asked; files split to the size limit

Only the refusal of a call is kept and escalated, the same in the three mocks; the runners' loop and run files are no longer touched.

Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex schema: spawn_agent's task_name and fork_turns, recorded in a function call, are declared

Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* toolspec: codex is wired by the codex batch with the options it implements; the ADR keeps only the user's decision

The transitional codex schema accepted exec_command options the mock does not read, which adr/fail-fast-unimplemented forbids; removing them would refuse every recorded call. codex-mock does not validate until batch/codex declares its exec options on the shared schema (turnloop.Validator on turnHost and subHost).

Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr(tool-calls-validated): name the mocks, the schema package and what a failing call does

Sloprail-Cites-User: Every mock (claude, codex, cursor) checks each tool call its scenario script asks for against that harness's recorded tool schema (internal/toolspec) before playing it, in replays and ordinary runs; a call that fails the check fails the run, except where a recording shows the real harness answering that call itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* toolspec review: Bash timeout leaves the claude schema (no recording grounds its answer), the answered-kind tests check the recorded call and answer, cursor's refusals belong to the run's main session, a refused claude run streams no result

Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ems (#230)

* codex replay: hook payloads are sorted only within a group of concurrent hooks, the groups keep their order (audit #207 item 3)

A continued Stop follows the Stop it continues (stop_hook_active is in the group key).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* replay: per-run values keep their key behind a placeholder instead of being dropped (audit #207 item 4)

Core Rules.MaskKeys: the key stays, its value is <key>; the codex rules mask usage, model, the transcript paths, ids and timings, and only the mock's own script parameter is dropped.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: spawn_agent answers at once, add wait_agent (replay-driven)

The mock's spawn_agent always answers {agent_id, nickname}; the sub-agent runs as a background task on the session's task registry and wait_agent (hooked multi_agent_v1wait_agent) waits on it through the core's subagents.Wait (first target to finish, statuses completed/running/not_found, an empty status on timeout). The adapter maps tools.multi_agent_v1__wait_agent to a unified ToolWait; the replay's scratch repository is on branch main; the old background parameter is refused. Six recordings leave the exception list.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the mock-not-modeled deviations #187 makes false are removed or corrected

The mock's spawn_agent answers at once and wait_agent is modelled, as the recordings show (foreground-subagent-result, nested-subagents-nowait, task-stream-frames, background-agent): the deviations saying the mock waits inside spawn_agent, has no wait tool or takes a background parameter are removed or reworded.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: foreground-subagent-result's statement is harness-neutral

The dispatching agent receives the foreground sub-agent's final result: Codex provides it with spawn_agent and wait_agent, so the cell stays supported; the Claude trailer wording leaves the statement.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: task-stream-frames/codex discloses the recorded gap, no frame for a background task

A harness-lacks deviation backed by runs/task-stream-frames (item_4: the shell command asked to run detached is an ordinary command_execution).

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex replay: the model's answers are steps of the script, played in recorded order

A turn a Stop hook continues is followed by the model's next steps; the unified agent keeps Final (the claude adapter reads it) and answers among the calls are ToolAnswer steps. stops leaves the exception list; hook-exit-codes keeps its entry (it is green on macOS and red on Linux: its recorded command is ls of a missing path, whose message and exit code are the host's).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: an exec_command option the mock does not implement is refused, not ignored (audit #207 item 6)

A working directory other than the run's, and a command whose output is longer than max_output_tokens (Codex truncates what the model is told; the mock does not), are refused loudly. A tty is carried out as far as the recordings show (lines end in CRLF); shell and login only choose the shell, which the mock does not model (every scenario command is run by /bin/sh).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: the hand-written replay test goes; the generated replay is what hook-exit-code-semantics and pretooluse-refusal are proved by (audit #207 item 9)

TestReplayOfRecordedRuns replayed stops and hook-exit-codes with its own normalisation, deleting fields and taking its commands from the recording, while the generated replay (which compares the whole stream and payloads) listed them as failing.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* replay: compare with every sample; two empty event streams are not a green replay (audit #207 item 12)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex replay: sub-agent rollouts are matched by thread id, leftovers and unanswered scripts are refused (audit #207 item 5)

A second commentary before one call (the first was silently overwritten) is refused; each spawn_agent is matched to its sub-agent's rollout by the receiver thread id its stream item names (not by position) and a rollout no spawn names is refused; a script with no recorded output, an output of no script, and a failed script that made several calls (which ran is not recorded) are refused. A second final answer is not an overwrite any more: answers are steps.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex replay: a method call on a value the adapter does not follow is refused unless it only looks around; wait targets are filled in structurally (audit #207 item 11)

The evaluator accepted any call on an opaque receiver; only the methods the recorded scripts use to look at what they were given (filter, test, stringify, includes, toLowerCase) are allowed. The generated script no longer rewrites its own text with sed (a literal AGENT1 or IDPLACE in a recorded string was hit): wait targets are {spawned: k} objects that jq replaces by the k-th spawn receipt, and the call id is set by jq; the limit of nine spawns is gone.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* wait_agent: a background command of the session is not a sub-agent to wait for

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: an exec_command shell other than zsh is refused (audit #207 item 6)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: an exec_command runs by the shell it names, zsh -c or zsh -lc, and a missing zsh is refused (audit #207 item 6)

The mock ran every command by /bin/sh whatever shell and login the call named. It now runs it by zsh (-c, or -lc for login:true) as the harness does; zsh not being installed, or a login with no shell, is refused loudly, never replaced by /bin/sh; a shell other than zsh is still refused. The codex e2e CI jobs install zsh.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* tests: the zsh e2e test fails when zsh is absent instead of skipping

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* module map: exec_options.go with the tools; spawn receipts are parsed in the subagents module (module-coverage, module-leaks)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the codex mock-not-modeled deviations the recordings and the mock prove false are corrected

The mock fires SessionEnd, SubagentStart, SubagentStop, PreCompact and PostCompact, and reads a SubagentStop hook's block/exit 2 (runs/subagent-stop-block-loop): the hook-exit-code-semantics and hook-matcher-filter deviations that said otherwise are corrected, the subagent-lifecycle-hooks one removed.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* ADR tests-fail-on-missing-tool: a test whose required tool is missing fails, it never skips

Sloprail-Cites-User: Fail, never skip

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* tests: a missing required tool fails the test (capcells jq/yq/git/bash, python3, the real claude when asked for); CI installs python3, jq and yq

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: hook-exit-codes re-recorded with a portable failing command; its exception entry goes

Only the ls of a missing path (whose message and exit code are the host's: macOS 1, GNU 2) is replaced, by false; every other command, hook and step of the scenario is unchanged. Recorded with the pinned codex 0.159.3 (capture.sh); the old sample is dropped. The capture re-froze the plugins doc page's hash in MANIFEST. The replay is green on any host.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #230: wait_agent refuses bad input, a failed sub-agent start is errored not completed, the script fails loudly on a missing receipt, reasons and comments

wait_agent decodes strictly (an unknown field or a wrong type is a failed result naming it); a sub-agent whose rollout could not be created is told as errored with why; the generated script errors when a wait's spawn position has no receipt; three notReplaying reasons name wait_agent; the exec_options and wait comments are reflowed and true.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: noninteractive-run-output-schema re-recorded without --output-schema; its exception entry goes

The mock refuses --output-schema (it implements none of it), so the recording is the plain run: no args, no schema.json, recorded with the pinned codex 0.159.3 (capture.sh), the old sample dropped. The cell's mock-not-modeled deviation no longer cites the run for a schema-shaped message it no longer shows. The replay adapter writes no hooks.json for a run with no hooks (an empty file is not a hooks file).

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: subagent-transcripts-v2 re-recorded in the default multi-agent mode; its exception entry goes

The mock refuses --enable multi_agent_v2, so the run is recorded without it (args removed, the prompt asks for the sub-agent and wait tools; the name is kept: runs are not renamed). The sub-agent's path in the tree of agents is null in the default mode, as the mock leaves it, and the final answer reaches the session also as a subagent_notification message, which the mock does not record: the cell's two deviations are corrected to the recording and the test follows. Recorded with the pinned codex 0.159.3 (capture.sh).

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: the session is told when a sub-agent it started has ended, as the recording shows

A <subagent_notification> user message (the sub-agent's id and how it ended: its final answer, or why it failed) goes into the owner's rollout after the tool output it was just given, once, whether or not the agent waited for it (recorded: runs/subagent-transcripts-v2). The cell's deviation for the gap is removed, the mock having closed it; a test compares the notification, its place and its content with the recording's.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: noninteractive-run-output-schema is restored as recorded, with its exception entry and its spec sentence

The re-recording without --output-schema lost the recorded schema behaviour (the real run's message conforms to it), which the cell cites. The original run, its setup (args, schema.json) and sample are back, the plain re-recording's sample dropped; the replay adapter keeps writing no hooks.json for a run with none.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* review of #230: a failed sub-agent start fails the wait fast, no unrecorded errored status; a runner test; the v2 name's history noted

The core wait no longer has an errored status (no recording shows its shape): a wait that names a sub-agent which ended without an answer returns subagents.AgentFailed, the codex wait_agent fails with a clear error saying the mock refuses it rather than inventing it, the failed start is also reported on stderr, and no notification is written for it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* ADR tests-fail-on-missing-tool: only what the user's words decide (a missing tool fails the test, never skips)

Sloprail-Cites-User: Fail, never skip

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: subagent-transcripts/codex: the missing sidecar is the harness's lack, not the mock's

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* sub-agent end notices: which sub-agents and when is the core's (subagents.TakeNotices), the codex adapter only encodes, marked sr:provides

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: a wait for several sub-agents, recorded: it returns at the first, the others are told of one at a time (runs/foreground-subagent-wait-many)

The recording shows what the doc does not: a wait naming two sub-agents returns at the first to finish, lists only it (not a target still running, and the stream's completed wait item names only it), and the second is told of by a notification of its own after the agent's answer, which goes on. The mock now: lists only finished targets; tells the session of one ended sub-agent at a time, at the point the model is asked again (not between the calls of one script, the adapter marking those with the mock's own parameter more) or when the turn would end, which then goes on with it (a running sub-agent gets answerLatency to end first, one just started takes modelCall to start: the mock has no model to take that time); a sub-agent the recording shows unfinished hangs instead of ending; the replay evaluator reads const [a, b] = await Promise.all([...]). The harness-lacks deviation for the doc's all-wait is added to the cell.

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: the SubagentStop variants (continue:false over a block, plain text ignored) recorded and tested; the SessionStart matcher is tested on the resume and fork sources

Two new runs recorded with the pinned codex 0.159.3 (subagent-stop-continue-false: of two SubagentStop hooks one blocks and one says continue:false, both ran and the sub-agent was not run again; subagent-stop-plain-text: plain stdout on exit 0 is ignored); both already replay green, so the mock models them, and tests compare their hook logs with the recordings. The block-loop test is also marked as proving the exit codes. A test drives per-source SessionStart matchers on a fresh, a resumed and a forked session (doc-grounded: hooks#sessionstart). The hook-exit-code-semantics cell lists the three SubagentStop runs.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* the tests-fail-on-missing-tool ADR leaves the batch (it stays on side/tests-fail-on-missing-tool); the skip-to-fail conversions and the CI tool installs stay

Sloprail-Cites-Tool: no script or other rule enforces it. Either widen the rule's match to include test files (and any non-Go test scripts) or link a rule that checks tests for skip-on-missing-tool.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(scenario): a turn may carry a gate the loop holds it back for

An assistant line's "gate" names what must have happened before the turn's call is
taken, so a script orders agents by events, not by delays. turnloop calls the host's
Gater between the turn and its call.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(codex-mock): order agents by script gates, not by invented durations

The 1s model-call and 4s answer-latency delays are gone. The replay adapter derives each
step's gate from the recording's times; the mock counts calls started (hooks fired) and
finished per agent and releases a gate by the event it names.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(turnloop): the loop decides when the agent is told of what ended

A host may be a Noticer: the loop tells it after the last call of a model's script (a call's More
marks that another follows), and instead of ending the turn, with no Stop hook run and no block
counted. Host interfaces move to host.go.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* refactor(codex-mock): the core decides when a sub-agent's end is told

The adapter keeps only the notification's text and the user message it is recorded as; the
Stop hook's stop_hook_active comes from the loop's continuing. A call's `more` is a field of
the call, not of its input.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): refusing apply_patch, ask/allow leaving no record, and compaction/sub-agent matchers

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the wait-many deviation names the doc section and quotes its sentence

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* revert: the mock-not-modeled corrections wait for the user's words (patch handed back)

Sloprail-Cites-Tool: file-guards evaluated in 15m7.673s (31 rules, concurrency 6)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec(codex): narrow three mock-not-modeled deviations the recordings prove false

The codex mock fires PreCompact, PostCompact, SubagentStart, SubagentStop and SessionEnd, matches on the compaction trigger and sub-agent type (runs/subagent-stop-*), and refuses apply_patch by a PreToolUse hook; the deviations now name only what is still not modeled.

Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: transcript_path on every event, tool hooks and SessionEnd included, recorded (runs/session-transcript-file-hooks) and tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: session-transcript-file/codex lists the hooks recording and drops its kind-less deviation

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* refactor: the agents' gate ordering (progress, spawn log, hold) lives in internal/subagents

The codex adapter only reports a gate's mistakes in its own words. progress.go joins the
subagents module's home.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): several hooks' added context, SessionStart context after a compaction, PostCompact matcher marker

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: a PostToolUse block with added context, recorded (runs/posttool-block-context); context on a resumed session tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-additional-context/codex lists the block-with-context recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(replay): the codex adapter replays later runs of the harness (then-NN resume and fork)

A recording's later runs are read into core.Recording.Then: each run's records are cut from its
thread's rollout by its own prompt (a resumed thread holds the runs one after the other, a forked
one a copy of the history first). The mock is run once per run against the same CODEX_HOME, from
the run's directory, with <SESSION> the first run's session, and each run's script starts after
the steps its session already holds. session-resume and session-fork leave notReplaying.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: SessionStart context on a resumed session, recorded (runs/session-resume-context) and tested against the recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-additional-context/codex lists the resumed-session recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(replay): the codex adapter replays a session the harness compacted on its own

A compacted record becomes a compact step of the script (the mock has no tokens: the run's token-limit
option is the one setup args installed), and a compaction spends a step. manual-compaction-auto and
session-start-compact-continue-false leave notReplaying; the two runs a hook stopped keep their entry with
the reason now true.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: context after a mid-turn compaction and a /clear prompt, recorded (runs/manual-compaction-auto-context, runs/session-start-clear-prompt) and tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-additional-context/codex lists the compaction and /clear recordings; source clear is the harness's lack

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(hooks): hooks run in the background (Later) deliver their context at the next safe point

The agent does not wait for a hook marked async. Background hooks run one after another in the
order started; what they printed is due once the agent has gone through the call after the one it had
begun when they started, or when its turn would end, and delivery waits for the hook to finish: an event,
not a time. turnloop's Noticer returns every notice that is due.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: async hooks deliver late, login:false with no shell is the default shell

SessionStart, UserPromptSubmit and PostToolUse hooks marked async run in the background and add their
context when it is due (recorded: runs/hook-async-context); their logs land in whatever order they finish,
which the replay does not compare. login:false with no shell is accepted, as the real harness accepts it
(recorded: runs/exec-login-false); login:true with no shell stays refused.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: async hooks' late context recorded twice (runs/hook-async-context, runs/exec-login-false) and tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(hooks): one sr:capability marker for hook-additional-context

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): apply_patch matcher aliases and the SessionEnd matcher by the binary; sub-agent stop variants marked for the lifecycle capability

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: bash-tool-result/codex lists the login:false recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* refactor(hooks): one core function (ContextDue) decides when a hook's context reaches the agent

The sync path (a hook the agent waits for) and the background path (Later.Due) both ask it, and the one
sr:capability hook-additional-context marker is on it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): a failed command with no output is told as an empty output, and PostToolUse names its PreToolUse's tool_use_id

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: a command that ran to its end is told framed, as recorded; a user's SIGINT interrupts the turn

The agent is told "Script completed / Wall time / Output:" then the output, for a failure as for a
success (recorded: runs/shell-exit-status). SIGINT during a turn fires the Interrupt hook, answers the
running call "aborted by user after <time>", tells the agent the user interrupted, records the turn as
aborted, prints nothing more, fires no PostToolUse or Stop hook, still ends the session and exits 1
(recorded: runs/interrupt-hook). The replay adapter sends SIGINT once the command has started.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: an interrupted turn recorded (runs/interrupt-hook, capture.sh interrupt-after) and tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/codex lists the interrupt recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: Interrupt hook answers, PreToolUse exit-1 stderr and SessionEnd steering recorded and tested; the replay starts its child through procexec

runs/interrupt-hook-output (what an Interrupt hook answers cannot prevent the interruption or be shown), runs/session-end-steering (SessionEnd hooks cannot steer the run), and the PreToolUse exit-1 stderr case of runs/hook-unstartable. procexec.Spec gets Stdout and OnStart so the interrupted replay no longer calls exec.Command.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* revert: hook-matcher-filter's interrupt-hook run waits for the user's words (patch handed back)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec(codex): the mock fires Interrupt; cite the interrupt and session-end-steering runs

runs/interrupt-hook shows a SIGINT during a turn fires the Interrupt hook, and the mock now does the same, so the deviation saying it fires none is corrected; the new runs are listed in their cells.

Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the SessionEnd failure-report deviation begins Doc and recording conflict, names the section and quotes it

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(turnloop): interrupt.go joins the turn module's home

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* revert: the Interrupt work moves to side/codex-interrupt (it contradicts adr/modeled-surface)

Mock SIGINT handling, the Interrupt hook, the interrupted replay, runs/interrupt-hook and runs/interrupt-hook-output with their tests, capture.sh interrupt-after and procexec's Stdout/OnStart go with it. session-end-steering, the framing of a command's result and the other fixes stay. The deviation in hook-exit-code-semantics.yaml is handed back (patch outside the repo).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec(codex): Interrupt stays not modeled on this batch (its work is on side/codex-interrupt)

The Interrupt modeling moved to side/codex-interrupt pending the user's decision on adr/modeled-surface, so the deviation names Interrupt as not modeled again.

Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec(codex): cite the runs that prove the mock fires compaction, SubagentStart and SessionEnd

The narrowed exit-code deviation now lists the recordings that show those events firing (subagent-start-refused, manual-compaction-auto-blocked, manual-compaction-auto-post-stopped, session-end-hook-failure); the Interrupt wording is main's again.

Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(tasks): a gate waits for a sub-agent's task to be registered by an event, not by polling

Registry.Await wakes on Add; awaitEnd no longer spins on Find.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(scenario): a gate with a field it does not have is refused

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(codex replay): an exec_command option the mock does not implement is refused, not ignored

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test: the gate test no longer claims the turn loop; the SessionEnd failure test shows no report; the SubagentStart exit-2 test is marked for the exit-code capability

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: a sub-agent that ends with no message is null, SubagentStart's systemMessage is not shown, tool matchers on spawn_agent and wait_agent (recorded and tested)

runs/subagent-start-systemmessage and runs/subagent-stop-no-message; the replay reads trim as a look-around method; a sub-agent with an empty final message stops with last_assistant_message null and is completed with null in the wait. events/collab.go and filechange.go leave the scenario home for the subagents and tools modules.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: subagent-lifecycle-hooks/codex lists the systemMessage and no-message recordings

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: a call a hook rejects throws the hook's reason as a script error, framed as recorded

PreToolUse refusals and PostToolUse blocks are told "Script failed / Wall time / Output: / Script error: <reason>"
(runs/posttool-block, runs/pretool-decisions, runs/stops), whole, only the time masked.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): a PreToolUse hook refuses an exec_command with its options

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): the exec_command refusal and the resume/fork matcher tests are marked for the capabilities they prove

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): a SubagentStop matcher that matches the sub-agent's type runs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): the null last_assistant_message is present and explicit, recorded and mocked

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(codex-mock): the Interrupt work returns: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1

Brings back, from side/codex-interrupt (undoing de9ec6e): the mock's SIGINT handling, the Interrupt hook,
the interrupted replay with procexec's Stdout and OnStart, runs/interrupt-hook and runs/interrupt-hook-output with
their tests, and capture.sh interrupt-after. adr/modeled-surface carries the exception; the deviation in
hook-exit-code-semantics counts Interrupt as modeled.

Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): the exec option test does not depend on zsh being installed

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex-mock: an Interrupt hook over 3 seconds is clamped with the recorded warning; SubagentStop's continue:false is read by the hooks module

runs/interrupt-hook-timeout (the warning "clamping Interrupt hook timeout to 3s in <hooks file>"; hooks over their limit still
ran to their end). agent_stop.go goes: hooks.Interpret(SubagentStop, o).Halt decides.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-timeout/codex states the Interrupt timeout conflict between the docs and the recording

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr(modeled-surface): the abort bullet states the exception once, with the user's sentence, and no longer says a mock never interrupts

Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* codex: a SubagentStop hook's exit 2 beats its continue:false, recorded (runs/subagent-stop-exit2-continue-false) and tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-exit-code-semantics/codex lists the SubagentStop exit-2 recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(codex-mock): --ephemeral: no session file is kept and every hook payload says transcript_path null, replayed from the stream

The mock keeps the session in a scratch directory for the script and removes it; the replay adapter reads a run
that kept no rollout off its event stream (commands and one answer) and passes --ephemeral to the mock.
runs/ephemeral-no-transcript is the recording.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload and session-transcript-file (codex) list the ephemeral recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): the Interrupt clamp warning is asserted, and resuming a run by id is marked for the noninteractive capability

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr(modeled-surface): the Interrupt exception is its own modeled bullet, the user's sentence verbatim

Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): --ephemeral is implemented, so it is no longer among the refused flags

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* refactor(codex-mock): the ephemeral session's home and transcript live in ephemeral.go (files back under 150 lines); tests assert more

TestSubagentStartCannotRefuse asserts the exit-2 reason is surfaced nowhere; the ephemeral test is marked for noninteractive-run and the SessionEnd timeout test for hook-timeout.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr(modeled-surface): the Interrupt rule is its own decision, outside the out-of-model list

Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* revert: hook-exit-code-semantics' SubagentStop exit-2 run listing is re-added with the user's quote (patch handed back)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec(codex): drop deviations the recordings prove false (sub-agents, --ephemeral); cite the runs

runs/background-agent shows sub-agent payloads with an agent_id, runs/ephemeral-no-transcript shows --ephemeral modeled; the SubagentStop exit-2 run is listed again.

Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): an ephemeral run leaves nothing on disk: no sessions under the configuration directory and no scratch directory

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr(session-persistence): a mock persists a session exactly as the real harness does for the same flags

Sloprail-Cites-User: our system relies on persistent traj - so mock should work same way as real one, i suppose? unless both aaxctually persist it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: session-transcript-file/codex states the --ephemeral case: no rollout, transcript_path null on every event

Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): the background command's death is read from its process, the sub-agent payload test is marked for hook-common-payload, lifecycle_recorded_test.go is back at 400 lines

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr(modeled-surface): the abort out-list entry excepts codex-mock, and the Interrupt decision names its code

Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): a sub-agent's hook payloads carry its own transcript_path and the run's cwd, recorded and mocked

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr(modeled-surface): the abort out-list entry restates the user's rule (unless a recording drives it) and points to the Interrupt decision

Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): the SubagentStart/SubagentStop payload test is marked for hook-common-payload

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(codex-mock): the sub-agent start/stop payloads carry the parent session id and the run's directory, recorded and mocked

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* fix(codex-mock): --ephemeral=false is read as a value; the PostToolUse guard's two conditions are separate

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(codex-mock): validate a script's tool calls against a codex tool schema (adr/tool-calls-validated)

Needs batch/core (turnloop.Validator, internal/toolspec): rebase onto it, and add
codex-mock/internal/runner/toolschema.go and validate.go to internal/toolspec/module.yaml's home.

Schema: Bash (exec_command: command, workdir, yield_time_ms, max_output_tokens, shell zsh, login, tty),
spawn_agent (message; script is the mock's own), wait_agent, apply_patch, each parameter grounded in a
recorded run (replay.RolloutCalls reads the model's calls out of every rollout).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* toolspec module: the codex schema and validator wiring join its home

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* test(codex-mock): the old spawn background parameter and the test-only task_name are refused by the schema

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* fix(codex-mock): an ephemeral session keeps no sub-agent rollout either; run.go back at 150 lines

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…eleted (#233)

* feat(rules): snapshots-read-only lets a run directory that is not on main be deleted; only runs on main are read-only

Sloprail-Cites-User: A run directory that is not on main may be deleted; only runs already on main are read-only.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(rules): tests-fail-on-missing-tool: a Go test whose required external tool is missing fails; only an A10N_*_TEST gate may skip

Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): snapshots-read-only's not-on-main case asserts the gate's own permitted decision on the deletion

Sloprail-Cites-Tool: the permit of t1 (run not on main, which reaches refuse.sh and exits 0) is never asserted on a GateChecked event
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): snapshots-read-only refuses a Write to a recorded sample and lets setup/ and capture.sh be written

Sloprail-Cites-Tool: Add a case where the agent writes or edits a file under snapshots/ (refused, reason asserted), plus a permitted write to setup/ or capture.sh.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): the tests-fail-on-missing-tool ADR links only its own rule; the snapshots-read-only Write case is named for what it does

Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip.
Sloprail-Cites-Tool: That link therefore enforces nothing for this ADR and contradicts it
Sloprail-Cites-Tool: the folder name claims 'write-and-edit' but its agent.sh only issues Write tool calls
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(tools): skipcheck refuses a (*testing.T) skip in a function that reaches os/exec.LookPath, resolved by types; only an A10N_*_TEST gate is exempt

Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): tests-fail-on-missing-tool calls tools/skipcheck instead of matching Go text; a build or run failure is an error; cases for the gate-with-error, far-skip, alias, dot-import, method-value and helper bypasses

Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): each tests-fail-on-missing-tool case asserts the rule's own outcome on the events in its own test.sh

Sloprail-Cites-Tool: test.sh asserts no event of its owning rule tests-fail-on-missing-tool
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): skipcheck refuses every testing skip in a package that reaches os/exec.LookPath anywhere (variables, TestMain), checks root-level tests, and errors on a missing or empty directory; cases for the root test and the checker that cannot build

Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* docs(adr): tests-fail-on-missing-tool states that a test package that looks a tool up calls no skip outside the opt-in gate, which is what skipcheck enforces

Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip.
Sloprail-Cites-Tool: Checklist: (1) a test whose LookPath tool is missing fails and never t.Skips for that
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* docs(adr): tests-fail-on-missing-tool states the package-wide skip rule in the user's words

Sloprail-Cites-User: In a test package that looks up an external tool, no test may skip unless the skip sits behind an A10N_*_TEST gate.
Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): skipcheck judges a package with its _test package, counts interface and type-parameter skips, refuses a test that sets an A10N_*_TEST name; header documents what counts as a lookup

Sloprail-Cites-User: In a test package that looks up an external tool, no test may skip unless the skip sits behind an A10N_*_TEST gate.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* fix(tools): skipcheck drops the A10N Setenv check the ADR does not decide, and is split to stay under the file-size limit

Sloprail-Cites-User: In a test package that looks up an external tool, no test may skip unless the skip sits behind an A10N_*_TEST gate.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* test(rules): the snapshots-read-only cases assert the gate's actual reason, not the capture.sh pointer every refusal carries

Sloprail-Cites-Tool: none asserts the actual reason
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* cursor: replay every recording through the mock's own replay command (#158)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: review fixes (named gaps in the reasons, temp dir cleanup, empty hook log not green, print-script exit 2, multi-sample test)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor session-start-hook: cite runs/session-resume; drop proves markers from internals unit tests

The no-start-hook-on-resume claim rests on runs/session-resume, now in the cell's runs. The decision_test.go markers for pretooluse-refusal and hook-timeout sat on unit tests of Interpret; both cells keep their e2e proves.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: background-agent discloses the subagentStart/Stop doc and recording conflict; print-waits quotes the subagentStop sentence (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): hooks-all-matching-run/cursor cites the combined-refusal run and says the refusals are merged

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec: drop the cell note (notes field removed in #176)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): stop-block-cap/cursor cites the hooks doc's loop_limit default and discloses it conflicts with the print-mode recording (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): empty-tool-result-placeholder/cursor cites the exit-0 empty result

background-bash-start records ls exiting 0 with empty stdout: stream empty strings, hook {"output":"","exitCode":0}, no transcript tool result. Stays unsupported, with the sharper reason.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): manual-compaction/cursor cites the recorded auto compactions and says what they show

The cell cited only the headless /compress run. runs/compaction-transcript-continuity
records two real compactions with a preCompact payload (trigger auto, no hook after,
nothing stopping them); the hooks doc calls preCompact observational. Stays
unsupported: the request, the hook after and the stop are all absent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): manual-compaction/cursor says no compaction hook fires on /compress

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor): transcript user record opens with the empty timestamp element; transcript cells say what the recordings show (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec: tidy the transcript-file deviation wording (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor): pin the mock's transcript_path at the first preToolUse and beforeShellExecution (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec: disclose the doc's null-if-disabled transcript_path against the recordings (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(session-transcript-file): codex's recorded file-before-start behaviour is the mock's too, not a deviation (part of #161)

The entry had no kind and described behaviour the mock matches, so it declared nothing; the codex run still cites it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: task-stream-frames and hook-common-payload say what the recordings show for sub-agents

Drop the stale 'mock has no sub-agents' deviation of hook-common-payload/cursor for a
harness-lacks one the subagent-lifecycle-hooks recording grounds; prove the sub-agent half
of task-stream-frames/cursor and the sub-agent payload identity with e2e tests.

Part of #161.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: hook-common-payload discloses the sub-agent's own transcript path; task-stream-frames proves the background command's frames

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: disclose the subagentStart/Stop doc conflict; assert the recording has no sub-agent frame of its own

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: attribute subagentStart/Stop doc fields per event; assert neither fires

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: sub-agent payload test asserts Shell cwd is empty and no other event carries cwd

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: stream tests assert the recorded Task result and notification order on the recording too

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: hold the removal of the stale 'mock has no sub-agents' deviation until the user's words arrive

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: the task stream test asserts the recorded interleaved order of a turn's calls, for recording and mock

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: disclose the Task call's missing postToolUse as a doc/recording conflict and pin it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: drop the mock-not-modeled claims the recordings and tests prove false (the mock runs sub-agents)

Removes hook-common-payload's 'mock has no sub-agents' and task-stream-frames' 'mock runs no sub-agent / Task answered as an error'; the true no-await-tool clause stays.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: time a file tool's call; postToolUse reports a positive duration as recorded (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* posttooluse-payload/cursor: drop the false 'duration of 0 for a file tool' clause

runs/file-tools records positive durations (45.586, 1.695, 0.759, 27.441 ms).

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: file-size split; postToolUse test no longer claims distinct tool_use_ids (the recording repeats one across a turn's calls)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: move Edit beside the file tools; toolexec.go back under the size limit

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor-mock): fire beforeReadFile on each successful Read, as recorded (file-tools/cursor)

The mock carried the beforeReadFile docs marker but never fired the hook, and
the replay compare stripped it from the recording. It now fires after the
Read's preToolUse and before postToolUse with file_path, content and
attachments; the replay compares it; the cell drops its deviation.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a read of an existing file fires beforeReadFile between preToolUse and postToolUse, as recorded

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* refactor(cursor-mock): move the tool Result types to result.go, under the file-size limit

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(cursor): the mock fires beforeReadFile, so two deviations no longer say it does not

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a beforeReadFile hook is matched on the tool name

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* snapshots(cursor): runs/before-read-refusal, beforeReadFile hooks blocking reads

Six reads: exit 2, a JSON deny, invalid JSON, a crash (fails open), a failClosed
crash, and JSON with an unknown permission. Recorded with capture.sh at the
pinned cursor-agent 2026.09.28.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor-mock): a beforeReadFile hook that refuses blocks the read, as recorded (file-tools/cursor)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(cursor): beforeReadFile refuses a read as recorded, so pretooluse-refusal's deviation no longer lists it as unmodeled

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): the beforeReadFile refusal test also proves pretooluse-refusal/cursor

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(cursor): pretooluse-refusal cites the beforeReadFile refusal run and doc

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a blocked read gives the agent the failure hook's text, as recorded

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor-mock): the invalid-response wording is a beforeReadFile's alone, as recorded; the doc-only matcher test claims no cell

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a failed command's duration is the time it took

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: every tool_call frame carries toolCallId and hookAdditionalContexts (the after-tool hooks' context, on the completed frame)

Recorded in every tool_call frame of runs/*; the contexts in runs/additional-context.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: Task frames carry the harness's args and the preToolUse payload is the call as the model made it; five recordings replay green

subagent-lifecycle-hooks, foreground-subagent-result, subagent-stop-block-loop, subagent-transcripts and subagent-worktree-isolation leave notReplaying.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): the background Task test proves the receipt is at once, the recorded stream order and the payload fields (background-agent/cursor)

The capability-rigor judge found the test only asserted a completed taskToolCall
frame exists. It now also asserts the stream order of the Task call's frames,
the end notice and the result against runs/background-agent, that the receipt
precedes the agent's own command and the end notice follows it, and the
sub-agent's afterShellExecution output and sandbox and the sessionStart and
sessionEnd fields against the recording.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record a beforeReadFile hook timing out (failClosed or not) and a run with a second root added; replay tests

runs/before-read-timeout (cursor-agent 2026.09.28): a beforeReadFile hook that outruns its 1 s timeout is killed and the read goes through; the failClosed one blocks the read with a postToolUseFailure saying it failed closed and timed out after 1000ms. runs/multiroot-workspace: a run with --add-dir still tells every hook one workspace_roots entry. The mock accepts --add-dir (ignored, as recorded). Both runs are cited by their cells and replayed by e2e tests.

Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: a mock replay replays every captured sample of a recorded run

Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record --add-dir against reads inside it and outside every root; the mock ignores --add-dir because the recordings show no difference

runs/add-dir-access and runs/no-add-dir-access (cursor-agent 2026.09.28): the same three Reads (a file of a sibling directory, a file outside every root, a file of the project), with and without --add-dir for the sibling. All reads succeed in both, with the same hooks, and hook payloads name one workspace root. The replay now maps <TMP> (the directory the workspace sits in) so such paths replay.

Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: replay-every-sample states the decision only

Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: --add-dir is accepted only with --force, the mode its recordings cover

The recordings of --add-dir (runs/add-dir-access, no-add-dir-access, multiroot-workspace) are all of -p --force. Without --force the mock refuses the flag, naming it, until a recording covers that.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: split decision.go and toolhost.go back under the 150-line limit (hooks/combine.go, runner/toolhost_after.go)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* module: toolhost_after.go is in internal/toolcall's home beside toolhost.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record --print alongside -p and test it; the depth-limit test asserts the recorded Task hooks and the call left in the transcript

runs/print-long-form (cursor-agent 2026.09.28): -p --print is accepted and the run is as with -p. TestDepthLimitStopsNesting now also asserts the two recorded Task preToolUse hooks and that every recorded agent said it had no Task tool, and that the mock makes the Task call at the limit (it is in the transcript) with no hook seeing it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record matchers on Task and on beforeReadFile; the replay plays a recorded Task call

runs/hook-matchers-task-read (cursor-agent 2026.09.28): a preToolUse hook matched on Task runs for the Task call and one matched on Shell does not; a beforeReadFile hook matched on Read runs for a read and one matched on Shell does not; the postToolUse hook matched on Task never runs (no postToolUse fires for a Task call). The replay helper plays a recorded Task call as a Task whose sub-agent only replies.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload/cursor says a sub-agent is told from the main agent only by its own ids

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: take replay-every-sample out of the batch (kept on side/adr-replay-every-sample)

The adr-well-formed and adr-grounded judges refuse the user-worded decision sentence: it has a gap clause (not checkable, not present tense) and names no replay command or check. Rewording it needs the user's words, so the ADR waits on its own branch.

Sloprail-Cites-Tool: The only Decision bullet fails 'Checkable' and 'Present tense, current state only'

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: pretooluse-refusal/cursor discloses the agent_message doc and recording conflict

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: session-start-hook/cursor says its payload carries no start-kind field

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): the symlinked-cwd run proves the sessionStart hook fires once with no transcript path

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor drops the matcher run, to be cited with its capture

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor cites runs/hook-matchers-task-read, the capture of matchers on Task and beforeReadFile

Sloprail-Cites-Tool: captured runs/hook-matchers-task-read/samples/20261005-213709

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record Grep and Delete calls with matchers on each; the mock models both, as recorded

runs/hook-matchers-grep-delete (cursor-agent 2026.09.28): a preToolUse or postToolUse hook matched on Grep or Delete runs for that tool and not for the other, with the recorded hook inputs and outputs, grepToolCall and deleteToolCall frames with their results. The mock runs both tools; a Grep with no match and a Delete of a file it cannot read fail rather than guess (unrecorded).

Sloprail-Cites-Tool: captured runs/hook-matchers-grep-delete/samples/20261005-215835

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: pretooluse-refusal/cursor lists the agent_message conflict after the existing deviations

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record an MCP tool call with matchers on MCP:<tool>; the mock calls the project's stdio MCP server, as recorded

runs/hook-matchers-mcp (cursor-agent 2026.09.28, --approve-mcps, a local stdio MCP server written in sh): hooks name the tool MCP:echo; a matcher on it runs the hook and one on MCP:other does not; the agent first reads the tool schema in a getMcpToolsToolCall no hook sees, then makes the mcpToolCall. The mock starts the server of .cursor/mcp.json for the call and refuses MCP calls without --approve-mcps (unrecorded).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record a Task call with an invalid model; the mock answers it as recorded, with no sub-agent

runs/foreground-subagent-failure (cursor-agent 2026.09.28): the Task call is not started, its preToolUse hooks fire, and it completes with "Invalid model selection ... could not be resolved to a valid subagent model" and the account's model list, with no after-tool hook. The mock answers the same for any model but default and lists only default.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a background sub-agent launched by a sub-agent outlives it and its end is announced on the run's stream, as recorded; the nested background test asserts the recorded hooks, order and notification

runs/nested-subagents-background: the launching sub-agent reports STARTED while the background one goes on, its command's hooks and the sessionEnd come after, and the main stream carries a task_notification naming it. The mock owned the background sub-agent by its launcher, so it was never announced.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): deny beats ask on beforeShellExecution and refusal messages are joined on beforeShellExecution and beforeReadFile

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a blocked read's completed frame carries no args, as recorded; models inherit and default are the valid ones for a Task; before-read-refusal replays green

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: setup prepare.sh and args are installed and passed (a flag the mock does not model is refused), MCP, Grep and Delete calls are mapped, a hook script's own log tag joins its event's group

prepare.sh runs in the repository under the run's home as the capture runs it; setup/args go on the mock's command line, only the flags it models; a recorded CallDynamicTool is a mcp__<server>__<tool> call with the model's arguments (its lookup left to the mock); Grep and Delete are unified tools; the model catalogue after 'Allowed model slugs:' is not behaviour. A closed.sh-style log line ran concurrently with its event's payload, so it is grouped with that event: before-read-refusal no longer depends on timing.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: the new runs enter the exception list with their gaps, six greens leave it

session-resume-unknown replays green with args; no-add-dir-access, multiroot-workspace, print-long-form, hook-matchers-mcp, hook-matchers-grep-delete and foreground-subagent-failure are listed with the gap each shows.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a Read's result names the file resolved, Grep, Delete and MCP frames carry their toolCallId in their args, as recorded; the print-flag and multiroot runs read a file, so they replay

The new runs print-long-form and multiroot-workspace are recorded with a Read instead of a shell command; no-add-dir-access and hook-matchers-grep-delete replay green and leave the exception list.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: assistant text is shown as one frame when the next call starts or at the end of the turn, a refused Task completes with only its error; MCP calls carry the model's description and their own hook id; hook-matchers-mcp and foreground-subagent-failure replay green

Recorded (runs/foreground-subagent-failure): a call refused before it started shows no frame at its call, so the text before it goes out with the rest at the end. The replay adapter reads the model-written mcpToolCall description from the stream. Both runs leave the exception list.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter and pretooluse-refusal back to main, to be corrected again in a commit that carries the citation

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the beforeReadFile deviations of hook-matcher-filter and pretooluse-refusal match the recordings (the mock fires beforeReadFile); the agent_message doc conflict and the matcher runs are cited

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: the MCP server is run through internal/procexec; files split under the size limit and kept in their modules

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: the user's ~/.cursor/hooks.json is a hook source beside the project's, a hook in both runs once per source, as recorded

runs/hooks-all-matching-run-same-hook-two-sources. The e2e helper installs a run's user-hooks.json in its home and reads hook logs whose concurrent writes ran together.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hooks-all-matching-run/cursor says the mock reads the user source beside the project source, as recorded

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: assistant and tool_call frames carry model_call_id and timestamp_ms (the text at the end of the turn neither), as recorded; tests for them and for the common payload fields

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: replay_test.go back under the file-size limit; out.go belongs to the scenario module with frames.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): the transcript file exists when the first beforeShellExecution names it, as recorded

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-exit-code-semantics and hook-common-payload say where cursor differs from their statements (exit 0 with invalid JSON blocks; transcript_path null with transcripts on)

Sloprail-Cites-Tool: Cursor cell lists no deviation for exit 0 with invalid JSON on a permission hook

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: frame printing joins frames.go (no module change); the tests prove the sub-agent hooks that never fire, the notification detail as all the sub-agent said, and the transcript path on later hooks; hook-exit-code-semantics back to main for a cited re-correction

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-exit-code-semantics/cursor lists beforeReadFile among the events the mock fires and says an exit 0 with invalid JSON blocks, as the recordings show

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: a hook log line of several objects run together is read as the objects (hooks of one event append to one log side by side)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record workspaceOpen firing in print mode; the mock fires it before the session start with the recorded payload

runs/workspace-open (cursor-agent 2026.09.28): the hook fires once, first, with only the event name, the Cursor version, the workspace roots and the user email (null), no session, model or transcript path.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the cursor cells say what the recordings show of workspaceOpen, multiroot workspaces and model_id/model_params

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload/cursor drops the model_id/model_params deviation: it is a mock gap (the thinking hook is not fired), not a harness difference

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* afterAgentThought: a script's thinking block is a thought of the turn (core), the shared loop tells a Thinker host before the turn's message and calls, cursor fires the hook with the model's fields and its own generation

Optional: a host that does not implement turnloop.Thinker (claude, codex) never hears of a thinking block. A sub-agent and a turn that a finished background task starts are model requests of their own; the result frame keeps the run's first.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: the thoughts a recording holds are fed to the mock, response by response, and no longer dropped

The unified Call and Agent carry an optional Thinking; the adapter puts the i-th afterAgentThought of a conversation on its i-th response when the recording holds one per response, and refuses a recording that holds them for some only (no guessing, nothing dropped). The generation's number and name are masked, a thought and the preToolUse beside it are concurrent. nested-subagents replays green.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a step's text and its call are one transcript record, as recorded; tests for the transcript's records, the foreground/background contrast; the cells say workspaceOpen names no session and a background Task is modeled

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: a thought is put on the response its generation names, so a run whose model thought in some responses only is replayed (symlinked-cwd)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the cursor cells list afterAgentThought and workspaceOpen among the events the mock fires, and cite the run that replays with thoughts

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record matchers on afterAgentThought, workspaceOpen and sessionEnd with a thinking model; the mock tests a thought hook's matcher against AgentThought, names the run's model in the init frame and sessionEnd, and refuses models it has no record of

runs/hook-matchers-thought (cursor-agent --model cursor-grok-4.5-high): the afterAgentThought hook matched on AgentThought and the one with no matcher run for each thought, the one matched on nothing does not; workspaceOpen and sessionEnd run every hook whatever its matcher. The sessionEnd payload and the init frame name the model the run was started with.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor cites the thought matcher run and discloses that no postToolUse follows a Task call

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a Grep with any parameter but the pattern fails; unrecorded sub-agent models are refused as not modeled; --add-dir is gated on --force alone; the replay masks the clock and the service ids instead of dropping them, and tool_call frames say when a call began and ended

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor drops runs/nested-subagents, which has no matchers

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: StrReplace, GetDynamicTools and AwaitShell are modeled from the recordings that use them; file-tools and nested-subagents-depth replay and leave the exception list, the other four are listed with what is left

A StrReplace is an edit whose hooks see the whole file it makes and whose result carries the diff; a catalogue search finds nothing (the mock has no catalogue); a wait with no task named lasts as long as it was told. A hidden read before a write carries the write call's id.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a catalogue search is answered only where a recording shows it finding nothing, any other is refused as not modeled

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a refusal of something not modeled fails the run, a sub-agent's too; nested-subagents-depth is listed again, as its recording does not hold what the sub-agent's search answered

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: files back under the size limit and in their modules; a test that the result agentId is the sub-agent's own id and not the args'

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: the background Task test checks it returned at once; think.go joins the hooks module and refusal.go the scenario module

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: each mock's replay test replays every captured sample of a recorded run

Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: the exception list is accepted as it stands and may only shrink

Sloprail-Cites-User: Accept the cursor mock's initial replay exception list as it stands; from then on it may only shrink.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: foreground-subagent-result/cursor says the mock does not model a sub-agent failing mid-run

Sloprail-Cites-User: The cursor mock does not model a sub-agent failing mid-run; no headless run produces one.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload says a sub-agent event carries an identity of that sub-agent: its own session id, or the main session's together with an agent id

Sloprail-Cites-User: All lgtm

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor): pin the recorded afterFileEdit edits and a JSON allow on beforeReadFile; module for the refusal

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: take replay-every-sample out of batch/cursor (it lands from side/adr-replay-every-sample with batch/rules-pin)

Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload/cursor discloses the attachments entries the mock does not model

Sloprail-Cites-User: The mock's payloads always carry an empty attachments list; the entries a real cursor-agent adds for an attached file or rule (a type and a file_path) are not modeled, since no headless run can attach one.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor): a Task call at the depth limit is answered as an unknown tool

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(cursor-mock): refuse from the structured message, not a scan of printed frames; test the unrecorded-model, Grep-keys and --yolo --add-dir cases

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(cursor-mock): --add-dir is refused with --yolo too (only --force is recorded); test it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: declare the tools the mock plays (Edit, Grep, Delete, GetDynamicTools, AwaitShell, MCP tools) in the validated schema; toolspec: prefix and open tools

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* toolspec: a prefix tool may carry a Valid rule (mcp__<server>__<tool>, both non-empty); drop the unreachable Grep-keys path

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: pin the thought's null transcript path and the file tools' result frames; split replay's measure out of canon.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ure, resume frames, rigor tests (#158, #207) (#229)

* claude-mock: a run whose result is an error exits 1 and fires no Stop; recording runs/run-failure (an unrecognised model)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* snapshots: the docs re-frozen with the recording (capture.sh)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* run failure: the core says a failed result fires no Stop (scenario.Result.Failed); run-failure cited by noninteractive-run

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: a failing SubagentStart hook still hands back the report, driven by runs/hookerrors

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: an isolated sub-agent's hand-back trailer names its branch (worktreeBranch)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: the worktree trailer lines in agent_worktree.go (file-size)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: the hand-back tests compare the mock's output with the recording's

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: a terminated background shell is really gone (background-bash-reaped-at-exit)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: a foreground sub-agent's refused tool call is denied (foreground-subagent-result)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: -r and -c, the short forms of --resume and --continue

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: a notification turn ends with its own result frame (runs/bgagent)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: hook payloads sort only within one firing of one event, firings keep their order (#207 item 3)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: only the handlers' own log lines are sorted; payload lines keep their order (review)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: scoped drops (the mock's script key only from Agent inputs, assistant bookkeeping only when empty); no capability cell may name what is dropped (#207 item 8)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: init is also dropped after a compaction (comment, reasons); a safe type assertion

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude: stream frames carry a timestamp; the replay scrubs it instead of dropping the key (#207, replay-fidelity)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: timestamp is scrubbed, not dropped, so no cell excuse for it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a session resumed by name, path or --continue streams its SessionStart hook frames under the lookup's own id

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: resume by name, fork transcripts, no-resume with names and forks

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: the unknown resume's payload names the id, file and directory; no-resume invariant marker

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: hook payloads carry prompt_id and permission_mode, in the recorded key order

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: prompt/permission decisions live in the core, notifications continue the prompt, compaction starts one; split files; no-resume shapes tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: prompting config in the hooks module home; one capability marker

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: refuse output formats it does not produce; tests for compaction exit codes, session ids of forks/resumes, unknown resume; module-home files

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: drop the text-format no-resume test (text output is refused now)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude: deviations for what the mock refuses (--agent, text/json, stdin, partial messages, SIGTERM 143)

Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: fork hooks name the fork's transcript; no-resume equals an unknown session's refusal

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: a compaction streams only the compacted session's SessionStart hook frames (runs/compact)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: refuse --include-partial-messages and a piped stdin, as the deviations say

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: the result frame carries num_turns, stop_reason, terminal_reason, permission_denials, queued_turn_count, result_index, api_error_status, is_error, session_id

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: seven recordings replay green; removed from the exception list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: the stream ends with a result frame (not a fixed field order)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: the result frame's fields are derived from the run (turns since the last result, refused calls, result index, failure), a refused call's result carries tool_result_meta; two more recordings replay green

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: flag helpers in flags.go (file-size); the batch's noninteractive-run script prose

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* hooks: blockers that finish together are acted on last-configured first (all-hooks replays green, deterministic)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* hooks: ActedBlock back to the recorded semantics (the blocker that finished last, by Done); all-hooks stays on the exception list as on main (its order is a race); a test over built outcomes

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: every sample of a recorded run is replayed, each with its own model turns

Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* recordings: all-hooks-close-first and -second, two blocking hooks finishing 75 ms apart (A last / B last); docs re-frozen by capture.sh

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: compaction payloads carry the common fields; one hook in two settings files runs once

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: the idle ceiling stops the agent's process; the default ceiling is ten minutes

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a sub-agent streams its prompt, tool calls and results with parent_tool_use_id, and a task_progress frame before each call (runs/isolated-worktree, fg-subagent-bash); replay scrubs end_time

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: refuse --input-format and --max-budget-usd rather than ignore them; --max-turns and --include-hook-events refused as unknown flags (tested)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude: background-bash-reaped-at-exit declares the refused piped stdin

Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: the foreground agent's task frames exclude its progress frames; bgbash replays green and leaves the exception list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* hooks: ActedBlock by finish time: the last finisher, of those finishing within 50 ms the last configured, grounded in the all-hooks recordings; all-hooks replays green and leaves the list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: --max-turns ends the run with the error result, no Stop, exit 1 (recording runs/max-turns); recording runs/include-hook-events (not implemented: the flag stays refused)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* recording plugin-hooks-same-command (a plugin's copy of a hook stays separate from the project's: two runs); test proves the mock does the same

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: setup args (modelled flags passed on, refused flags a by-design skip) and prepare.sh are replayed; recorded exit 0 or 1; max-turns, plugin-hooks-same-command, plugin-hooks, transcript-retention replay green and leave the list; the three new entries are not added

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: Read calls, a recorded API failure and the run directory in calls are replayed; a successful non-Bash tool result streams without is_error (recorded: 83 results); tool-errors, tool-invalid-input, matcher, run-failure and all-hooks-close-second replay green and leave the list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: all-hooks-close-first dropped (a race: A, B, A; recorded in the ActedBlock grounding); a refused-flag recording is asserted refused by the mock, not skipped; a run with no sample is not a recording

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* batch/claude: new runs cited by their cells; denormalize.go and stream_tool.go split (file-size)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: unit tests follow the mapped Read tool and the modelled --model flag

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: worktree hooks ignore their matcher; PostToolUseFailure matches the tool's name

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock e2e: compact-nohooks output, startup hook frames, unknown resume on stderr and stdout (RunSplit); close-first re-captured, replays accept a recorded outcome where samples disagree (Reconcile)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude: hook-two-settings-files recorded (the same hook in two settings files runs once; replays green); bgbash after-result order tested; the new runs cited by hooks-all-matching-run

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude: noninteractive-run's refused list names --agent too, as the user's words do

Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a command hook's args run its program directly with them, shell is accepted and ignored (recordings hook-args, hook-shell); a test per recording; the hook-two-settings-files test proves hooks-all-matching-run too

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: --include-hook-events streams every main-thread hook's frames (PostToolUse ahead of the result frame; recording runs/include-hook-events replays green); a hook shell other than the recorded bash is refused

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: tests that a sub-agent's own PostToolUse context and a resumed session's SessionStart context are added (recording subagent-post-ctx); replay leaves a sub-agent's spend (tokens, duration, resolved model id) out of the comparison, so 6 recordings replay green and leave the exception list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* replay: a sub-agent's spend is masked, not dropped: its keys stay compared, their values and the usage line's counts become <MASKED>

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: runs of several steps (a resume) against one session store, so recording resume-session-start-ctx replays green and backs the test that a resumed session's SessionStart context is added; the mock no longer streams that context as a frame (the real stream has none); a resumed start's measures are masked

capture.sh also takes later steps as flat files (then-NN-prompt.txt, then-NN-args): the structure gate allows no deeper path under setup/.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* hook-additional-context/claude cites runs subagent-post-ctx and resume-session-start-ctx

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: e2e tests of prompt_id, SessionStart matchers by source, fork/resume context, the failure payload's tool and input; RunSplit uses procexec; files under the size limit

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: hooks of a session with a scratchpad are told scratchpad_dir (recording scratchpad-dir; after cwd, as recorded), none carries effort; replay installs a setup's env; RunSplit is a test helper again (e2etest.go is untouched)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* hook-common-payload/claude cites run scratchpad-dir

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: A10N_MOCK_NO_RESUME names the session a forking --resume resumed, not the fork's new id; tested

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a settings deny rule for an exact Bash command refuses the call as recorded (runs/permission-denied): permission_denied frame, the denial in the result, no PostToolUse; any other deny rule is refused

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: an unknown flag fails on stderr in claude's words before any frame (runs/invalid-flag); --bare and --agent are refused by name (runs/bare); the cell cites the runs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: allowManagedHooksOnly in a settings file is refused by name, not ignored

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a scenario that streams system/api_retry is refused by name (no model API to retry; a retry cannot be recorded on demand)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* ADR refused-flag-replay: a recorded run made with a flag the mock refuses replays as a check that the mock refuses it

Sloprail-Cites-User: A recorded run made with a flag the mock refuses replays as a check that the mock refuses that flag; it is not an exception-list entry.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* noninteractive-run/claude declares that system/api_retry events are not modeled

Sloprail-Cites-User: The claude mock does not model system/api_retry events; it refuses them until a recording drives them.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a turn that ends while a background agent works streams its result with the later turn's, after the notification turn's init frame (recording bgagent); the flag-refusal wording in the sr-agent flags test follows the new unknown-option message

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that a background shell killed at exit is reported after the result frame, as recorded (runs/bgbash)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a sub-agent's task_progress names an Agent call by its description and a Read or Edit by the file, as recorded (runs/meta, file-tools); test of nested foreground agents' task frames against runs/meta

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that a failed background shell is reported failed, against the new recording runs/bgbash-failed; the cell cites it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that an unknown resume ends its session before the result frame and that frame is all of stdout (runs/resume-unknown); the unknown-resume test keeps to the recorded form

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that the print wait for background agents defaults to ten minutes

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a scenario that calls the Monitor or Workflow tool is refused by name

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude cells declare what the mock refuses or does not model: --bare, --agent and other deny rules (noninteractive-run); the workflow and Monitor tools and the wall-time ceiling (print-waits-for-background-agents)

Sloprail-Cites-User: All lgtm

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: the refused Monitor and Workflow tool names cite the tools reference that lists them

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: the per-user temp folder claude-<uid> is the same on both sides (CI runs under another uid than the captures); the background-bash unit test no longer races a command that ends at once

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* tasks.Results holds the results of turns that end while a background agent works (core, next to AwaitAfterTurn); the deny-rule refusal lives in tool_hooks.go with the tool call; the refused-flag ADR names the code it governs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec(claude): hook-common-payload cites scratchpad-dir (runs in name order)

The cell lists the scratchpad-dir recording among its runs, and keeps the --agent deviation the user asked for.

Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them.

Sloprail-Cites-User: So those should not be limited or what's the problem? As long as capability describes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: hook_additional_context records carry the text as the agent receives it (rendered system reminder, renderedRole system), as recorded; one sr:capability marker for print-waits-for-background-agents

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* ADR refused-flag-replay: scope says every mock's replay test follows it

Sloprail-Cites-User: A recorded run made with a flag the mock refuses replays as a check that the mock refuses that flag; it is not an exception-list entry.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* replay masks end_time instead of dropping it (the mock's task_updated patch carries it, as runs/isolated-worktree shows); the test checks the key

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: runner.go under the size limit (settings load and invoker setup in runner_setup.go); flagError notes only the long-flag wording is recorded; one import block in control_records.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: --include-hook-events streams PostToolUseFailure's and a foreground sub-agent's SubagentStart/SubagentStop frames as recorded (runs/include-hook-events-more), in raw print mode Stop's too; it is refused when the run reaches an unrecorded hook (compaction, worktree, background sub-agent)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: prepare.sh runs with a claude that accepts only plugin commands (the mock reads a local marketplace itself; CI has no claude), so plugin-hooks and plugin-hooks-same-command replay there; the unknown-resume test's hook no longer sleeps near SessionEnd's budget

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: --include-hook-events also refuses the subagent_start control record, a scenario-written tool_result's PostToolUse and a background task's notification UserPromptSubmit (frames not recorded); only a manual compaction refuses on SubagentStop; the early-return sub-agent writes its held start frame; worktree and other cases in TestT001_19

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: stream_scan.go within the size limit; the noninteractive-run cell cites runs/include-hook-events-more

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* the deny rule's refusal is the core's (toolcall.DeniedByRule), the claude settings only parse their syntax; stream_turn.go within its ceiling; gofmt

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* the deny-rule decision sits beside noninteractive-run's one marker (internal/scenario), not under a second one

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…re, replay lists only shrink (#207), layering and snapshots-current fixes (#232)

* fix(rules): module-boundaries fails closed on a failed module lookup; test injects it

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): module-coverage fails closed on a failed ADR or module lookup; tests inject them

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): modules.sh load_modules and module_home_files report a failed lookup

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): module-distinct fails closed and hands its modules to jq by file; test injects the failures

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): module-leaks fails closed and hands candidates to jq by file; test injects the failures

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): concern-undeclared fails closed and hands ADRs and modules to jq by file; test injects the failures

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): module rules read through here-strings, not a printf into grep -q or awk (a pipefail SIGPIPE race); cases name their rule literally

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): module cases show the recovery and the nearest permit; mock judges read their input

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): judge rules show a real refusal, its recovery and the nearest permit (mock judges decide from input)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): the new judge-rule cases name their rule literally

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-covered refuses when its lookups fail, with a case (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-grounded refuses when a lookup fails; the catalog goes through files, not argv; with cases (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-rigor refuses when a lookup fails; the catalog, pairs and markers go through files, not argv; with cases (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): one refuses-when-a-lookup-fails case per capability rule, asserting the FileGuardChecked refusal and its reason (#204)

capability-covered's case asserts through the FileGuardChecked event; capability-grounded and capability-rigor fold
their separate prepare/subjects/inputs cases into one sr-checks run case each, as capability-reconciled's cases do.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): capability-grounded's case injects each prepare lookup with a marker only prepare.sh carries

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): the ADR loader, linked-adrs and confine refuse on a failed lookup; ADRs go to jq by file, not argv

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): invariant-covered, -rigor and -grounded refuse on a failed lookup

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): adr-linked, -grounded, -well-formed and -matches-sloprails refuse on a failed lookup; large values go by file

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): each rule of the invariant and ADR slice refuses on an injected lookup failure

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): the slice's cases also run through sr-checks run; invariant-covered permits when implemented and proven

adr-linked reads an ADR's sections with a here-string, not printf piped to grep -q: under pipefail a grep
that exits early made printf fail, which read as a missing section.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): behavioural cases for the nine rules of the invariant and ADR slice

Each rule gets a case beyond its lookup-failure one: a refusal with the rule's real reason, a permit, the
match's boundary, and the recovery the refusal names. The judges are mocks through SR_CHECKS_JUDGE_MOCKS
(adr-conformance, adr-grounded, adr-matches-sloprails, adr-well-formed, invariant-grounded, invariant-rigor),
deciding from their rendered prompts; the cited cases get the user's words from one scripted agent turn.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): the new cases name their rule literally in the jq that asserts its FileGuardChecked event

rule-tests-rigorous's script floor looks for one jq selecting the owning rule by name.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(checks): invariant-covered and capability-covered error on tooling failure, not a cached fail

Their spec helpers returned an empty SPEC when `ls spec/<kind>/*.yaml` failed and
harnesses() listed nothing on an incomplete tree, so the rules read "no specs" /
"no claude-mock/" and refused with a plain fail, which the engine caches as a
terminal verdict (seen on PR #178). Since sloprail #266 a refusal carrying
"error": true is never cached.

- _lib/changeset.sh: refuse_error (reason + "error": true); an unset SR_TREE and a
  failed git grep over markers are tooling errors.
- _lib/spec.sh: load_spec refuses with an error when yq/jq are missing, spec/<kind>
  is not in the tree, the listing fails or holds no *.yaml. Invalid YAML and every
  content finding stay plain refusals.
- capability-covered: no *-mock/ dir in the tree is an error.
- rule tests: errors-on-an-incomplete-tree for each rule (healthy -> pass, missing
  or empty spec dir -> error, genuine finding -> plain refusal).

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(checks): recovery and the other plain findings in the covered-rule error cases

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(checks): load_spec at top level in the pairs scripts; test the missing spec dir and mock dir

A refusal inside $(...) only ends the subshell, so rigor_pairs/reconcile_pairs no
longer call load_spec: their callers load $SPEC at top level (inputs-ready.sh and
prepare.sh now do). spec.sh avoids a pipefail-prone grep. New cases: capability-rigor
errors-on-a-missing-spec-dir (rigor and reconciled scripts), and a missing *-mock/
dir in capability-covered's case.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(checks): capability-rigor missing-spec-dir case asserts the engine's event

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(checks): no swallowed failure in the capability-rigor case

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(checks): an empty harness list is an error in reconciled, its subjects and snapshots-current

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): module-coverage tells an absent base ADR from an unreadable base

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): tooling failures in the module, capability, invariant and ADR rules are errors, not cached verdicts; the cases assert it

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): a failing subjects.sh is the engine's report, not a check's error flag

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-grounded, capability-covered and cells.sh refuse on a failed changed-files, harness or shape lookup (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): cases for every remaining capability lookup that refuses, with a skip-counting jq shim (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): no swallowed failure in the prepare case; capability-rigor's inputs-ready refusal and recovery are a case (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): load_touched_markers fails on an unreadable marker table, diff or file, and capability-rigor refuses (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-covered lists mocks itself (one refusal); changed-files matches are SIGPIPE-proof (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): cases for the marker/diff/pairs refusals, a stray z-mock, a >64KB changed list; header on the not-ready case (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): the >64KB case asserts its status instead of swallowing it (#204)

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-grounded narrowing list is built in a checked step; drop the unreachable refuse

The harness list was built inside a process substitution, where a failed jq is invisible (an empty file reads as null and the judge would get no providers). It is now built first, with a refusal, and a case injects the inner jq. touched_harnesses cannot fail (load_touched already ran and refuses), so its dead `|| refuse` is replaced by a comment.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-grounded keeps its refusal on the touched-harnesses read, with a case that fires it

The previous commit dropped `|| refuse` on touched_harnesses; the grounded-rule-changes judge rightly objected (it loosens a rule). The refusal stays, and an awk shim (touched_harnesses ends in an awk over the table) makes it fire in refuses-when-a-prepare-lookup-fails; without the refusal that case fails.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): capability-rigor's samples listing no longer refuses a run with no samples (#204)

The listing loop ended on a false [ -f ] &&, which under pipefail failed the substitution and refused a legitimately empty run. An if keeps only a real jq failure fatal. A case runs prepare.sh over a change citing a run with no samples.

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): the late #216 fixes land as errors; capability-covered's missing-mock message is the one #225 and the other rules use

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): layering also checks a package's test imports (adr-matches-sloprails finding), with a case

Sloprail-Cites-Tool: leaves out TestImports and XTestImports
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): the layering case asserts the sr-checks status instead of swallowing it

Sloprail-Cites-Tool: refuses-a-test-importing-another-mock is not a rigorous sr-test case
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): snapshots-current's drifted-doc case shows the recovery

Sloprail-Cites-Tool: never shows the agent fixing it
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): snapshots-current refuses when no *-mock directory exists, with a case

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): the object-id, layering and capability-grounded tooling failures are errors; yaml_str_json removed; module-leaks' BAD line is noted as a verdict

Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(rules): replay-exceptions-only-shrink: every replay exception list may only shrink, and no reason moves to a weaker category

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Sloprail-Cites-User: why fucking not? do
Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(rules): a replay list moved to another replay directory of its mock is compared with its old self; cases for the move

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): replay-exceptions-only-shrink sees a list renamed or named from another file, knows the reason categories, and reads the entries once into variables; the ADR names the categories

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): replay-exceptions-only-shrink keeps to the user's words: a created list is empty-only (a git mv rename is compared), no legacy exemption; the ADR is stated without time words

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* docs(rules): replay-exceptions-only-shrink's header names the three bypasses its other-file and reason-category checks close, with the output that showed each passing

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.
Sloprail-Cites-Tool: MISFIRE-PREFIX untriaged y weakened to y with no category: the old rule exited 0
Sloprail-Cites-Tool: MISFIRE-INIT init in zz_test.go adds a notReplaying entry: the old rule exited 0
Sloprail-Cites-Tool: MISFIRE-MOVE list moved to exceptions_test.go with a new entry: the old rule exited 0
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): capability-grounded's lookup cases use a judge mock that decides from its prompt, and harness-lacks-cited.sh has a refusal and recovery case

Sloprail-Cites-Tool: so they are the always-pass mocks the criterion calls proof of nothing
Sloprail-Cites-Tool: is never exercised by any case, so its refusal and recovery are untested
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): the harness-lacks case asserts capability-grounded's own outcome, refused and then passed, beside the script's reason

Sloprail-Cites-Tool: test.sh asserts no event of its owning rule capability-grounded
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(rules): the harness-lacks case reads the outcome with one jq over the events, not a constant filter

Sloprail-Cites-Tool: refuses-a-harness-lacks-deviation-with-no-citation is not a rigorous sr-test case: - test.sh runs jq
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* docs(adr): replay-exceptions-only-shrink's concern has no time word

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* docs(adr): replay-exceptions-only-shrink cites the user's words for the four reason categories

Sloprail-Cites-User: A replay exception's reason must start with one of adapter:, mock gap:, untriaged: or flaky:; a reason with none of these is refused.
Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): replay-exceptions-only-shrink refuses a new line naming the list in the generated test and a key with a tab, and matches any <mock>-mock/e2e path

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI
Sloprail-Cites-User: why fucking not? do
Sloprail-Cites-User: A replay exception's reason must start with one of adapter:, mock gap:, untriaged: or flaky:; a reason with none of these is refused.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

n-sviridenko and others added 9 commits October 6, 2026 18:34
…ness) pair; adr-well-formed asks only scope (#237, #238) (#242)

* feat(rules): capability-rigor judges per (capability, harness) pair; the shared sr:capability marker touches no harness's pair

Sloprail-Cites-User: A capability is judged per harness: one harness's missing tests never refuse a change that touches only another harness.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(rules): capability-covered checks only what a change touches: a capability's shape, the pairs whose cell, recording, markers or deviation ADRs changed, and the markers touched

Sloprail-Cites-User: A capability is judged per harness: one harness's missing tests never refuse a change that touches only another harness.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* feat(rules): adr-well-formed asks only that a Decision names where it applies, never for detail the user did not give

Sloprail-Cites-User: An ADR records the user's decision in their own words; adr-well-formed may ask only that it names where the decision applies, never for detail the user did not give.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(rules): per-harness scoping keeps its fail-closed tail, back-checks a marker whose cell was touched, and the lookup cases follow the pair-scoped subjects

Sloprail-Cites-User: A capability is judged per harness: one harness's missing tests never refuse a change that touches only another harness.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* test(rules): the shared-statement scenario of judges-per-harness asserts the refusal's reason

Sloprail-Cites-Tool: checks54-finding-A: the shared-statement scenario asserts a refusal with no reason
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* fix(rules): capability-covered checks the harness set for every capability when a mock dir is added or deleted

Per-pair scoping skipped the cell-presence check for untouched capabilities,
so a new <h>-mock/ with no cells, or a deleted one whose cells remained,
passed. The base and head mock-dir sets are compared now; a changed set (or
an unreadable base) puts every capability in scope for cell presence. New
case judges-a-changed-harness-set covers both directions.

Sloprail-Cites-Tool: a new mock dir with no cell: capability-covered did not refuse

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* test(capcells): the fixture's change adds the capability file, so it touches every pair it judges

The four tests ran covered.sh over an empty changeset. Under per-pair scoping an empty change touches no pair, so the
bare false, the false without docs or runs, the pending cell and the supported cell without a proving test were never
judged: the fixture, not the rule, was out of scope. The change now adds spec/capabilities/cap.yaml, which touches every
(capability, harness) pair; the assertions are unchanged.

Sloprail-Cites-Tool: cells_test.go:120: a supported cell without a proving test must be refused, got 0

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* test(subjects): capability-rigor subjects are (capability, harness) pairs; the expectations say so

The capability-rigor expectations still named bare capability ids ('a ', 'b '), but its subjects are <id>/<harness>.
They now expect 'a/claude ', 'b/claude '. The key lookups used the bare id too, so "b's proving test changed, b's key
moves" compared two empty keys (and the two "same key" checks compared empty with empty, vacuously); they look the pair
up now and compare real keys. capability-grounded keeps its bare ids.

Sloprail-Cites-Tool: FAIL: cap-rigor: a run of b touches b alone: 'b/claude ' != 'b '

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ty list (#250)

* fix(replay-exceptions-only-shrink): the one-line empty map is the empty list

gofmt writes an empty map as `var notReplaying = map[string]string{}`, so the two-line form the rule demanded is not gofmt-clean, and the rule refused the empty list it exists to reach (#239 emptied claude's list). A one-line empty map is now read as the declaration and its closing brace: emptying a list to {} passes, and an entry added back to a {} list is refused, naming it.

Sloprail-Cites-Tool: gofmt-writes-empty-map-q9: var notReplaying = map[string]string{}

Sloprail-Cites-Tool: with one "run": "reason" per line (file-guard "replay-exceptions-only-shrink")

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* replay-exceptions-only-shrink cases: a delete( of the list in the generated test is refused, as the comment says

The case's comment claims a delete( and a maps.Copy( are refused, but it only wrote the maps.Copy(. It now also adds a delete(notReplaying, ...) line and asserts the refusal.

Sloprail-Cites-Tool: rigor-finding-q9: case refuses-a-generated-test-that-gains-a-line-naming-the-list: its comment says

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…e cited docs mention (#234)

* capability-rigor: judge the statement's scope, not every behaviour the cited docs mention

The judge demanded an sr:proves test for every behaviour anywhere in a cell's cited doc pages, so each
run surfaced new "the doc says X, no test" items, even for doc behaviour outside what the capability
claims. It now judges whether the provider's tests prove the cell's statement (the capability's scope)
plus the declared deviations. The docs and runs ground the statement and are read for its meaning; they
are not a checklist, and doc behaviour outside the statement is not demanded. A recorded run replaying
green proves what it recorded. Fail-closed: a statement clause that no test and no replay proves still
refuses.

Three sr-test cases (mock judge: it proves the rubric reaches the prompt and the verdict path, how the
real model reads it is for sr-eval): a doc-only behaviour outside the statement is permitted; a
statement clause with no test is refused; a clause proven only by a green replay is permitted.

Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* capability-rigor cases: assert the rule's event in each test.sh, the refusal with its reason

Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* capability-rigor cases: the permit case's failure text names no refusal

Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* capability-rigor: only a run that replays green counts as proof

prepare.sh reads each harness's replay exception list (notReplaying in
<h>-mock/e2e/*/replay_allowlist_test.go; a list that cannot be parsed, or a harness with neither a list nor a
replay test, refuses) and marks every run replays: true|false with its notReplayingReason. The template says
only a replays="true" run proves a clause; a run on the exception list proves nothing. The mock judge honours
the flag, and a new case (refuses-a-clause-shown-only-by-a-run-that-does-not-replay) installs a notReplaying
entry for the run. The cases share one scope-lib.sh and judge.sh in capability-rigor/test-lib/, and the
file-guard.yaml comment is rewrapped.

Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* capability-rigor cases: the two refusals recover once a test asserts the clause

Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* capability-rigor: list every finding in the first refusal (harness-mocks#236)

The template tells the judge to finish the checklist and name every unproven clause of every subject in the
one refusal, one finding each, never only the first. The mock judge names all of them, and two cases cover it:
lists-every-unproven-clause-at-once (beta, gamma and delta in one refusal, fixed in one pass) and
reruns-after-an-unrelated-change-raise-nothing-new (the same single finding after an unrelated change). The two
refusal cases from the replay change now recover once a test asserts the clause, as rule-tests-rigorous asked.

Not done: handing the judge the previous stored verdict's findings and the diff since. A verdict is a pure
function of its cache key (rule hash, subject files, fingerprint); prepare's output is not in the key, the
store is an internal compressed format with no reader for a script (sr-checks show --json takes about a minute
per range), and mocked judges are not cached, so it cannot be proven here.

Sloprail-Cites-User: capability-rigor lists every finding for a subject in its first refusal, and a re-run raises new findings only for what changed since.
Sloprail-Cites-Tool: (the rubric says the fix is a test), but no case adds a test for beta

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* capability-rigor cases: the older sandboxes have a replay exception list, and the samples-listing injection lets the list's own jq -R through

Sloprail-Cites-Tool: c/claude: no replay exception list or replay test of claude-mock: cannot tell which runs replay green

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* capability-rigor cases: per-pair fixtures give each mock an empty replay exception list

After the rebase onto the per-pair rules, judges-per-harness and the doc-problem scenario of subjects.sh refused for the missing list instead of the behaviour they test.

Sloprail-Cites-Tool: c/claude: no replay exception list or replay test of claude-mock: cannot tell which runs replay green

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* docs-follow-recordings: the sandbox has an empty replay exception list

capability-rigor's prepare now needs each harness's replay exception list (or a replay test) to tell which runs replay green; the suite's sandbox had neither, so 'a drifted page is read' got the refusal instead of the page. The sandbox gets an empty notReplaying list.

Sloprail-Cites-Tool: docs-follow-fail-q9: FAIL: capability-rigor prepare: a drifted page is read

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* capability-covered: list every pending cell of every capability, touched by the change or not

adr/capability-once says the rule lists every pending cell on each run, but the per-pair scoping skipped untouched capabilities and pairs before the pending case, so only the pending cells of touched pairs were listed. The pending list is now taken for every capability before the scope filters; the checks themselves stay scoped. A Go test covers an untouched capability with a pending cell, still listed when the change only adds another capability's file.

Sloprail-Cites-Tool: adr-pending-q9: The ADR's Decision says `file-guard/capability-covered` lists every `pending` cell on each run, but

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* noninteractive-run/claude declares that system/api_retry events are not modeled

Sloprail-Cites-User: The claude mock does not model system/api_retry events; it refuses them until a recording drives them.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a turn that ends while a background agent works streams its result with the later turn's, after the notification turn's init frame (recording bgagent); the flag-refusal wording in the sr-agent flags test follows the new unknown-option message

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that a background shell killed at exit is reported after the result frame, as recorded (runs/bgbash)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a sub-agent's task_progress names an Agent call by its description and a Read or Edit by the file, as recorded (runs/meta, file-tools); test of nested foreground agents' task frames against runs/meta

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that a failed background shell is reported failed, against the new recording runs/bgbash-failed; the cell cites it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that an unknown resume ends its session before the result frame and that frame is all of stdout (runs/resume-unknown); the unknown-resume test keeps to the recorded form

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: test that the print wait for background agents defaults to ten minutes

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: a scenario that calls the Monitor or Workflow tool is refused by name

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude cells declare what the mock refuses or does not model: --bare, --agent and other deny rules (noninteractive-run); the workflow and Monitor tools and the wall-time ceiling (print-waits-for-background-agents)

Sloprail-Cites-User: All lgtm

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: the refused Monitor and Workflow tool names cite the tools reference that lists them

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude replay: the per-user temp folder claude-<uid> is the same on both sides (CI runs under another uid than the captures); the background-bash unit test no longer races a command that ends at once

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* tasks.Results holds the results of turns that end while a background agent works (core, next to AwaitAfterTurn); the deny-rule refusal lives in tool_hooks.go with the tool call; the refused-flag ADR names the code it governs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec(claude): hook-common-payload cites scratchpad-dir (runs in name order)

The cell lists the scratchpad-dir recording among its runs, and keeps the --agent deviation the user asked for.

Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them.

Sloprail-Cites-User: So those should not be limited or what's the problem? As long as capability describes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: hook_additional_context records carry the text as the agent receives it (rendered system reminder, renderedRole system), as recorded; one sr:capability marker for print-waits-for-background-agents

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* ADR refused-flag-replay: scope says every mock's replay test follows it

Sloprail-Cites-User: A recorded run made with a flag the mock refuses replays as a check that the mock refuses that flag; it is not an exception-list entry.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* replay masks end_time instead of dropping it (the mock's task_updated patch carries it, as runs/isolated-worktree shows); the test checks the key

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: runner.go under the size limit (settings load and invoker setup in runner_setup.go); flagError notes only the long-flag wording is recorded; one import block in control_records.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: --include-hook-events streams PostToolUseFailure's and a foreground sub-agent's SubagentStart/SubagentStop frames as recorded (runs/include-hook-events-more), in raw print mode Stop's too; it is refused when the run reaches an unrecorded hook (compaction, worktree, background sub-agent)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: prepare.sh runs with a claude that accepts only plugin commands (the mock reads a local marketplace itself; CI has no claude), so plugin-hooks and plugin-hooks-same-command replay there; the unknown-resume test's hook no longer sleeps near SessionEnd's budget

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor: replay every recording through the mock's own replay command (#158)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: review fixes (named gaps in the reasons, temp dir cleanup, empty hook log not green, print-script exit 2, multi-sample test)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor session-start-hook: cite runs/session-resume; drop proves markers from internals unit tests

The no-start-hook-on-resume claim rests on runs/session-resume, now in the cell's runs. The decision_test.go markers for pretooluse-refusal and hook-timeout sat on unit tests of Interpret; both cells keep their e2e proves.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: background-agent discloses the subagentStart/Stop doc and recording conflict; print-waits quotes the subagentStop sentence (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): hooks-all-matching-run/cursor cites the combined-refusal run and says the refusals are merged

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec: drop the cell note (notes field removed in #176)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): stop-block-cap/cursor cites the hooks doc's loop_limit default and discloses it conflicts with the print-mode recording (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): empty-tool-result-placeholder/cursor cites the exit-0 empty result

background-bash-start records ls exiting 0 with empty stdout: stream empty strings, hook {"output":"","exitCode":0}, no transcript tool result. Stays unsupported, with the sharper reason.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): manual-compaction/cursor cites the recorded auto compactions and says what they show

The cell cited only the headless /compress run. runs/compaction-transcript-continuity
records two real compactions with a preCompact payload (trigger auto, no hook after,
nothing stopping them); the hooks doc calls preCompact observational. Stays
unsupported: the request, the hook after and the stop are all absent.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(spec): manual-compaction/cursor says no compaction hook fires on /compress

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor): transcript user record opens with the empty timestamp element; transcript cells say what the recordings show (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec: tidy the transcript-file deviation wording (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor): pin the mock's transcript_path at the first preToolUse and beforeShellExecution (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec: disclose the doc's null-if-disabled transcript_path against the recordings (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(session-transcript-file): codex's recorded file-before-start behaviour is the mock's too, not a deviation (part of #161)

The entry had no kind and described behaviour the mock matches, so it declared nothing; the codex run still cites it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: task-stream-frames and hook-common-payload say what the recordings show for sub-agents

Drop the stale 'mock has no sub-agents' deviation of hook-common-payload/cursor for a
harness-lacks one the subagent-lifecycle-hooks recording grounds; prove the sub-agent half
of task-stream-frames/cursor and the sub-agent payload identity with e2e tests.

Part of #161.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: hook-common-payload discloses the sub-agent's own transcript path; task-stream-frames proves the background command's frames

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: disclose the subagentStart/Stop doc conflict; assert the recording has no sub-agent frame of its own

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: attribute subagentStart/Stop doc fields per event; assert neither fires

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: sub-agent payload test asserts Shell cwd is empty and no other event carries cwd

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: stream tests assert the recorded Task result and notification order on the recording too

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: hold the removal of the stale 'mock has no sub-agents' deviation until the user's words arrive

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: the task stream test asserts the recorded interleaved order of a turn's calls, for recording and mock

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: disclose the Task call's missing postToolUse as a doc/recording conflict and pin it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor: drop the mock-not-modeled claims the recordings and tests prove false (the mock runs sub-agents)

Removes hook-common-payload's 'mock has no sub-agents' and task-stream-frames' 'mock runs no sub-agent / Task answered as an error'; the true no-await-tool clause stays.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: time a file tool's call; postToolUse reports a positive duration as recorded (part of #161)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* posttooluse-payload/cursor: drop the false 'duration of 0 for a file tool' clause

runs/file-tools records positive durations (45.586, 1.695, 0.759, 27.441 ms).

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: file-size split; postToolUse test no longer claims distinct tool_use_ids (the recording repeats one across a turn's calls)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: move Edit beside the file tools; toolexec.go back under the size limit

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor-mock): fire beforeReadFile on each successful Read, as recorded (file-tools/cursor)

The mock carried the beforeReadFile docs marker but never fired the hook, and
the replay compare stripped it from the recording. It now fires after the
Read's preToolUse and before postToolUse with file_path, content and
attachments; the replay compares it; the cell drops its deviation.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a read of an existing file fires beforeReadFile between preToolUse and postToolUse, as recorded

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* refactor(cursor-mock): move the tool Result types to result.go, under the file-size limit

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(cursor): the mock fires beforeReadFile, so two deviations no longer say it does not

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a beforeReadFile hook is matched on the tool name

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* snapshots(cursor): runs/before-read-refusal, beforeReadFile hooks blocking reads

Six reads: exit 2, a JSON deny, invalid JSON, a crash (fails open), a failClosed
crash, and JSON with an unknown permission. Recorded with capture.sh at the
pinned cursor-agent 2026.09.28.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor-mock): a beforeReadFile hook that refuses blocks the read, as recorded (file-tools/cursor)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(cursor): beforeReadFile refuses a read as recorded, so pretooluse-refusal's deviation no longer lists it as unmodeled

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): the beforeReadFile refusal test also proves pretooluse-refusal/cursor

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* spec(cursor): pretooluse-refusal cites the beforeReadFile refusal run and doc

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a blocked read gives the agent the failure hook's text, as recorded

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* fix(cursor-mock): the invalid-response wording is a beforeReadFile's alone, as recorded; the doc-only matcher test claims no cell

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* test(cursor-mock): a failed command's duration is the time it took

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx

* cursor-mock: every tool_call frame carries toolCallId and hookAdditionalContexts (the after-tool hooks' context, on the completed frame)

Recorded in every tool_call frame of runs/*; the contexts in runs/additional-context.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: Task frames carry the harness's args and the preToolUse payload is the call as the model made it; five recordings replay green

subagent-lifecycle-hooks, foreground-subagent-result, subagent-stop-block-loop, subagent-transcripts and subagent-worktree-isolation leave notReplaying.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): the background Task test proves the receipt is at once, the recorded stream order and the payload fields (background-agent/cursor)

The capability-rigor judge found the test only asserted a completed taskToolCall
frame exists. It now also asserts the stream order of the Task call's frames,
the end notice and the result against runs/background-agent, that the receipt
precedes the agent's own command and the end notice follows it, and the
sub-agent's afterShellExecution output and sandbox and the sessionStart and
sessionEnd fields against the recording.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record a beforeReadFile hook timing out (failClosed or not) and a run with a second root added; replay tests

runs/before-read-timeout (cursor-agent 2026.09.28): a beforeReadFile hook that outruns its 1 s timeout is killed and the read goes through; the failClosed one blocks the read with a postToolUseFailure saying it failed closed and timed out after 1000ms. runs/multiroot-workspace: a run with --add-dir still tells every hook one workspace_roots entry. The mock accepts --add-dir (ignored, as recorded). Both runs are cited by their cells and replayed by e2e tests.

Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: a mock replay replays every captured sample of a recorded run

Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record --add-dir against reads inside it and outside every root; the mock ignores --add-dir because the recordings show no difference

runs/add-dir-access and runs/no-add-dir-access (cursor-agent 2026.09.28): the same three Reads (a file of a sibling directory, a file outside every root, a file of the project), with and without --add-dir for the sibling. All reads succeed in both, with the same hooks, and hook payloads name one workspace root. The replay now maps <TMP> (the directory the workspace sits in) so such paths replay.

Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: replay-every-sample states the decision only

Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: --add-dir is accepted only with --force, the mode its recordings cover

The recordings of --add-dir (runs/add-dir-access, no-add-dir-access, multiroot-workspace) are all of -p --force. Without --force the mock refuses the flag, naming it, until a recording covers that.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: split decision.go and toolhost.go back under the 150-line limit (hooks/combine.go, runner/toolhost_after.go)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* module: toolhost_after.go is in internal/toolcall's home beside toolhost.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record --print alongside -p and test it; the depth-limit test asserts the recorded Task hooks and the call left in the transcript

runs/print-long-form (cursor-agent 2026.09.28): -p --print is accepted and the run is as with -p. TestDepthLimitStopsNesting now also asserts the two recorded Task preToolUse hooks and that every recorded agent said it had no Task tool, and that the mock makes the Task call at the limit (it is in the transcript) with no hook seeing it.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record matchers on Task and on beforeReadFile; the replay plays a recorded Task call

runs/hook-matchers-task-read (cursor-agent 2026.09.28): a preToolUse hook matched on Task runs for the Task call and one matched on Shell does not; a beforeReadFile hook matched on Read runs for a read and one matched on Shell does not; the postToolUse hook matched on Task never runs (no postToolUse fires for a Task call). The replay helper plays a recorded Task call as a Task whose sub-agent only replies.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload/cursor says a sub-agent is told from the main agent only by its own ids

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: take replay-every-sample out of the batch (kept on side/adr-replay-every-sample)

The adr-well-formed and adr-grounded judges refuse the user-worded decision sentence: it has a gap clause (not checkable, not present tense) and names no replay command or check. Rewording it needs the user's words, so the ADR waits on its own branch.

Sloprail-Cites-Tool: The only Decision bullet fails 'Checkable' and 'Present tense, current state only': the clause "a mock that replays only the newest sample is a gap to close, not a convention" is a plan or description of existing code, not a rule

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: pretooluse-refusal/cursor discloses the agent_message doc and recording conflict

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: session-start-hook/cursor says its payload carries no start-kind field

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): the symlinked-cwd run proves the sessionStart hook fires once with no transcript path

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor drops the matcher run, to be cited with its capture

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor cites runs/hook-matchers-task-read, the capture of matchers on Task and beforeReadFile

Sloprail-Cites-Tool: captured runs/hook-matchers-task-read/samples/20261005-213709

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record Grep and Delete calls with matchers on each; the mock models both, as recorded

runs/hook-matchers-grep-delete (cursor-agent 2026.09.28): a preToolUse or postToolUse hook matched on Grep or Delete runs for that tool and not for the other, with the recorded hook inputs and outputs, grepToolCall and deleteToolCall frames with their results. The mock runs both tools; a Grep with no match and a Delete of a file it cannot read fail rather than guess (unrecorded).

Sloprail-Cites-Tool: captured runs/hook-matchers-grep-delete/samples/20261005-215835

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: pretooluse-refusal/cursor lists the agent_message conflict after the existing deviations

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record an MCP tool call with matchers on MCP:<tool>; the mock calls the project's stdio MCP server, as recorded

runs/hook-matchers-mcp (cursor-agent 2026.09.28, --approve-mcps, a local stdio MCP server written in sh): hooks name the tool MCP:echo; a matcher on it runs the hook and one on MCP:other does not; the agent first reads the tool schema in a getMcpToolsToolCall no hook sees, then makes the mcpToolCall. The mock starts the server of .cursor/mcp.json for the call and refuses MCP calls without --approve-mcps (unrecorded).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record a Task call with an invalid model; the mock answers it as recorded, with no sub-agent

runs/foreground-subagent-failure (cursor-agent 2026.09.28): the Task call is not started, its preToolUse hooks fire, and it completes with "Invalid model selection ... could not be resolved to a valid subagent model" and the account's model list, with no after-tool hook. The mock answers the same for any model but default and lists only default.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a background sub-agent launched by a sub-agent outlives it and its end is announced on the run's stream, as recorded; the nested background test asserts the recorded hooks, order and notification

runs/nested-subagents-background: the launching sub-agent reports STARTED while the background one goes on, its command's hooks and the sessionEnd come after, and the main stream carries a task_notification naming it. The mock owned the background sub-agent by its launcher, so it was never announced.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): deny beats ask on beforeShellExecution and refusal messages are joined on beforeShellExecution and beforeReadFile

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a blocked read's completed frame carries no args, as recorded; models inherit and default are the valid ones for a Task; before-read-refusal replays green

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: setup prepare.sh and args are installed and passed (a flag the mock does not model is refused), MCP, Grep and Delete calls are mapped, a hook script's own log tag joins its event's group

prepare.sh runs in the repository under the run's home as the capture runs it; setup/args go on the mock's command line, only the flags it models; a recorded CallDynamicTool is a mcp__<server>__<tool> call with the model's arguments (its lookup left to the mock); Grep and Delete are unified tools; the model catalogue after 'Allowed model slugs:' is not behaviour. A closed.sh-style log line ran concurrently with its event's payload, so it is grouped with that event: before-read-refusal no longer depends on timing.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: the new runs enter the exception list with their gaps, six greens leave it

session-resume-unknown replays green with args; no-add-dir-access, multiroot-workspace, print-long-form, hook-matchers-mcp, hook-matchers-grep-delete and foreground-subagent-failure are listed with the gap each shows.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a Read's result names the file resolved, Grep, Delete and MCP frames carry their toolCallId in their args, as recorded; the print-flag and multiroot runs read a file, so they replay

The new runs print-long-form and multiroot-workspace are recorded with a Read instead of a shell command; no-add-dir-access and hook-matchers-grep-delete replay green and leave the exception list.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: assistant text is shown as one frame when the next call starts or at the end of the turn, a refused Task completes with only its error; MCP calls carry the model's description and their own hook id; hook-matchers-mcp and foreground-subagent-failure replay green

Recorded (runs/foreground-subagent-failure): a call refused before it started shows no frame at its call, so the text before it goes out with the rest at the end. The replay adapter reads the model-written mcpToolCall description from the stream. Both runs leave the exception list.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter and pretooluse-refusal back to main, to be corrected again in a commit that carries the citation

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the beforeReadFile deviations of hook-matcher-filter and pretooluse-refusal match the recordings (the mock fires beforeReadFile); the agent_message doc conflict and the matcher runs are cited

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: the MCP server is run through internal/procexec; files split under the size limit and kept in their modules

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: the user's ~/.cursor/hooks.json is a hook source beside the project's, a hook in both runs once per source, as recorded

runs/hooks-all-matching-run-same-hook-two-sources. The e2e helper installs a run's user-hooks.json in its home and reads hook logs whose concurrent writes ran together.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hooks-all-matching-run/cursor says the mock reads the user source beside the project source, as recorded

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: assistant and tool_call frames carry model_call_id and timestamp_ms (the text at the end of the turn neither), as recorded; tests for them and for the common payload fields

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: replay_test.go back under the file-size limit; out.go belongs to the scenario module with frames.go

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: session-transcript-file keeps its codex section as on main; the kind-less deviation removal is handed to batch/codex as a patch

Sloprail-Cites-Tool: index 08642bc7..49cd9654 100644

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor-mock): the transcript file exists when the first beforeShellExecution names it, as recorded

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-exit-code-semantics and hook-common-payload say where cursor differs from their statements (exit 0 with invalid JSON blocks; transcript_path null with transcripts on)

Sloprail-Cites-Tool: Cursor cell lists no deviation for exit 0 with invalid JSON on a permission hook

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: frame printing joins frames.go (no module change); the tests prove the sub-agent hooks that never fire, the notification detail as all the sub-agent said, and the transcript path on later hooks; hook-exit-code-semantics back to main for a cited re-correction

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-exit-code-semantics/cursor lists beforeReadFile among the events the mock fires and says an exit 0 with invalid JSON blocks, as the recordings show

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: a hook log line of several objects run together is read as the objects (hooks of one event append to one log side by side)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record workspaceOpen firing in print mode; the mock fires it before the session start with the recorded payload

runs/workspace-open (cursor-agent 2026.09.28): the hook fires once, first, with only the event name, the Cursor version, the workspace roots and the user email (null), no session, model or transcript path.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the cursor cells say what the recordings show of workspaceOpen, multiroot workspaces and model_id/model_params

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload/cursor drops the model_id/model_params deviation: it is a mock gap (the thinking hook is not fired), not a harness difference

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* afterAgentThought: a script's thinking block is a thought of the turn (core), the shared loop tells a Thinker host before the turn's message and calls, cursor fires the hook with the model's fields and its own generation

Optional: a host that does not implement turnloop.Thinker (claude, codex) never hears of a thinking block. A sub-agent and a turn that a finished background task starts are model requests of their own; the result frame keeps the run's first.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: the thoughts a recording holds are fed to the mock, response by response, and no longer dropped

The unified Call and Agent carry an optional Thinking; the adapter puts the i-th afterAgentThought of a conversation on its i-th response when the recording holds one per response, and refuses a recording that holds them for some only (no guessing, nothing dropped). The generation's number and name are masked, a thought and the preToolUse beside it are concurrent. nested-subagents replays green.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a step's text and its call are one transcript record, as recorded; tests for the transcript's records, the foreground/background contrast; the cells say workspaceOpen names no session and a background Task is modeled

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: a thought is put on the response its generation names, so a run whose model thought in some responses only is replayed (symlinked-cwd)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: the cursor cells list afterAgentThought and workspaceOpen among the events the mock fires, and cite the run that replays with thoughts

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: record matchers on afterAgentThought, workspaceOpen and sessionEnd with a thinking model; the mock tests a thought hook's matcher against AgentThought, names the run's model in the init frame and sessionEnd, and refuses models it has no record of

runs/hook-matchers-thought (cursor-agent --model cursor-grok-4.5-high): the afterAgentThought hook matched on AgentThought and the one with no matcher run for each thought, the one matched on nothing does not; workspaceOpen and sessionEnd run every hook whatever its matcher. The sessionEnd payload and the init frame name the model the run was started with.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor cites the thought matcher run and discloses that no postToolUse follows a Task call

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a Grep with any parameter but the pattern fails; unrecorded sub-agent models are refused as not modeled; --add-dir is gated on --force alone; the replay masks the clock and the service ids instead of dropping them, and tool_call frames say when a call began and ended

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-matcher-filter/cursor drops runs/nested-subagents, which has no matchers

Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor: StrReplace, GetDynamicTools and AwaitShell are modeled from the recordings that use them; file-tools and nested-subagents-depth replay and leave the exception list, the other four are listed with what is left

A StrReplace is an edit whose hooks see the whole file it makes and whose result carries the diff; a catalogue search finds nothing (the mock has no catalogue); a wait with no task named lasts as long as it was told. A hidden read before a write carries the write call's id.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a catalogue search is answered only where a recording shows it finding nothing, any other is refused as not modeled

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: a refusal of something not modeled fails the run, a sub-agent's too; nested-subagents-depth is listed again, as its recording does not hold what the sub-agent's search answered

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: files back under the size limit and in their modules; a test that the result agentId is the sub-agent's own id and not the args'

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor-mock: the background Task test checks it returned at once; think.go joins the hooks module and refusal.go the scenario module

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: each mock's replay test replays every captured sample of a recorded run

Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* cursor replay: the exception list is accepted as it stands and may only shrink

Sloprail-Cites-User: Accept the cursor mock's initial replay exception list as it stands; from then on it may only shrink.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: foreground-subagent-result/cursor says the mock does not model a sub-agent failing mid-run

Sloprail-Cites-User: The cursor mock does not model a sub-agent failing mid-run; no headless run produces one.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload says a sub-agent event carries an identity of that sub-agent: its own session id, or the main session's together with an agent id

Sloprail-Cites-User: All lgtm

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor): pin the recorded afterFileEdit edits and a JSON allow on beforeReadFile; module for the refusal

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* adr: take replay-every-sample out of batch/cursor (it lands from side/adr-replay-every-sample with batch/rules-pin)

Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* spec: hook-common-payload/cursor discloses the attachments entries the mock does not model

Sloprail-Cites-User: The mock's payloads always carry an empty attachments list; the entries a real cursor-agent adds for an attached file or rule (a type and a file_path) are not modeled, since no headless run can attach one.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* test(cursor): a Task call at the depth limit is answered as an unknown tool

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(cursor-mock): refuse from the structured message, not a scan of printed frames; test the unrecorded-model, Grep-keys and --yolo --add-dir cases

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* fix(cursor-mock): --add-dir is refused with --yolo too (only --force is recorded); test it

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG

* claude-mock: --include-hook-events also refuses the subagent_start control record, a scenario-written tool_result's PostToolUse and a background task's notification UserPromptSubmit (frames not recorded); only a manual compaction refuses on SubagentStop; the early-return sub-agent writes its held start frame; worktree and other cases in TestT001_19

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: declare the tools the mock plays (Edit, Grep, Delete, GetDynamicTools, AwaitShell, MCP tools) in the validated schema; toolspec: prefix and open tools

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: stream_scan.go within the size limit; the noninteractive-run cell cites runs/include-hook-events-more

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a Shell call's frames carry what Cursor reports of it (parsed command by a shell parser, limits and flags, ids, description); five runs replay green

hook-matchers, hooks-together, session-end-hook-output, stop-block-cap-always and plugin-hooks leave notReplaying; the rest of the hooks half have their remaining gaps named.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a shell command's own hook names no transcript; replay places a thought by request, not by the response number alone

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a Task call needs only its prompt; one without it is never started, its completed frame has no args (agent-input-validation, -description replay green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: several texts before a call, and the answers that end a turn the harness goes on from (a notification, a Stop block), are steps of the script; a Stop block streams its feedback frame, the error notice once, and the cap's informational and notification frames, the override counting a turn; a background agent is announced with its launch and has no prompt frame. cap, stops and hook-exit-codes leave the exception list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: a tool the adapter has no mapping for goes on under its own name (the mock runs or refuses it); a wakeup's wall-clock time, seconds-to-go and scheduledFor are masked. schedule-wakeup-limits leaves the exception list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor replay: hooks that name a call by an id of their own are matched to the call by its input, a file call that failed has no args, no hook log is no hook ran, the flaky hand-written comparison of additional-context goes

A call's hooks name it by an id other than its call id in 17 of the recordings: the adapter puts the id on the call whose command, prompt, pattern or path the hook saw, and refuses an id no call matches; the mock names the call's hooks by it (hook_tool_use_id, mock-only). A completed Read or Edit that errored carries no args. A recording with no payloads.jsonl whose hooks were configured, and whose stream compares green, proves the hooks never ran. additional-context, hook-exit-codes, hook-fail-closed, hook-matchers-after-events, hook-timeout, hook-timeout-early-output, hooks-all-matching-run-same-hook-two-sources, pretool-refusal, pretool-refusal-combined, pretool-refusal-file-tools, session-start-continue-false, stop-block-continuation, tool-failure and user-prompt-submit-hook leave notReplaying, with symlinked-cwd, noninteractive-force-write and foreground-subagent-failure, which the shell frames and the hook ids turn green too.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: a run that took no model turn ends with a result frame that has no terminal_reason or api_error_status (prompt-blocked, prompt-blocked-json, prompt-blocked-suppressed replay green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: a child's environment names the harness's messaging socket and secret, Bash its executable, a SessionStart hook its env file; --name names a session it starts; a named session's prompts carry its title; a replay runs an earlier run of the recording (prepare.sh) and the resume, continue, fork and no-persistence flags (nested-session-env, subprocess-session-env, resume-*, no-session-persistence replay green)

Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: a /compact step replays (the compaction, its summary text, the preserved messages, an unwritten logical parent and the summarizer's own output are read from the recording); /compact is the harness's own command: no UserPromptSubmit, no Stop, its result says local_command; the summary streams as a synthetic user frame; SessionStart:compact names the model; a first-run --session-id is the replay's own. compact and compact-nohooks leave the exception list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: a background shell a sub-agent launched is announced owned_by_subagent, as recorded (runs/fg-subagent-bash); the entry leaves the exception list

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a foreground command's hooks name it by an id of its own, its tool_input carries the model's timeout (none for a background one); the shell sees CURSOR_REQUEST_ID and CURSOR_RIPGREP_PATH, hooks CURSOR_RIPGREP_PATH; replay does not compare what Cursor leaves unsettled (the first beforeShellExecution's transcript path); a refused command reports the time it took (shell-exit-status, symlinked-cwd, noninteractive-force-write, noninteractive-no-force, nested-session-env, subprocess-session-env and, as a result, 12 hooks-half entries replay green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor: shell syntax the recordings show is modeled (double-quoted words, < >> 2>&1, command substitution), the rest is refused by name; stop-hook-payload re-recorded and closed

runs/shell-syntax records a double-quoted word (string), < (fd 0), >> (fd 1), 2>&1 (operator >&, number target, all-quiet) and $(...) (command_substitution, its commands listed after the outer). Variables, && and ||, here-documents, globs, subshells, loops, assignments, backticks and background jobs have no recording: the validator refuses them (Param.Unmodeled). runs/stop-hook-payload is re-recorded with a prompt that avoids the catalogue search.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: a Stop hook that blocks streams its feedback to the agent as a synthetic user message, and its error notice once per run; a sub-agent re-run after a blocked SubagentStop streams no second prompt; a replay makes the answers a Stop hook sent the model on from (stops, hook-exit-codes, cap-sub replay green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a turn's closing text waits for the run's end (a finished background shell's turn follows); a background shell's pid and id are masked, its result in the order Cursor words it; its executionTime is measured; commands naming the harness's ripgrep are refused as not modeled (background-bash-start, bg-bash-reaped-at-exit, task-notifications-bg, task-notifications-inturn replay green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: a run, earlier or later, starts in a directory of the repository (cwd, symlink) with the project files of that directory (symlinked-cwd, forkresume replay green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: allowlist after merging sprint/cursor-a (the union of both halves' removals)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: --tools names the tools a run has; an Edit's omitted replace_all is filled in as recorded, the stream naming the input as sent; a replay maps Write, Edit and Glob and plays the inputs as the model sent them (file-tools replays green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor replay: the harness's own skill files a run read are laid out from the recorded read, a read's content_length counts UTF-16 units; schedule-wakeup-ask closes

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a sub-agent's conversation takes the id a replayed reply quotes (mock-only agent_id), a background Task reports its measured duration; background-agent, print-waits-for-background-agents replay green; shell ids masked as numbers (task_id, taskId)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* fix the merge: core.Agent carries the agent's own id (the cursor adapter reads it), a test closed

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: a sample recorded when a setup file read otherwise than it does now (the model's Read of hook.sh shows it) is replayed with the file as the model read it (hookmix replays green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a response's Task calls start after its other calls; consecutive text-only responses of one turn are one answer (nested-subagents-background replays green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude replay: adapter.go and denormalize.go within the size limit

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a background shell's file has the header and footer Cursor writes, a wait on a named shell is answered, a response's Task goes after its other calls; core: steps and exit statuses in the unified recording

task-stream-frames stays listed with what is left: the calls' preToolUse order and the Read hooks a wait fires.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: a command a foreground sub-agent's end ended still gives the parent its follow-up turn; a process listing is compared as a listing, not by which processes (foreground-subagent-bash-ends-with-response replays green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* claude-mock: a foreground Bash command of the main agent that runs 3 seconds or more is a task (task_started when it crosses that, task_notification when it ends), as recorded in the new runs bash-long-foreground and bash-short-foreground (hook-timeout replays green)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor replay: a run of several steps is replayed step by step over one home, exit statuses are compared (core), a later step may resume the session or be refused; session-fork closes

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* cursor-mock: record the catalogue search a sub-agent at the depth limit made (runs/catalogue-search-task: matches []), so nested-subagents-depth replays green

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(replay): the codex adapter gives the mock a run's -c agents.max_depth, -C and --ephemeral, and installs a project layer's hooks

hooks-all-matching-run-same-hook-two-files, nested-subagents-nowait, compaction-transcript-continuity and hook-command-subdir
replay green and leave notReplaying. The mock makes a relative -C absolute, as the harness does.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(replay): a sub-agent's own sub-agents are replayed (nested-subagents, nested-subagents-limit)

The loader attaches sub-agents recursively, taking a sub-agent's receipts from its own rollout (its spawns are not in the
run's stream), and the scenario gives each its own script, named across the whole tree.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(replay): a refused run, a run's own command line, a resume of an unknown session and a stopped compaction replay

The adapter reads the recorded command line (--json, --skip-git-repo-check, --dangerously-bypass-*) and gives the mock the same
flags, skips the git setup for a run recorded outside a repository, and expects the recorded exit status. A run the harness
refused (no rollout, non-zero exit) replays as that refusal; a turn aborted with no interrupt is a stopped compaction.
noninteractive-run-git-check-refused, noninteractive-run-no-git-check, session-resume-unknown, manual-compaction-auto-blocked and
manual-compaction-auto-post-stopped leave notReplaying.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(replay): a run's env file and its prepare.sh (with codex plugin and sed stubs) are replayed: nested-session-env, plugin-hooks

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(replay): the model's apply_patch calls are replayed (file-tools, file-tools-failure)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(replay): a hook's own non-JSON lines are kept in the compared log (bg-bash-reaped-at-exit)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* feat(codex-mock): write_stdin polls a command left running; a resumed script cell is replayed (task-stream-frames, subagent-stop-block-loop-cap)

The mock keeps a yielded command's output and item, and …
…41d413 (#264)

Verdicts are keyed by their input only (sloprail #290, #298): the store moves to
schema v2026-10-07, which the old pin refuses. The old pin's store gc also
crashes on dictionary training (fixed on main).


Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…e recording's gates, not timing (#263)

* fix(claude-mock): the order of a parent's and its sub-agents' frames comes from the recording's gates, not timing

Replays of bgagent-concurrent-limit and bgagent-nested-launcher swapped a
parent's text or hook with a sub-agent's frame under load. The gates that
order agents' steps had gaps, each closed here from the recording's own times:

- a step waits for the calls of sub-agents below its sub-agents (the inner
  agent's tool_use was ahead of the main agent's answer), via ChildCalls.Via;
- a step waits for a sub-agent's call the sample shows carried out ahead of it
  (the shell's task_started ahead of the answer), via ChildCalls.Executed;
- a step of an agent waits for the calls of the agents above its parent to
  have finished (main's PostToolUse ahead of the inner agent's PreToolUse),
  via Gate.AncestorDone;
- a sub-agent launched by a sub-agent's call is started after that call's
  PostToolUse when the payloads show it, as the main agent's already was.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* refactor: split the gate and spawn-log files under the 150-line limit

file-size refused gates.go, gate.go and spawn_log.go (176-178 lines). Moved by
responsibility, no code changed: the gate walk helpers (gates_walk.go), the hold
functions (gate_hold.go), the carried-out-call progress (progress_exec.go) and a
spawned agent's progress (spawn_progress.go).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
… their finish order (#262)

An event's hooks run together, so the log holds them in finish order, which
the recording does not promise.


Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Only the main session runs sr-checks run; sloprail/sloprail#294 ships the gate (off by default).


Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
Sloprail-Cites-User: you can have it in this plugin disabled by default and then enabled explicitly in our tool repos

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…es; adr/replay-every-sample (#253)

* mocks: a flaky: entry is replayed three times by every mock's generated replay test, and fails when none is green

claude, codex and cursor generated_replay_test.go run a flaky: entry through replayUntilGreen (flakyRuns = 3) and fail it when no run is green, instead of passing it on one green run or skipping it. Each package has a flaky_replay_test.go for replayUntilGreen.

Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* replay-exceptions-only-shrink: the replay test package is pinned whole, the replay files are read from the syntax tree, and the generated tests are pinned to canonical copies

tools/replaycheck reads each mock's notReplaying list and generated replay test from the syntax tree and type information instead of their text, and compares replayUntilGreen and TestGeneratedReplay with the canonical copies in the rule's folder (one per mock: claude, codex, cursor), so a flaky: entry is still run three times and failed when never green. A mock's replay_allowlist_test.go and generated_replay_test.go are no longer deleted, renamed or moved: that would end the replay without removing an entry. A reason that starts with none of adapter:, mock gap:, untriaged: or flaky: is still refused. Cases for the package pin, the text a grep would miss and the map outside its file; the move cases go, as a move is now refused.

Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI

Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.

Sloprail-Cites-User: A replay exception's reason must start with one of adapter:, mock gap:, untriaged: or flaky:; a reason with none of these is refused.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* adr: a mock's replay replays every captured sample of a recorded run (linked to the replay-exceptions-only-shrink rule that pins the replay tests)

The decision is the user's own words; the pin on generated_replay_test.go is what keeps a mock's replay test from being weakened to fewer samples.

Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* replay-exceptions-only-shrink: the timed replays may be skipped on macOS and run three times on Linux; any other -skip of a replay is refused

The workflows and a mock's Makefile are judged too: a go test -skip there may name only TestGeneratedReplay/all-hooks-close-(first|second) (directly, or through TIMED_REPLAYS defined as exactly that), and a workflow that skips them must also run them on their own with -count=3. Any other skip, a widened TIMED_REPLAYS, or a skip with no run is refused. Cases for both directions.

Sloprail-Cites-User: Fail, never skip

Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* replay-exceptions-only-shrink: the timed-replay skip is the user's decision and requires Linux, three runs and every PR; adr/replay-every-sample drops its link to the pin rule

adr/replay-exceptions-only-shrink states the timed-pair exception in the user's words (macOS e2e jobs skip it; the Linux timed-replays job runs it three times on every PR). The rule now also requires the workflow to be triggered by pull_request and the job that runs the timed pair to be on ubuntu, with cases for both. adr/replay-every-sample goes back to file-guard/adr-conformance: the pin rule does not check which samples a replay covers.

Sloprail-Cites-User: Timing-sensitive replays whose recorded gaps a CI runner's timers cannot keep may run on Linux only, three times each, as long as they run on every PR.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah

* adr/replay-exceptions-only-shrink: the skip bullet states the policy in the user's words

adr-grounded refused the bullet for deciding more than the cited words (it named
the tests, the platform and the job). It now says only what the user decided:
every other replay skip is refused, and timing-sensitive replays may run on
Linux only, three times each, on every PR. The names stay in the rule's script.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
Sloprail-Cites-User: Timing-sensitive replays whose recorded gaps a CI runner's timers cannot keep may run on Linux only, three times each, as long as they run on every PR.
Sloprail-Cites-User: Yes, ban other skips

---------

Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant