Repository navigation
review: harness-mocks changes since the 2026-10-03 demo (do not merge) - #240
Draft
n-sviridenko wants to merge 40 commits into
Draft
n-sviridenko wants to merge 40 commits into
n-sviridenko wants to merge 40 commits into
Conversation
…l pinned binaries (#139) * snapshots: pin harness versions; a binary bump no longer re-judges every capability The harness binary a recording was made with is per run (run.yaml); the MANIFEST's top-level `version` (and each page's `version`) are gone. A doc page is frozen by its own sha256 and `fetched` date, and MANIFEST `pin` only names which binary capture.sh runs next. touched.sh treats only a page whose sha256 changed as in question, so bumping the binary or re-freezing a page with an unchanged hash invalidates nothing; verdict keys carry page hashes only. The three MANIFESTs are migrated (hashes untouched, fetched = the date each hash was first committed). tools/harness-bin installs an exact claude (npm), codex (npm) or cursor-agent (versioned download) into a cache of its own, never the global install, and `path` refuses a binary that does not report exactly the version. capture.sh runs that binary with auto-update off, links the user's login as before, and refuses to add a sample to a run recorded at another version unless --rerecord. `capture.sh pin <version>` installs and sets the pin. Sloprail-Cites-User: this ownership of different versions of harnesses probably also has to be a part of the infrastructure of mocks Sloprail-Cites-User: if the version is just updated, it doesn't mean that we need to rescan all the features. So there has to be a mechanism on how to kind of freeze it. Sloprail-Cites-User: we can kind of pin the things Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * harness-bin: refuse an escaping symlink in a tarball; capture.sh all checks the pinned binary before dropping samples, pin never overwrites a MANIFEST Review fixes: untar accepted absolute and ../ symlink targets; `all` removed every run's samples before discovering the pinned binary was missing; `pin` fell back to overwriting an existing MANIFEST (and its docs) when yq failed. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * adr(pinned-harness-versions): keep to the decisions the user stated and the rules check Sloprail-Cites-User: if the version is just updated, it doesn't mean that we need to rescan all the features. So there has to be a mechanism on how to kind of freeze it. Sloprail-Cites-User: we can kind of pin the things Sloprail-Cites-User: this ownership of different versions of harnesses probably also has to be a part of the infrastructure of mocks Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * adr(pinned-harness-versions): link only the rules that decide it; drop the unverifiable cache sentence Sloprail-Cites-User: if the version is just updated, it doesn't mean that we need to rescan all the features. So there has to be a mechanism on how to kind of freeze it. Sloprail-Cites-User: we can kind of pin the things Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
The marker was written as "# sr-mark: ci-verify"; the gate sloprail/gate/ci-verify-required looks for "sr:ci verify" (written by `sr-mark apply ci --verify=<path>:<line>`), so it kept reporting that no committed file shows CI verifying the file-guards. Same step, same comment position; only the marker text changes. Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…on (#149) * ci(claude-mock): release darwin binaries, checksums.txt and --version The release now builds linux and darwin (amd64, arm64) with the tag embedded as main.version, and uploads checksums.txt. a10n-claude-mock --version prints the embedded version (dev when unstamped). The mock never answered --version for scenarios, so nothing existing is shadowed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * docs(adr): release-version — a mock's --version is its release, stamped at build time Sloprail-Cites-User: confirm the 3 release-version bullets Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(sloprail): adr-conformance also judges the claude-mock release workflow adr/release-version's first decision (the release stamps the tag into the binary) lives in .github/workflows/publish-claude-mock.yml, which the rule's match never covered. Sloprail-Cites-User: Make the rule also watch `publish-claude-mock.yml` Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
#155) * fix(ci): the e2e git identity is test-process environment, not the machine's global git config A workflow step wrote the global git config; copied onto a developer's machine it rewrote their real ~/.gitconfig (commits of PR #142 were authored ci-local@sloprail.invalid). The test-e2e target now exports GIT_AUTHOR_*/GIT_COMMITTER_*, the step is gone, and tools/repoguard fails on any workflow/Makefile/script/Go source that writes the global or system git config. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(repoguard): see commands split over lines, --file on a home gitconfig, more file kinds Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(repoguard): comments never join a command, real line numbers, XDG config via --file, deleted tracked files Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(repoguard): a Go split joins only inside an open call; every line is also judged alone Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(repoguard): judge each shell command of a line alone; join a multi-line slice literal Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(repoguard): a read inside a substitution or comment no longer exempts a write; one report per command Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… (pilot) (#140) * codex: replay every recording through the mock's own `replay` command (pilot) The mock replays a recording itself: `a10n-codex-mock replay [--print-script] <run-dir>` turns the recorded run's model turns into a scenario script in the format any scenario has (in memory, nothing stored), runs the mock on the run's own setup, and has the shared core (internal/replay) compare the mock's whole event stream and hook payloads with the recording's. Exit 0 green, 1 differs, 2 not replayable yet. The table-driven test only enumerates snapshots/runs/* and calls the same function once per recording folder. Recordings that do not replay green are on an explicit allow-list with a reason; an entry whose run is gone, or that now replays green, fails the test (the list only shrinks). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * replay: the core speaks one unified format, every heuristic is the harness adapter's The core (internal/replay) takes a recording as normalised turns (agents, calls, a unified tool vocabulary) and compares normalised output; it holds no heuristic. Reading a rollout into turns, the mock's script format, the dropped/rewritten fields and the ordering of hook payloads move to the codex adapter (codex-mock/internal/replay), behind core.Adapter so claude and cursor can add theirs. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * replay: the adapter starts children through procexec, takes its environment from the entrypoint Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * replay: an unparsable recording line fails the replay; the codex mock refuses feature switches The adapter reads the recording's stream, hook log and rollouts strictly: a line that is not a JSON object is an *Unbuildable naming the file and line, not a line left out of the comparison. The codex mock refuses --enable / --disable (it implements no feature switch), so a run asked for `--enable multi_agent_v2` fails loudly instead of passing as the default mode. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * codex-mock: the feature-switch refusal is a file of its own (exec.go stays under the size limit) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * scenario module: codex-mock/feature_switches.go is in its home beside exec.go Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * replay adapter: run.yaml and shell command lines are parsed, not scanned; the allowlist only shrinks - load.go reads run.yaml with a YAML parser into a typed struct, and a rollout belongs to a thread by the thread id its file name carries (exact), not by a substring - rules.go loses its shell regexes: a command line is parsed with mvdan.cc/sh/v3/syntax (shell.go) to unwrap `/bin/<shell> -c[l] <cmd>` and to recognise the replayed `sleep N` job; the ps listing parser stays (ps.go) and has unit tests - TestAllowlistOnlyShrinks (CI job "replay allowlist only shrinks", ALLOWLIST_BASE=origin/main) fails when notReplaying gains an entry compared with the base branch Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * adr fail-fast-unimplemented: a mock refuses what it does not implement; the codex mock does The codex mock refuses every flag it implements nothing of (--enable/--disable, -o, --output-schema, --ephemeral, --sandbox, --profile, --color, --thread-source, --ignore-*, --strict-config, --approve-for-me) and any -c key but agents.max_depth, with an error that names it; the noninteractive-run deviations say so; claude-mock/run_flags.go is the ADR's one legacy exception (the claude and cursor mocks follow). Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * adr fail-fast-unimplemented: keep only what the user decided, and name the code it governs Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * adr replay-exceptions-only-shrink: the replay exception list may only shrink Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * replay-exceptions-only-shrink: a file-guard enforces the ADR, replacing the CI test The file-guard compares the notReplaying map's keys at the base and head and refuses an added one. The ADR names the map and the guard; the Go test and its CI job go (one mechanism). Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * replay-exceptions-only-shrink: sr-test cases, and an emptied list is a legitimate shrink Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * tests: the refusal cases assert their reason; every refused flag is covered, --full-auto too Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * replay-exceptions-only-shrink tests: each case proves both the permit and the refusal, with the reason Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * tests assert the rule's events and reason directly; noninteractive-run names stdin-as-prompt and resume --last as not modeled Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * replay-exceptions-only-shrink tests: show the recovery after the refusal Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * codex-mock: a test that `exec resume --last` is refused and nothing runs Sloprail-Cites-User: yes - refuse all unimplemented - general rule: fail fast in mocks - bc its business critical Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * modules: one home per file: the command line and the output frames leave internal/scenario internal/cli owns each mock's entry points and flag refusal, internal/frames the output event frames; internal/scenario keeps driving the script and routing records. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * Revert "modules: one home per file: the command line and the output frames leave internal/scenario" The split made module-leaks refuse (the A10N_MOCK_* protocol read outside the scenario home) and module-distinct still named other files; the module layout is fixed in its own PR (issue filed). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ngs must be clean (#142) * Judges list every failure; deviations need user words; recordings must be clean - judge templates: a fail verdict lists EVERY failing item, one numbered line each - capability-grounded: user words for a new deviation or a newly unsupported cell; undisclosed doc/recording conflict fails; a deviation negating the cell's defining clause means supported: false (ADR capability-grounding updated) - snapshots-current + capture.sh: unparsed hook-log lines, harness error lines/frames and a mode the mock does not imitate refuse a recording (_lib/recording.sh, expected.yaml declares legitimate ones) - tests/rules.sh covers all of it; run in CI Sloprail-Cites-User: #142 (review) <-- fix Sloprail-Cites-User: (c) "When a doc and a recording disagree, the cell must say so as a deviation." Sloprail-Cites-User: (d) "A deviation that contradicts what the cell claims means the cell is not supported." Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #142: drop the every-failure paragraph and rules.sh, quote only for what the mock does not model - the "Every failure, in one verdict" section leaves the judge templates: the engine keeps only the first line of a multi-line fail reasoning (judge.go reasonFromVerifierOutput), so asking for a list would not reach the agent - tests/rules.sh and its CI step go; the tests move to sr-test cases in a later PR - added-or-removed.sh: a user quote is required only for a new modeled-surface deviation (the mock lacks what the harness does); a deviation or supported:false grounded by a recording that shows the harness lacking it needs none (ADR capability-grounding) Sloprail-Cites-User: #142 (review) <-- fix Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #142: drop the recording checks (_lib/recording.sh, expected.yaml, capture.sh hooks) Recordings are trusted as recorded: snapshots-current and the three capture.sh go back to what main has. Sloprail-Cites-User: #142 (review) <-- fix Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #142: deviations carry a kind; user words only for what the mock does not model - capability.cue: a deviation has an optional `kind`: mock-not-modeled | harness-lacks - capability-grounded/added-or-removed.sh (deterministic): words are required when a mock-not-modeled deviation (or one with no kind) is added or changed, or a cell turns supported: false without a cited recording; a harness-lacks deviation and a recorded unsupported cell need none - ADR capability-grounding and the rule's comments describe it Sloprail-Cites-User: #142 (review) <-- fix Sloprail-Cites-User: (b) "A deviation where the mock doesn't model something the harness does needs my quote; one backed by a recording of the harness doesn't." Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * chore(spec): give each deviation its kind (migration) adr modeled-surface -> kind: mock-not-modeled (97 less 3 unsure), adr capability-grounding -> kind: harness-lacks (47 less 2 unsure). Five entries are left without a kind for the user to decide (the mock differs from the harness in part, or what the entry names is unclear): hooks-all-matching-run cursor[0], session-transcript-file codex[1], subagent-transcripts codex[0], codex[1], cursor[1]. Content otherwise unchanged. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #142: a kind: harness-lacks deviation must cite a doc or a run harness-lacks-cited.sh (a script check ahead of the judge in capability-grounded) refuses a cell with a harness-lacks deviation that cites neither docs nor runs: that deviation is waived from the user's words because a cited doc or recording shows it (ADR capability-grounding). Sloprail-Cites-User: #142 (review) <-- fix Sloprail-Cites-User: (b) "A deviation where the mock doesn't model something the harness does needs my quote; one backed by a recording of the harness doesn't." Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * chore(spec): classify the five deviations left without a kind hooks-all-matching-run cursor[0], subagent-transcripts codex[0] and codex[1]: mock-not-modeled; subagent-transcripts cursor[1]: harness-lacks; session-transcript-file codex[1] is dropped (the mock does the same as the harness, so it is not a deviation). The user decided each. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #142: kind is required on every deviation capability.cue: `kind` is required (every deviation is classified), ADR capability-grounding says so. Sloprail-Cites-User: Every deviation says whether the mock doesn't model it or the harness lacks it Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): cells whose deviation says the harness lacks the defining clause are unsupported Rule (d): codex agent-input-validation, bash-tool-result, compaction-transcript-continuity, task-stream-frames, transcript-record-envelope; cursor plugin-hooks, task-stream-frames become supported: false, each keeping the doc and recorded runs that show the harness lacking it. The harness-lacks deviation becomes the reason; the mock-not-modeled ones go (nothing is mocked). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): the remaining cells #142's rules flag - (d) cursor agent-input-validation and foreground-subagent-bash-ends-with-response are supported: false (their docs and runs kept; the harness-lacks text is the reason) - (c) cursor noninteractive-run and print-waits-for-background-agents, codex session-transcript-file disclose the doc-versus-recording conflict - codex transcript-record-envelope: the reason names turn_context's cwd and the item events' thread id - cursor hooks-all-matching-run[0]: kind harness-lacks (Cursor itself runs the hook once per source) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): codex session-transcript-file is unsupported (rule (d)) Both of its deviations negate a clause of the statement (keyed by working directory; the file does not exist when the start hook runs), so the cell is supported: false, with its doc and recorded run. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * codex-mock, cursor-mock: markers of cells that became unsupported go to the supported cell the code serves - cursor background.go notificationFrame and its test: task-notifications/cursor - cursor agent_task.go keeps foreground-subagent-result/cursor; subagent_shell_test: task-notifications/cursor - codex collab.go and its test: foreground-subagent-result/codex Comments only; no behaviour change. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #142: a cell may carry notes, each harness's own way of providing the capability capability.cue: optional `notes` on a providers cell. A statement stays abstract; where a harness does the feature in its own way (a location, a timing), that is a note, not a deviation. Sloprail-Cites-User: maybe statement should be more abstract and deviated per harness? Sloprail-Cites-User: each hartness just has its own transcript resolution - otherwise how we'd strucutre transcript frature? Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): abstract statements, each harness's own way as notes; ten cells supported again agent-input-validation, bash-tool-result, compaction-transcript-continuity, task-stream-frames, transcript-record-envelope, plugin-hooks, foreground-subagent-bash-ends-with-response and session-transcript-file name no single harness's mechanism now; where a harness does the feature its own way (codex agent-input-validation, bash-tool-result, compaction-transcript-continuity, task-stream-frames, transcript-record-envelope, session-transcript-file; cursor agent-input- validation, task-stream-frames, plugin-hooks, foreground-subagent-bash-ends-with-response) the cell says so in `notes`, and the mock-not-modeled deviations dropped earlier are back. Stay unsupported: cursor compaction-transcript-continuity and transcript-record-envelope, codex foreground-subagent-bash-ends-with-response (they lack the feature). Sloprail-Cites-User: maybe statement should be more abstract and deviated per harness? Sloprail-Cites-User: each hartness just has its own transcript resolution - otherwise how we'd strucutre transcript frature? Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…#176) Sloprail-Cites-User: and why it introduced notes btw? this is an extra level of complexity and N cases to jduge etc Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ported again (#186) #142 moved these markers off task-stream-frames (codex, cursor) and foreground-subagent-bash-ends-with-response (cursor) while the cells were unsupported; the cells are supported again, so capability-covered flagged them as missing markers. Comment lines only, no behaviour change. Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ng and refuses nothing (#167) * feat(snapshots): docs follow recordings; a doc change re-judges nothing and refuses nothing No check reads a live doc page; capture.sh re-freezes the docs its cells cite with a recording; capability-grounded/rigor keys carry recordings, never doc hashes; doc_copy_var fixes DOC_ERROR: unbound variable; sr-test cases for the four rules. Sloprail-Cites-User: only recording changes matter - doc changes don't - bc recording changes when changed then docs are pulled Sloprail-Cites-User: but why the fuck we need this drift job? we only pull docs when needed ad-hoc or Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(snapshots-read-only): assert the one refusal, no refused event for the recovery and boundary Sloprail-Cites-User: only recording changes matter - doc changes don't - bc recording changes when changed then docs are pulled Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…158) (#184) * claude: replay every recording through the mock's own `replay` command (adapter + generated test) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: split adapter and load by responsibility (file-size) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: review fixes (explicit assistant drops, ids found by the adapter not the core, refuse two sayings, specific reasons) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* ci: run the mock tests three times to surface flaky ones (#151) Unit packages run with -race -count=3 (COUNT in each mock's Makefile). The e2e suites take minutes, so two extra replicas per mock run in parallel next to the required e2e job instead of -count=3 in one job: three runs under CI load, no extra wall time. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * ci: keep the repeat in the workflow, not the mock Makefiles A change under *-mock/ re-runs the whole-tree capability-covered guard, which is red on main for unrelated missing provides/proves markers. The unit repeat now calls go test directly and the e2e replicas run the Makefile's commands themselves. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…_use_result, wire_tool_inputs (#158) (#192) * claude-mock: stream frames carry session_id, parent_tool_use_id, tool_use_result and wire_tool_inputs Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: stream frame stamping in its own file (file-size) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: stream frame stamping lives with the records (module home) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…m (part of #161) (#195) * turnloop: a turn may have several tool calls; cursor-mock starts all before completing any scenario.Turn reads the tool_use lines of one turn (Tools; Tool stays the first); turnloop starts them all before completing any for a host that opts in through turnloop.Interleaver. claude-mock and codex-mock do not opt in and carry out the first call as before. cursor-mock opts in: shell started, Task started, shell completed, Task completed (runs/task-stream-frames). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * agent_task: reflow the finishSubagent comment Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
#177) The mock took 0 as no yield and ran the command to its end; Codex returns the receipt at once, whatever the command does (recorded: task-notifications-bg). The replay found it; task-notifications-bg leaves the exception list. Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…161) (#196) * cursor-mock: the thinking-frames test is not vacuous; a positive control shows the mock prints a script's assistant text (part of #161) The test set assistant frames aside from the replay, whose scripts carry no text, so "no assistant frames" held whatever the mock did. It now asserts the recording has assistant frames, the replay has no thinking frames, and a new test shows a script's text is an assistant frame in the stream. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor-mock: a test pins the tool_call frame bodies the force-write recording shows (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… wording, drifted-page cache, escaped path, trailing-slash lookup, tests) (#191) Header and judge prompts no longer say a doc is verified against its freeze; doc_section/slug removed; a drifted page is cached once; trailing-slash URLs found; all-wiring and slash tests; the capability-grounded add-e case asserts the exact reason and recovers. Sloprail-Cites-User: Apply the #167 reviewer's follow-ups: fix the stale snapshots header, dead helpers, judge wording about live docs, re-fetching drifted pages, the unescaped regex, and the trailing-slash doc lookup, with tests. Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ing (#203) * cursor noninteractive-run: the mock-not-modeled deviation claimed the mock emits no assistant frames; it prints them for script text, and emits no thinking frames (part of #161) Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor noninteractive-run: the mock stream lacks the recorded assistant frames too, not only thinking frames (part of #161) Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…egexes (#170) * codex replay: read the model's JavaScript with a parser, not regexes The adapter parses each exec call's script (goja's parser) and evaluates what decides the tool calls: literals, constants, objects, arrays, templates, store/load and the receipts of earlier calls. A call that depends on a result, or JS the adapter does not read, makes the recording unbuildable instead of being guessed. user-prompt-submit-hook-bg now replays green and leaves the exception list. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #170: scope names, refuse tool calls inside functions and aliases, read bracket keys Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * review of #170 (2): refuse var, tagged templates, defaulted params, mutating methods, exponent numbers, dropped arguments Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #170 (3): receiver before arguments, optional chains, dead zone, lone surrogates, whole-number yields, a spawn with no arguments agent-input-validation replays green and leaves the exception list. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #170 (4): refuse reads of null, a spawn argument that is not an object Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #170 (5): a binary operator's right side runs unless it is ||, && or ??; the base of a ?. always runs Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #170 (6): a ?. on a known non-null value does not short-circuit; tests for the chain's end and operands Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #170 (CI judge): the harness's other tool options go to the mock as given, not dropped by the adapter Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #170 (8): refuse non-finite options and another working directory; tests for the pass-through Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #170 (9): sub-agent scripts get the run's directory, a workdir that is not text is refused Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…tch both ways (#153) (#189) * feat(rules): capability-reconciled, cells and sr:provides adapters match both ways (#153) A deterministic file-guard scoped to the (capability, harness) pairs a change touches: a supported cell needs its sr:provides code and a cited recording that exists; a sr:provides marker needs an existing, supported cell. Replaces the sr:provides halves of capability-covered (whole tree, forward and back). Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): rigorous cases for capability-reconciled; cases for capability-covered Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): capability-reconciled covers a supported cell citing no run Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(capcells): the sr:provides half moved to capability-reconciled; capability-covered tests assert sr:proves Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-reconciled sees a cited recording deleted or replaced Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * docs(adr): capability-once names the cell's deviations field and links capability-reconciled Sloprail-Cites-User: The capability-once ADR names the deviations field and links the capability-reconciled rule. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * docs(adr): capability-once says each deviation names an ADR that exists Sloprail-Cites-User: each deviation names its ADR, and that ADR must exist Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): capability-reconciled cases assert the refusal's own reason text Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-covered re-runs when an ADR is deleted or renamed, with a case Sloprail-Cites-User: each deviation names its ADR, and that ADR must exist Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-reconciled fails closed when the touched pairs cannot be worked out; the catalog goes through a file, not argv Sloprail-Cites-User: sohuld just making sure that reconcillation is both ways Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… limited in width (session-end flake cause) (#201) * codex replay: the scratch repository, the machine's text, the scratch root, at most a few replays at once The recording's scratch repository has one empty init commit; its commit id, the harness's npm package root and the scratch directory (<TMP> is the root the repository, CODEX_HOME and TMPDIR sit in) are not behaviour; sub-agents' scripts sit beside the repository, as the real run's has none. Hooks run under wall-clock limits (SessionEnd: one second by default), so replays are limited in how many run at once: sixteen at once killed a recorded half-second hook, one alone never did. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * replay width is not decided here: the session-end entries stay, as flaky with the cause found Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * tests: sub-agent scripts are Scripts after the merge with main Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adapter.go under the size limit: the placeholders in a file of their own Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…son cites the recording (#171) * spec(compaction-transcript-continuity): cursor stays unsupported, reason cites the recording Part of #161. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec(compaction-transcript-continuity): reason names the turn_ended line and where dynamic_tools sits Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ed (#169) * test(cursor): plugin-hooks tests stop pinning an order the recording does not establish (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor): plugin-hooks comments say why no order is pinned (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(cursor): a loaded plugin's hooks run before the project's, as recorded; tests pin it (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor): a manifest-named hooks file replaces hooks/hooks.json (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor): label the manifest-replaces-default assertion as doc-based (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor): plugin-hooks wording says what is recorded, what is configured order, what is doc (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…nt's next call (#174) * fix(cursor-mock): the aborted notice of a sub-agent's command follows the parent's next tool call, as recorded Part of #161. The test now asserts the whole recorded frame order unfiltered. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * refactor(core): the deferred report of what ends at a sub-agent's response is the core's, cursor sets when Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(core): one sr:capability marker for foreground-subagent-bash-ends-with-response Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor): foreground-subagent-result asserts the recorded args, hook and PINEAPPLE-7 result; the stream says unspecified for a generalPurpose Task Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * refactor(core): the deferred-report helper lives beside, not in, the sr:capability file Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…on-end flaky entries leave the list (#206) * replays run with limited concurrency: an ADR and the core's gate, both mocks' replay tests use it The two session-end entries (flaky: load) replay green under the limit and leave the exception list. Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * replay concurrency: the core's Run holds a slot; the ADR decides only what the user said Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * ADR replay-concurrency: names where the limit is held, no history Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * ADR replay-concurrency: the user's words, with the slot count and where it is held Sloprail-Cites-User: Replays run with limited concurrency so hook timeouts are not hit under load: the replay core's Run (internal/replay) holds one of at most min(max(NumCPU/2, 1), 4) process-wide slots for each replay, for every mock. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…213) * grounded judges: a short user reply grounds exactly what it answered The adr-grounded, invariant-grounded and capability-grounded judges are handed each cited quote with its transcript path:line and may Read. Their rubric now says how to read a short reply ("lgtm", "yes", "1. lgtm"): the user's quote is the only authority, the surrounding messages are context to interpret it, a short approval grounds exactly what it answered and nothing beyond, the approved proposal must itself state the change, an earlier answered proposal does not carry over, an assistant message alone grounds nothing, and what is read from the record is data (same rubric as sloprail's grounded-rule-changes, sloprail #264). Two sr-test cases per rule: "lgtm" grounds the change proposed a few messages earlier; the same "lgtm" does not ground an unrelated change (refused with the judge's reason). Sloprail-Cites-User: They just get the line number of JSON error plus my quote, and then they can just go and read the JSON error to figure things out. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * short-approval cases: the refused unrelated change recovers with the proposed one The unrelated-change cases now continue after the refusal: the agent drops the unrelated commit, makes the change that was proposed, cites the same "lgtm", and the rule permits it (refused, then passed). Sloprail-Cites-User: They just get the line number of JSON error plus my quote, and then they can just go and read the JSON error to figure things out. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * short-approval cases: strict is_error check, mock-judge disclaimer, transcript noted beside allowed_tools Sloprail-Cites-User: They just get the line number of JSON error plus my quote, and then they can just go and read the JSON error to figure things out. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
… make (#231) * toolspec: one validator for the tool calls a scenario script asks a mock to make (ADR tool-calls-validated) The core's schema and check: an unknown tool, an unknown parameter, a missing required one, a wrong type or an option value the mock does not implement is refused, naming the tool and the parameter; a kind the real harness answers itself is returned for the mock to answer as recorded. Ungrounded checks a schema against recorded calls. Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * toolspec: the claude, codex and cursor runners validate every tool call a script asks for, against a schema each declares The shared turn loop (codex, cursor) and claude's stream scan check each call before it is played; a refusal in a sub-agent's script fails the whole run. Schemas are grounded in the recordings (a test per harness), the kinds a harness answers itself are listed with their recorded run, and parameters the mock implements that no run shows are grounded in the docs. ToolSearch, recorded, is implemented. Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * toolspec: the shared turn loop validates through the host (turnloop.Validator), a refused call is kept for the run; the ADR says only what the user asked; files split to the size limit Only the refusal of a call is kept and escalated, the same in the three mocks; the runners' loop and run files are no longer touched. Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex schema: spawn_agent's task_name and fork_turns, recorded in a function call, are declared Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * toolspec: codex is wired by the codex batch with the options it implements; the ADR keeps only the user's decision The transitional codex schema accepted exec_command options the mock does not read, which adr/fail-fast-unimplemented forbids; removing them would refuse every recorded call. codex-mock does not validate until batch/codex declares its exec options on the shared schema (turnloop.Validator on turnHost and subHost). Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr(tool-calls-validated): name the mocks, the schema package and what a failing call does Sloprail-Cites-User: Every mock (claude, codex, cursor) checks each tool call its scenario script asks for against that harness's recorded tool schema (internal/toolspec) before playing it, in replays and ordinary runs; a call that fails the check fails the run, except where a recording shows the real harness answering that call itself. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * toolspec review: Bash timeout leaves the claude schema (no recording grounds its answer), the answered-kind tests check the recorded call and answer, cursor's refusals belong to the run's main session, a refused claude run streams no result Sloprail-Cites-User: we should validate the tools that are passed by the script of the mock, right? Uh, so that the player that gets the script and then does something should should validate the tools. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ems (#230) * codex replay: hook payloads are sorted only within a group of concurrent hooks, the groups keep their order (audit #207 item 3) A continued Stop follows the Stop it continues (stop_hook_active is in the group key). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * replay: per-run values keep their key behind a placeholder instead of being dropped (audit #207 item 4) Core Rules.MaskKeys: the key stays, its value is <key>; the codex rules mask usage, model, the transcript paths, ids and timings, and only the mock's own script parameter is dropped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: spawn_agent answers at once, add wait_agent (replay-driven) The mock's spawn_agent always answers {agent_id, nickname}; the sub-agent runs as a background task on the session's task registry and wait_agent (hooked multi_agent_v1wait_agent) waits on it through the core's subagents.Wait (first target to finish, statuses completed/running/not_found, an empty status on timeout). The adapter maps tools.multi_agent_v1__wait_agent to a unified ToolWait; the replay's scratch repository is on branch main; the old background parameter is refused. Six recordings leave the exception list. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: the mock-not-modeled deviations #187 makes false are removed or corrected The mock's spawn_agent answers at once and wait_agent is modelled, as the recordings show (foreground-subagent-result, nested-subagents-nowait, task-stream-frames, background-agent): the deviations saying the mock waits inside spawn_agent, has no wait tool or takes a background parameter are removed or reworded. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: foreground-subagent-result's statement is harness-neutral The dispatching agent receives the foreground sub-agent's final result: Codex provides it with spawn_agent and wait_agent, so the cell stays supported; the Claude trailer wording leaves the statement. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: task-stream-frames/codex discloses the recorded gap, no frame for a background task A harness-lacks deviation backed by runs/task-stream-frames (item_4: the shell command asked to run detached is an ordinary command_execution). Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex replay: the model's answers are steps of the script, played in recorded order A turn a Stop hook continues is followed by the model's next steps; the unified agent keeps Final (the claude adapter reads it) and answers among the calls are ToolAnswer steps. stops leaves the exception list; hook-exit-codes keeps its entry (it is green on macOS and red on Linux: its recorded command is ls of a missing path, whose message and exit code are the host's). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: an exec_command option the mock does not implement is refused, not ignored (audit #207 item 6) A working directory other than the run's, and a command whose output is longer than max_output_tokens (Codex truncates what the model is told; the mock does not), are refused loudly. A tty is carried out as far as the recordings show (lines end in CRLF); shell and login only choose the shell, which the mock does not model (every scenario command is run by /bin/sh). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: the hand-written replay test goes; the generated replay is what hook-exit-code-semantics and pretooluse-refusal are proved by (audit #207 item 9) TestReplayOfRecordedRuns replayed stops and hook-exit-codes with its own normalisation, deleting fields and taking its commands from the recording, while the generated replay (which compares the whole stream and payloads) listed them as failing. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * replay: compare with every sample; two empty event streams are not a green replay (audit #207 item 12) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex replay: sub-agent rollouts are matched by thread id, leftovers and unanswered scripts are refused (audit #207 item 5) A second commentary before one call (the first was silently overwritten) is refused; each spawn_agent is matched to its sub-agent's rollout by the receiver thread id its stream item names (not by position) and a rollout no spawn names is refused; a script with no recorded output, an output of no script, and a failed script that made several calls (which ran is not recorded) are refused. A second final answer is not an overwrite any more: answers are steps. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex replay: a method call on a value the adapter does not follow is refused unless it only looks around; wait targets are filled in structurally (audit #207 item 11) The evaluator accepted any call on an opaque receiver; only the methods the recorded scripts use to look at what they were given (filter, test, stringify, includes, toLowerCase) are allowed. The generated script no longer rewrites its own text with sed (a literal AGENT1 or IDPLACE in a recorded string was hit): wait targets are {spawned: k} objects that jq replaces by the k-th spawn receipt, and the call id is set by jq; the limit of nine spawns is gone. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * wait_agent: a background command of the session is not a sub-agent to wait for Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: an exec_command shell other than zsh is refused (audit #207 item 6) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: an exec_command runs by the shell it names, zsh -c or zsh -lc, and a missing zsh is refused (audit #207 item 6) The mock ran every command by /bin/sh whatever shell and login the call named. It now runs it by zsh (-c, or -lc for login:true) as the harness does; zsh not being installed, or a login with no shell, is refused loudly, never replaced by /bin/sh; a shell other than zsh is still refused. The codex e2e CI jobs install zsh. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * tests: the zsh e2e test fails when zsh is absent instead of skipping Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * module map: exec_options.go with the tools; spawn receipts are parsed in the subagents module (module-coverage, module-leaks) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: the codex mock-not-modeled deviations the recordings and the mock prove false are corrected The mock fires SessionEnd, SubagentStart, SubagentStop, PreCompact and PostCompact, and reads a SubagentStop hook's block/exit 2 (runs/subagent-stop-block-loop): the hook-exit-code-semantics and hook-matcher-filter deviations that said otherwise are corrected, the subagent-lifecycle-hooks one removed. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * ADR tests-fail-on-missing-tool: a test whose required tool is missing fails, it never skips Sloprail-Cites-User: Fail, never skip Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * tests: a missing required tool fails the test (capcells jq/yq/git/bash, python3, the real claude when asked for); CI installs python3, jq and yq Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: hook-exit-codes re-recorded with a portable failing command; its exception entry goes Only the ls of a missing path (whose message and exit code are the host's: macOS 1, GNU 2) is replaced, by false; every other command, hook and step of the scenario is unchanged. Recorded with the pinned codex 0.159.3 (capture.sh); the old sample is dropped. The capture re-froze the plugins doc page's hash in MANIFEST. The replay is green on any host. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #230: wait_agent refuses bad input, a failed sub-agent start is errored not completed, the script fails loudly on a missing receipt, reasons and comments wait_agent decodes strictly (an unknown field or a wrong type is a failed result naming it); a sub-agent whose rollout could not be created is told as errored with why; the generated script errors when a wait's spawn position has no receipt; three notReplaying reasons name wait_agent; the exec_options and wait comments are reflowed and true. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: noninteractive-run-output-schema re-recorded without --output-schema; its exception entry goes The mock refuses --output-schema (it implements none of it), so the recording is the plain run: no args, no schema.json, recorded with the pinned codex 0.159.3 (capture.sh), the old sample dropped. The cell's mock-not-modeled deviation no longer cites the run for a schema-shaped message it no longer shows. The replay adapter writes no hooks.json for a run with no hooks (an empty file is not a hooks file). Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: subagent-transcripts-v2 re-recorded in the default multi-agent mode; its exception entry goes The mock refuses --enable multi_agent_v2, so the run is recorded without it (args removed, the prompt asks for the sub-agent and wait tools; the name is kept: runs are not renamed). The sub-agent's path in the tree of agents is null in the default mode, as the mock leaves it, and the final answer reaches the session also as a subagent_notification message, which the mock does not record: the cell's two deviations are corrected to the recording and the test follows. Recorded with the pinned codex 0.159.3 (capture.sh). Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: the session is told when a sub-agent it started has ended, as the recording shows A <subagent_notification> user message (the sub-agent's id and how it ended: its final answer, or why it failed) goes into the owner's rollout after the tool output it was just given, once, whether or not the agent waited for it (recorded: runs/subagent-transcripts-v2). The cell's deviation for the gap is removed, the mock having closed it; a test compares the notification, its place and its content with the recording's. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: noninteractive-run-output-schema is restored as recorded, with its exception entry and its spec sentence The re-recording without --output-schema lost the recorded schema behaviour (the real run's message conforms to it), which the cell cites. The original run, its setup (args, schema.json) and sample are back, the plain re-recording's sample dropped; the replay adapter keeps writing no hooks.json for a run with none. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * review of #230: a failed sub-agent start fails the wait fast, no unrecorded errored status; a runner test; the v2 name's history noted The core wait no longer has an errored status (no recording shows its shape): a wait that names a sub-agent which ended without an answer returns subagents.AgentFailed, the codex wait_agent fails with a clear error saying the mock refuses it rather than inventing it, the failed start is also reported on stderr, and no notification is written for it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * ADR tests-fail-on-missing-tool: only what the user's words decide (a missing tool fails the test, never skips) Sloprail-Cites-User: Fail, never skip Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: subagent-transcripts/codex: the missing sidecar is the harness's lack, not the mock's Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * sub-agent end notices: which sub-agents and when is the core's (subagents.TakeNotices), the codex adapter only encodes, marked sr:provides Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: a wait for several sub-agents, recorded: it returns at the first, the others are told of one at a time (runs/foreground-subagent-wait-many) The recording shows what the doc does not: a wait naming two sub-agents returns at the first to finish, lists only it (not a target still running, and the stream's completed wait item names only it), and the second is told of by a notification of its own after the agent's answer, which goes on. The mock now: lists only finished targets; tells the session of one ended sub-agent at a time, at the point the model is asked again (not between the calls of one script, the adapter marking those with the mock's own parameter more) or when the turn would end, which then goes on with it (a running sub-agent gets answerLatency to end first, one just started takes modelCall to start: the mock has no model to take that time); a sub-agent the recording shows unfinished hangs instead of ending; the replay evaluator reads const [a, b] = await Promise.all([...]). The harness-lacks deviation for the doc's all-wait is added to the cell. Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: the SubagentStop variants (continue:false over a block, plain text ignored) recorded and tested; the SessionStart matcher is tested on the resume and fork sources Two new runs recorded with the pinned codex 0.159.3 (subagent-stop-continue-false: of two SubagentStop hooks one blocks and one says continue:false, both ran and the sub-agent was not run again; subagent-stop-plain-text: plain stdout on exit 0 is ignored); both already replay green, so the mock models them, and tests compare their hook logs with the recordings. The block-loop test is also marked as proving the exit codes. A test drives per-source SessionStart matchers on a fresh, a resumed and a forked session (doc-grounded: hooks#sessionstart). The hook-exit-code-semantics cell lists the three SubagentStop runs. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * the tests-fail-on-missing-tool ADR leaves the batch (it stays on side/tests-fail-on-missing-tool); the skip-to-fail conversions and the CI tool installs stay Sloprail-Cites-Tool: no script or other rule enforces it. Either widen the rule's match to include test files (and any non-Go test scripts) or link a rule that checks tests for skip-on-missing-tool. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(scenario): a turn may carry a gate the loop holds it back for An assistant line's "gate" names what must have happened before the turn's call is taken, so a script orders agents by events, not by delays. turnloop calls the host's Gater between the turn and its call. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(codex-mock): order agents by script gates, not by invented durations The 1s model-call and 4s answer-latency delays are gone. The replay adapter derives each step's gate from the recording's times; the mock counts calls started (hooks fired) and finished per agent and releases a gate by the event it names. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(turnloop): the loop decides when the agent is told of what ended A host may be a Noticer: the loop tells it after the last call of a model's script (a call's More marks that another follows), and instead of ending the turn, with no Stop hook run and no block counted. Host interfaces move to host.go. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * refactor(codex-mock): the core decides when a sub-agent's end is told The adapter keeps only the notification's text and the user message it is recorded as; the Stop hook's stop_hook_active comes from the loop's continuing. A call's `more` is a field of the call, not of its input. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): refusing apply_patch, ask/allow leaving no record, and compaction/sub-agent matchers Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: the wait-many deviation names the doc section and quotes its sentence Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * revert: the mock-not-modeled corrections wait for the user's words (patch handed back) Sloprail-Cites-Tool: file-guards evaluated in 15m7.673s (31 rules, concurrency 6) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec(codex): narrow three mock-not-modeled deviations the recordings prove false The codex mock fires PreCompact, PostCompact, SubagentStart, SubagentStop and SessionEnd, matches on the compaction trigger and sub-agent type (runs/subagent-stop-*), and refuses apply_patch by a PreToolUse hook; the deviations now name only what is still not modeled. Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: transcript_path on every event, tool hooks and SessionEnd included, recorded (runs/session-transcript-file-hooks) and tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: session-transcript-file/codex lists the hooks recording and drops its kind-less deviation Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * refactor: the agents' gate ordering (progress, spawn log, hold) lives in internal/subagents The codex adapter only reports a gate's mistakes in its own words. progress.go joins the subagents module's home. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): several hooks' added context, SessionStart context after a compaction, PostCompact matcher marker Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: a PostToolUse block with added context, recorded (runs/posttool-block-context); context on a resumed session tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-additional-context/codex lists the block-with-context recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(replay): the codex adapter replays later runs of the harness (then-NN resume and fork) A recording's later runs are read into core.Recording.Then: each run's records are cut from its thread's rollout by its own prompt (a resumed thread holds the runs one after the other, a forked one a copy of the history first). The mock is run once per run against the same CODEX_HOME, from the run's directory, with <SESSION> the first run's session, and each run's script starts after the steps its session already holds. session-resume and session-fork leave notReplaying. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: SessionStart context on a resumed session, recorded (runs/session-resume-context) and tested against the recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-additional-context/codex lists the resumed-session recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(replay): the codex adapter replays a session the harness compacted on its own A compacted record becomes a compact step of the script (the mock has no tokens: the run's token-limit option is the one setup args installed), and a compaction spends a step. manual-compaction-auto and session-start-compact-continue-false leave notReplaying; the two runs a hook stopped keep their entry with the reason now true. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: context after a mid-turn compaction and a /clear prompt, recorded (runs/manual-compaction-auto-context, runs/session-start-clear-prompt) and tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-additional-context/codex lists the compaction and /clear recordings; source clear is the harness's lack Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(hooks): hooks run in the background (Later) deliver their context at the next safe point The agent does not wait for a hook marked async. Background hooks run one after another in the order started; what they printed is due once the agent has gone through the call after the one it had begun when they started, or when its turn would end, and delivery waits for the hook to finish: an event, not a time. turnloop's Noticer returns every notice that is due. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: async hooks deliver late, login:false with no shell is the default shell SessionStart, UserPromptSubmit and PostToolUse hooks marked async run in the background and add their context when it is due (recorded: runs/hook-async-context); their logs land in whatever order they finish, which the replay does not compare. login:false with no shell is accepted, as the real harness accepts it (recorded: runs/exec-login-false); login:true with no shell stays refused. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: async hooks' late context recorded twice (runs/hook-async-context, runs/exec-login-false) and tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(hooks): one sr:capability marker for hook-additional-context Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): apply_patch matcher aliases and the SessionEnd matcher by the binary; sub-agent stop variants marked for the lifecycle capability Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: bash-tool-result/codex lists the login:false recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * refactor(hooks): one core function (ContextDue) decides when a hook's context reaches the agent The sync path (a hook the agent waits for) and the background path (Later.Due) both ask it, and the one sr:capability hook-additional-context marker is on it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): a failed command with no output is told as an empty output, and PostToolUse names its PreToolUse's tool_use_id Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: a command that ran to its end is told framed, as recorded; a user's SIGINT interrupts the turn The agent is told "Script completed / Wall time / Output:" then the output, for a failure as for a success (recorded: runs/shell-exit-status). SIGINT during a turn fires the Interrupt hook, answers the running call "aborted by user after <time>", tells the agent the user interrupted, records the turn as aborted, prints nothing more, fires no PostToolUse or Stop hook, still ends the session and exits 1 (recorded: runs/interrupt-hook). The replay adapter sends SIGINT once the command has started. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: an interrupted turn recorded (runs/interrupt-hook, capture.sh interrupt-after) and tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-matcher-filter/codex lists the interrupt recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: Interrupt hook answers, PreToolUse exit-1 stderr and SessionEnd steering recorded and tested; the replay starts its child through procexec runs/interrupt-hook-output (what an Interrupt hook answers cannot prevent the interruption or be shown), runs/session-end-steering (SessionEnd hooks cannot steer the run), and the PreToolUse exit-1 stderr case of runs/hook-unstartable. procexec.Spec gets Stdout and OnStart so the interrupted replay no longer calls exec.Command. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * revert: hook-matcher-filter's interrupt-hook run waits for the user's words (patch handed back) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec(codex): the mock fires Interrupt; cite the interrupt and session-end-steering runs runs/interrupt-hook shows a SIGINT during a turn fires the Interrupt hook, and the mock now does the same, so the deviation saying it fires none is corrected; the new runs are listed in their cells. Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: the SessionEnd failure-report deviation begins Doc and recording conflict, names the section and quotes it Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(turnloop): interrupt.go joins the turn module's home Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * revert: the Interrupt work moves to side/codex-interrupt (it contradicts adr/modeled-surface) Mock SIGINT handling, the Interrupt hook, the interrupted replay, runs/interrupt-hook and runs/interrupt-hook-output with their tests, capture.sh interrupt-after and procexec's Stdout/OnStart go with it. session-end-steering, the framing of a command's result and the other fixes stay. The deviation in hook-exit-code-semantics.yaml is handed back (patch outside the repo). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec(codex): Interrupt stays not modeled on this batch (its work is on side/codex-interrupt) The Interrupt modeling moved to side/codex-interrupt pending the user's decision on adr/modeled-surface, so the deviation names Interrupt as not modeled again. Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec(codex): cite the runs that prove the mock fires compaction, SubagentStart and SessionEnd The narrowed exit-code deviation now lists the recordings that show those events firing (subagent-start-refused, manual-compaction-auto-blocked, manual-compaction-auto-post-stopped, session-end-hook-failure); the Interrupt wording is main's again. Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(tasks): a gate waits for a sub-agent's task to be registered by an event, not by polling Registry.Await wakes on Add; awaitEnd no longer spins on Find. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(scenario): a gate with a field it does not have is refused Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(codex replay): an exec_command option the mock does not implement is refused, not ignored Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test: the gate test no longer claims the turn loop; the SessionEnd failure test shows no report; the SubagentStart exit-2 test is marked for the exit-code capability Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: a sub-agent that ends with no message is null, SubagentStart's systemMessage is not shown, tool matchers on spawn_agent and wait_agent (recorded and tested) runs/subagent-start-systemmessage and runs/subagent-stop-no-message; the replay reads trim as a look-around method; a sub-agent with an empty final message stops with last_assistant_message null and is completed with null in the wait. events/collab.go and filechange.go leave the scenario home for the subagents and tools modules. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: subagent-lifecycle-hooks/codex lists the systemMessage and no-message recordings Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: a call a hook rejects throws the hook's reason as a script error, framed as recorded PreToolUse refusals and PostToolUse blocks are told "Script failed / Wall time / Output: / Script error: <reason>" (runs/posttool-block, runs/pretool-decisions, runs/stops), whole, only the time masked. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): a PreToolUse hook refuses an exec_command with its options Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): the exec_command refusal and the resume/fork matcher tests are marked for the capabilities they prove Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): a SubagentStop matcher that matches the sub-agent's type runs Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): the null last_assistant_message is present and explicit, recorded and mocked Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(codex-mock): the Interrupt work returns: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1 Brings back, from side/codex-interrupt (undoing de9ec6e): the mock's SIGINT handling, the Interrupt hook, the interrupted replay with procexec's Stdout and OnStart, runs/interrupt-hook and runs/interrupt-hook-output with their tests, and capture.sh interrupt-after. adr/modeled-surface carries the exception; the deviation in hook-exit-code-semantics counts Interrupt as modeled. Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): the exec option test does not depend on zsh being installed Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex-mock: an Interrupt hook over 3 seconds is clamped with the recorded warning; SubagentStop's continue:false is read by the hooks module runs/interrupt-hook-timeout (the warning "clamping Interrupt hook timeout to 3s in <hooks file>"; hooks over their limit still ran to their end). agent_stop.go goes: hooks.Interpret(SubagentStop, o).Halt decides. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-timeout/codex states the Interrupt timeout conflict between the docs and the recording Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr(modeled-surface): the abort bullet states the exception once, with the user's sentence, and no longer says a mock never interrupts Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * codex: a SubagentStop hook's exit 2 beats its continue:false, recorded (runs/subagent-stop-exit2-continue-false) and tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-exit-code-semantics/codex lists the SubagentStop exit-2 recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(codex-mock): --ephemeral: no session file is kept and every hook payload says transcript_path null, replayed from the stream The mock keeps the session in a scratch directory for the script and removes it; the replay adapter reads a run that kept no rollout off its event stream (commands and one answer) and passes --ephemeral to the mock. runs/ephemeral-no-transcript is the recording. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-common-payload and session-transcript-file (codex) list the ephemeral recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): the Interrupt clamp warning is asserted, and resuming a run by id is marked for the noninteractive capability Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr(modeled-surface): the Interrupt exception is its own modeled bullet, the user's sentence verbatim Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): --ephemeral is implemented, so it is no longer among the refused flags Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * refactor(codex-mock): the ephemeral session's home and transcript live in ephemeral.go (files back under 150 lines); tests assert more TestSubagentStartCannotRefuse asserts the exit-2 reason is surfaced nowhere; the ephemeral test is marked for noninteractive-run and the SessionEnd timeout test for hook-timeout. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr(modeled-surface): the Interrupt rule is its own decision, outside the out-of-model list Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * revert: hook-exit-code-semantics' SubagentStop exit-2 run listing is re-added with the user's quote (patch handed back) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec(codex): drop deviations the recordings prove false (sub-agents, --ephemeral); cite the runs runs/background-agent shows sub-agent payloads with an agent_id, runs/ephemeral-no-transcript shows --ephemeral modeled; the SubagentStop exit-2 run is listed again. Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): an ephemeral run leaves nothing on disk: no sessions under the configuration directory and no scratch directory Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr(session-persistence): a mock persists a session exactly as the real harness does for the same flags Sloprail-Cites-User: our system relies on persistent traj - so mock should work same way as real one, i suppose? unless both aaxctually persist it Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: session-transcript-file/codex states the --ephemeral case: no rollout, transcript_path null on every event Sloprail-Cites-User: A harness-lacks deviation, or an unsupported cell backed by a recording, needs no quote from me. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): the background command's death is read from its process, the sub-agent payload test is marked for hook-common-payload, lifecycle_recorded_test.go is back at 400 lines Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr(modeled-surface): the abort out-list entry excepts codex-mock, and the Interrupt decision names its code Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): a sub-agent's hook payloads carry its own transcript_path and the run's cwd, recorded and mocked Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr(modeled-surface): the abort out-list entry restates the user's rule (unless a recording drives it) and points to the Interrupt decision Sloprail-Cites-User: The codex mock models Interrupt: a SIGINT during a turn fires the Interrupt hook, aborts the turn and exits 1, as runs/interrupt-hook records; a mock may interrupt a running tool when a recording drives it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): the SubagentStart/SubagentStop payload test is marked for hook-common-payload Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(codex-mock): the sub-agent start/stop payloads carry the parent session id and the run's directory, recorded and mocked Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * fix(codex-mock): --ephemeral=false is read as a value; the PostToolUse guard's two conditions are separate Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * feat(codex-mock): validate a script's tool calls against a codex tool schema (adr/tool-calls-validated) Needs batch/core (turnloop.Validator, internal/toolspec): rebase onto it, and add codex-mock/internal/runner/toolschema.go and validate.go to internal/toolspec/module.yaml's home. Schema: Bash (exec_command: command, workdir, yield_time_ms, max_output_tokens, shell zsh, login, tty), spawn_agent (message; script is the mock's own), wait_agent, apply_patch, each parameter grounded in a recorded run (replay.RolloutCalls reads the model's calls out of every rollout). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * toolspec module: the codex schema and validator wiring join its home Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * test(codex-mock): the old spawn background parameter and the test-only task_name are refused by the schema Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * fix(codex-mock): an ephemeral session keeps no sub-agent rollout either; run.go back at 150 lines Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…eleted (#233) * feat(rules): snapshots-read-only lets a run directory that is not on main be deleted; only runs on main are read-only Sloprail-Cites-User: A run directory that is not on main may be deleted; only runs already on main are read-only. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(rules): tests-fail-on-missing-tool: a Go test whose required external tool is missing fails; only an A10N_*_TEST gate may skip Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): snapshots-read-only's not-on-main case asserts the gate's own permitted decision on the deletion Sloprail-Cites-Tool: the permit of t1 (run not on main, which reaches refuse.sh and exits 0) is never asserted on a GateChecked event Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): snapshots-read-only refuses a Write to a recorded sample and lets setup/ and capture.sh be written Sloprail-Cites-Tool: Add a case where the agent writes or edits a file under snapshots/ (refused, reason asserted), plus a permitted write to setup/ or capture.sh. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): the tests-fail-on-missing-tool ADR links only its own rule; the snapshots-read-only Write case is named for what it does Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip. Sloprail-Cites-Tool: That link therefore enforces nothing for this ADR and contradicts it Sloprail-Cites-Tool: the folder name claims 'write-and-edit' but its agent.sh only issues Write tool calls Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(tools): skipcheck refuses a (*testing.T) skip in a function that reaches os/exec.LookPath, resolved by types; only an A10N_*_TEST gate is exempt Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): tests-fail-on-missing-tool calls tools/skipcheck instead of matching Go text; a build or run failure is an error; cases for the gate-with-error, far-skip, alias, dot-import, method-value and helper bypasses Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): each tests-fail-on-missing-tool case asserts the rule's own outcome on the events in its own test.sh Sloprail-Cites-Tool: test.sh asserts no event of its owning rule tests-fail-on-missing-tool Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): skipcheck refuses every testing skip in a package that reaches os/exec.LookPath anywhere (variables, TestMain), checks root-level tests, and errors on a missing or empty directory; cases for the root test and the checker that cannot build Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * docs(adr): tests-fail-on-missing-tool states that a test package that looks a tool up calls no skip outside the opt-in gate, which is what skipcheck enforces Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip. Sloprail-Cites-Tool: Checklist: (1) a test whose LookPath tool is missing fails and never t.Skips for that Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * docs(adr): tests-fail-on-missing-tool states the package-wide skip rule in the user's words Sloprail-Cites-User: In a test package that looks up an external tool, no test may skip unless the skip sits behind an A10N_*_TEST gate. Sloprail-Cites-User: A Go test whose required external tool (looked up with exec.LookPath) is missing fails; it never calls t.Skip for that. Only an explicit opt-in environment gate (A10N_*_TEST) may skip. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): skipcheck judges a package with its _test package, counts interface and type-parameter skips, refuses a test that sets an A10N_*_TEST name; header documents what counts as a lookup Sloprail-Cites-User: In a test package that looks up an external tool, no test may skip unless the skip sits behind an A10N_*_TEST gate. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * fix(tools): skipcheck drops the A10N Setenv check the ADR does not decide, and is split to stay under the file-size limit Sloprail-Cites-User: In a test package that looks up an external tool, no test may skip unless the skip sits behind an A10N_*_TEST gate. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * test(rules): the snapshots-read-only cases assert the gate's actual reason, not the capture.sh pointer every refusal carries Sloprail-Cites-Tool: none asserts the actual reason Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
* cursor: replay every recording through the mock's own replay command (#158) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor replay: review fixes (named gaps in the reasons, temp dir cleanup, empty hook log not green, print-script exit 2, multi-sample test) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor session-start-hook: cite runs/session-resume; drop proves markers from internals unit tests The no-start-hook-on-resume claim rests on runs/session-resume, now in the cell's runs. The decision_test.go markers for pretooluse-refusal and hook-timeout sat on unit tests of Interpret; both cells keep their e2e proves. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: background-agent discloses the subagentStart/Stop doc and recording conflict; print-waits quotes the subagentStop sentence (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): hooks-all-matching-run/cursor cites the combined-refusal run and says the refusals are merged Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec: drop the cell note (notes field removed in #176) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): stop-block-cap/cursor cites the hooks doc's loop_limit default and discloses it conflicts with the print-mode recording (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): empty-tool-result-placeholder/cursor cites the exit-0 empty result background-bash-start records ls exiting 0 with empty stdout: stream empty strings, hook {"output":"","exitCode":0}, no transcript tool result. Stays unsupported, with the sharper reason. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): manual-compaction/cursor cites the recorded auto compactions and says what they show The cell cited only the headless /compress run. runs/compaction-transcript-continuity records two real compactions with a preCompact payload (trigger auto, no hook after, nothing stopping them); the hooks doc calls preCompact observational. Stays unsupported: the request, the hook after and the stop are all absent. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(spec): manual-compaction/cursor says no compaction hook fires on /compress Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(cursor): transcript user record opens with the empty timestamp element; transcript cells say what the recordings show (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec: tidy the transcript-file deviation wording (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor): pin the mock's transcript_path at the first preToolUse and beforeShellExecution (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec: disclose the doc's null-if-disabled transcript_path against the recordings (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec(session-transcript-file): codex's recorded file-before-start behaviour is the mock's too, not a deviation (part of #161) The entry had no kind and described behaviour the mock matches, so it declared nothing; the codex run still cites it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: task-stream-frames and hook-common-payload say what the recordings show for sub-agents Drop the stale 'mock has no sub-agents' deviation of hook-common-payload/cursor for a harness-lacks one the subagent-lifecycle-hooks recording grounds; prove the sub-agent half of task-stream-frames/cursor and the sub-agent payload identity with e2e tests. Part of #161. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: hook-common-payload discloses the sub-agent's own transcript path; task-stream-frames proves the background command's frames Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: disclose the subagentStart/Stop doc conflict; assert the recording has no sub-agent frame of its own Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: attribute subagentStart/Stop doc fields per event; assert neither fires Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: sub-agent payload test asserts Shell cwd is empty and no other event carries cwd Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: stream tests assert the recorded Task result and notification order on the recording too Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: hold the removal of the stale 'mock has no sub-agents' deviation until the user's words arrive Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: the task stream test asserts the recorded interleaved order of a turn's calls, for recording and mock Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: disclose the Task call's missing postToolUse as a doc/recording conflict and pin it Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor: drop the mock-not-modeled claims the recordings and tests prove false (the mock runs sub-agents) Removes hook-common-payload's 'mock has no sub-agents' and task-stream-frames' 'mock runs no sub-agent / Task answered as an error'; the true no-await-tool clause stays. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor-mock: time a file tool's call; postToolUse reports a positive duration as recorded (part of #161) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * posttooluse-payload/cursor: drop the false 'duration of 0 for a file tool' clause runs/file-tools records positive durations (45.586, 1.695, 0.759, 27.441 ms). Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor-mock: file-size split; postToolUse test no longer claims distinct tool_use_ids (the recording repeats one across a turn's calls) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor-mock: move Edit beside the file tools; toolexec.go back under the size limit Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(cursor-mock): fire beforeReadFile on each successful Read, as recorded (file-tools/cursor) The mock carried the beforeReadFile docs marker but never fired the hook, and the replay compare stripped it from the recording. It now fires after the Read's preToolUse and before postToolUse with file_path, content and attachments; the replay compares it; the cell drops its deviation. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor-mock): a read of an existing file fires beforeReadFile between preToolUse and postToolUse, as recorded Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * refactor(cursor-mock): move the tool Result types to result.go, under the file-size limit Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec(cursor): the mock fires beforeReadFile, so two deviations no longer say it does not Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor-mock): a beforeReadFile hook is matched on the tool name Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * snapshots(cursor): runs/before-read-refusal, beforeReadFile hooks blocking reads Six reads: exit 2, a JSON deny, invalid JSON, a crash (fails open), a failClosed crash, and JSON with an unknown permission. Recorded with capture.sh at the pinned cursor-agent 2026.09.28. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(cursor-mock): a beforeReadFile hook that refuses blocks the read, as recorded (file-tools/cursor) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec(cursor): beforeReadFile refuses a read as recorded, so pretooluse-refusal's deviation no longer lists it as unmodeled Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor-mock): the beforeReadFile refusal test also proves pretooluse-refusal/cursor Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * spec(cursor): pretooluse-refusal cites the beforeReadFile refusal run and doc Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor-mock): a blocked read gives the agent the failure hook's text, as recorded Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * fix(cursor-mock): the invalid-response wording is a beforeReadFile's alone, as recorded; the doc-only matcher test claims no cell Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * test(cursor-mock): a failed command's duration is the time it took Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx * cursor-mock: every tool_call frame carries toolCallId and hookAdditionalContexts (the after-tool hooks' context, on the completed frame) Recorded in every tool_call frame of runs/*; the contexts in runs/additional-context. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: Task frames carry the harness's args and the preToolUse payload is the call as the model made it; five recordings replay green subagent-lifecycle-hooks, foreground-subagent-result, subagent-stop-block-loop, subagent-transcripts and subagent-worktree-isolation leave notReplaying. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor-mock): the background Task test proves the receipt is at once, the recorded stream order and the payload fields (background-agent/cursor) The capability-rigor judge found the test only asserted a completed taskToolCall frame exists. It now also asserts the stream order of the Task call's frames, the end notice and the result against runs/background-agent, that the receipt precedes the agent's own command and the end notice follows it, and the sub-agent's afterShellExecution output and sandbox and the sessionStart and sessionEnd fields against the recording. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record a beforeReadFile hook timing out (failClosed or not) and a run with a second root added; replay tests runs/before-read-timeout (cursor-agent 2026.09.28): a beforeReadFile hook that outruns its 1 s timeout is killed and the read goes through; the failClosed one blocks the read with a postToolUseFailure saying it failed closed and timed out after 1000ms. runs/multiroot-workspace: a run with --add-dir still tells every hook one workspace_roots entry. The mock accepts --add-dir (ignored, as recorded). Both runs are cited by their cells and replayed by e2e tests. Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr: a mock replay replays every captured sample of a recorded run Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record --add-dir against reads inside it and outside every root; the mock ignores --add-dir because the recordings show no difference runs/add-dir-access and runs/no-add-dir-access (cursor-agent 2026.09.28): the same three Reads (a file of a sibling directory, a file outside every root, a file of the project), with and without --add-dir for the sibling. All reads succeed in both, with the same hooks, and hook payloads name one workspace root. The replay now maps <TMP> (the directory the workspace sits in) so such paths replay. Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr: replay-every-sample states the decision only Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: --add-dir is accepted only with --force, the mode its recordings cover The recordings of --add-dir (runs/add-dir-access, no-add-dir-access, multiroot-workspace) are all of -p --force. Without --force the mock refuses the flag, naming it, until a recording covers that. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: split decision.go and toolhost.go back under the 150-line limit (hooks/combine.go, runner/toolhost_after.go) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * module: toolhost_after.go is in internal/toolcall's home beside toolhost.go Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record --print alongside -p and test it; the depth-limit test asserts the recorded Task hooks and the call left in the transcript runs/print-long-form (cursor-agent 2026.09.28): -p --print is accepted and the run is as with -p. TestDepthLimitStopsNesting now also asserts the two recorded Task preToolUse hooks and that every recorded agent said it had no Task tool, and that the mock makes the Task call at the limit (it is in the transcript) with no hook seeing it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record matchers on Task and on beforeReadFile; the replay plays a recorded Task call runs/hook-matchers-task-read (cursor-agent 2026.09.28): a preToolUse hook matched on Task runs for the Task call and one matched on Shell does not; a beforeReadFile hook matched on Read runs for a read and one matched on Shell does not; the postToolUse hook matched on Task never runs (no postToolUse fires for a Task call). The replay helper plays a recorded Task call as a Task whose sub-agent only replies. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-common-payload/cursor says a sub-agent is told from the main agent only by its own ids Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr: take replay-every-sample out of the batch (kept on side/adr-replay-every-sample) The adr-well-formed and adr-grounded judges refuse the user-worded decision sentence: it has a gap clause (not checkable, not present tense) and names no replay command or check. Rewording it needs the user's words, so the ADR waits on its own branch. Sloprail-Cites-Tool: The only Decision bullet fails 'Checkable' and 'Present tense, current state only' Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: pretooluse-refusal/cursor discloses the agent_message doc and recording conflict Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: session-start-hook/cursor says its payload carries no start-kind field Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor-mock): the symlinked-cwd run proves the sessionStart hook fires once with no transcript path Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-matcher-filter/cursor drops the matcher run, to be cited with its capture Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-matcher-filter/cursor cites runs/hook-matchers-task-read, the capture of matchers on Task and beforeReadFile Sloprail-Cites-Tool: captured runs/hook-matchers-task-read/samples/20261005-213709 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record Grep and Delete calls with matchers on each; the mock models both, as recorded runs/hook-matchers-grep-delete (cursor-agent 2026.09.28): a preToolUse or postToolUse hook matched on Grep or Delete runs for that tool and not for the other, with the recorded hook inputs and outputs, grepToolCall and deleteToolCall frames with their results. The mock runs both tools; a Grep with no match and a Delete of a file it cannot read fail rather than guess (unrecorded). Sloprail-Cites-Tool: captured runs/hook-matchers-grep-delete/samples/20261005-215835 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: pretooluse-refusal/cursor lists the agent_message conflict after the existing deviations Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record an MCP tool call with matchers on MCP:<tool>; the mock calls the project's stdio MCP server, as recorded runs/hook-matchers-mcp (cursor-agent 2026.09.28, --approve-mcps, a local stdio MCP server written in sh): hooks name the tool MCP:echo; a matcher on it runs the hook and one on MCP:other does not; the agent first reads the tool schema in a getMcpToolsToolCall no hook sees, then makes the mcpToolCall. The mock starts the server of .cursor/mcp.json for the call and refuses MCP calls without --approve-mcps (unrecorded). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record a Task call with an invalid model; the mock answers it as recorded, with no sub-agent runs/foreground-subagent-failure (cursor-agent 2026.09.28): the Task call is not started, its preToolUse hooks fire, and it completes with "Invalid model selection ... could not be resolved to a valid subagent model" and the account's model list, with no after-tool hook. The mock answers the same for any model but default and lists only default. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: a background sub-agent launched by a sub-agent outlives it and its end is announced on the run's stream, as recorded; the nested background test asserts the recorded hooks, order and notification runs/nested-subagents-background: the launching sub-agent reports STARTED while the background one goes on, its command's hooks and the sessionEnd come after, and the main stream carries a task_notification naming it. The mock owned the background sub-agent by its launcher, so it was never announced. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor-mock): deny beats ask on beforeShellExecution and refusal messages are joined on beforeShellExecution and beforeReadFile Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: a blocked read's completed frame carries no args, as recorded; models inherit and default are the valid ones for a Task; before-read-refusal replays green Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor replay: setup prepare.sh and args are installed and passed (a flag the mock does not model is refused), MCP, Grep and Delete calls are mapped, a hook script's own log tag joins its event's group prepare.sh runs in the repository under the run's home as the capture runs it; setup/args go on the mock's command line, only the flags it models; a recorded CallDynamicTool is a mcp__<server>__<tool> call with the model's arguments (its lookup left to the mock); Grep and Delete are unified tools; the model catalogue after 'Allowed model slugs:' is not behaviour. A closed.sh-style log line ran concurrently with its event's payload, so it is grouped with that event: before-read-refusal no longer depends on timing. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor replay: the new runs enter the exception list with their gaps, six greens leave it session-resume-unknown replays green with args; no-add-dir-access, multiroot-workspace, print-long-form, hook-matchers-mcp, hook-matchers-grep-delete and foreground-subagent-failure are listed with the gap each shows. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: a Read's result names the file resolved, Grep, Delete and MCP frames carry their toolCallId in their args, as recorded; the print-flag and multiroot runs read a file, so they replay The new runs print-long-form and multiroot-workspace are recorded with a Read instead of a shell command; no-add-dir-access and hook-matchers-grep-delete replay green and leave the exception list. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: assistant text is shown as one frame when the next call starts or at the end of the turn, a refused Task completes with only its error; MCP calls carry the model's description and their own hook id; hook-matchers-mcp and foreground-subagent-failure replay green Recorded (runs/foreground-subagent-failure): a call refused before it started shows no frame at its call, so the text before it goes out with the rest at the end. The replay adapter reads the model-written mcpToolCall description from the stream. Both runs leave the exception list. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-matcher-filter and pretooluse-refusal back to main, to be corrected again in a commit that carries the citation Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: the beforeReadFile deviations of hook-matcher-filter and pretooluse-refusal match the recordings (the mock fires beforeReadFile); the agent_message doc conflict and the matcher runs are cited Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: the MCP server is run through internal/procexec; files split under the size limit and kept in their modules Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: the user's ~/.cursor/hooks.json is a hook source beside the project's, a hook in both runs once per source, as recorded runs/hooks-all-matching-run-same-hook-two-sources. The e2e helper installs a run's user-hooks.json in its home and reads hook logs whose concurrent writes ran together. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hooks-all-matching-run/cursor says the mock reads the user source beside the project source, as recorded Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: assistant and tool_call frames carry model_call_id and timestamp_ms (the text at the end of the turn neither), as recorded; tests for them and for the common payload fields Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: replay_test.go back under the file-size limit; out.go belongs to the scenario module with frames.go Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor-mock): the transcript file exists when the first beforeShellExecution names it, as recorded Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-exit-code-semantics and hook-common-payload say where cursor differs from their statements (exit 0 with invalid JSON blocks; transcript_path null with transcripts on) Sloprail-Cites-Tool: Cursor cell lists no deviation for exit 0 with invalid JSON on a permission hook Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: frame printing joins frames.go (no module change); the tests prove the sub-agent hooks that never fire, the notification detail as all the sub-agent said, and the transcript path on later hooks; hook-exit-code-semantics back to main for a cited re-correction Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-exit-code-semantics/cursor lists beforeReadFile among the events the mock fires and says an exit 0 with invalid JSON blocks, as the recordings show Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor replay: a hook log line of several objects run together is read as the objects (hooks of one event append to one log side by side) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record workspaceOpen firing in print mode; the mock fires it before the session start with the recorded payload runs/workspace-open (cursor-agent 2026.09.28): the hook fires once, first, with only the event name, the Cursor version, the workspace roots and the user email (null), no session, model or transcript path. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: the cursor cells say what the recordings show of workspaceOpen, multiroot workspaces and model_id/model_params Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-common-payload/cursor drops the model_id/model_params deviation: it is a mock gap (the thinking hook is not fired), not a harness difference Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * afterAgentThought: a script's thinking block is a thought of the turn (core), the shared loop tells a Thinker host before the turn's message and calls, cursor fires the hook with the model's fields and its own generation Optional: a host that does not implement turnloop.Thinker (claude, codex) never hears of a thinking block. A sub-agent and a turn that a finished background task starts are model requests of their own; the result frame keeps the run's first. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor replay: the thoughts a recording holds are fed to the mock, response by response, and no longer dropped The unified Call and Agent carry an optional Thinking; the adapter puts the i-th afterAgentThought of a conversation on its i-th response when the recording holds one per response, and refuses a recording that holds them for some only (no guessing, nothing dropped). The generation's number and name are masked, a thought and the preToolUse beside it are concurrent. nested-subagents replays green. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: a step's text and its call are one transcript record, as recorded; tests for the transcript's records, the foreground/background contrast; the cells say workspaceOpen names no session and a background Task is modeled Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor replay: a thought is put on the response its generation names, so a run whose model thought in some responses only is replayed (symlinked-cwd) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: the cursor cells list afterAgentThought and workspaceOpen among the events the mock fires, and cite the run that replays with thoughts Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: record matchers on afterAgentThought, workspaceOpen and sessionEnd with a thinking model; the mock tests a thought hook's matcher against AgentThought, names the run's model in the init frame and sessionEnd, and refuses models it has no record of runs/hook-matchers-thought (cursor-agent --model cursor-grok-4.5-high): the afterAgentThought hook matched on AgentThought and the one with no matcher run for each thought, the one matched on nothing does not; workspaceOpen and sessionEnd run every hook whatever its matcher. The sessionEnd payload and the init frame name the model the run was started with. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-matcher-filter/cursor cites the thought matcher run and discloses that no postToolUse follows a Task call Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: a Grep with any parameter but the pattern fails; unrecorded sub-agent models are refused as not modeled; --add-dir is gated on --force alone; the replay masks the clock and the service ids instead of dropping them, and tool_call frames say when a call began and ended Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-matcher-filter/cursor drops runs/nested-subagents, which has no matchers Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor: StrReplace, GetDynamicTools and AwaitShell are modeled from the recordings that use them; file-tools and nested-subagents-depth replay and leave the exception list, the other four are listed with what is left A StrReplace is an edit whose hooks see the whole file it makes and whose result carries the diff; a catalogue search finds nothing (the mock has no catalogue); a wait with no task named lasts as long as it was told. A hidden read before a write carries the write call's id. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: a catalogue search is answered only where a recording shows it finding nothing, any other is refused as not modeled Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: a refusal of something not modeled fails the run, a sub-agent's too; nested-subagents-depth is listed again, as its recording does not hold what the sub-agent's search answered Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: files back under the size limit and in their modules; a test that the result agentId is the sub-agent's own id and not the args' Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: the background Task test checks it returned at once; think.go joins the hooks module and refusal.go the scenario module Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr: each mock's replay test replays every captured sample of a recorded run Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor replay: the exception list is accepted as it stands and may only shrink Sloprail-Cites-User: Accept the cursor mock's initial replay exception list as it stands; from then on it may only shrink. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: foreground-subagent-result/cursor says the mock does not model a sub-agent failing mid-run Sloprail-Cites-User: The cursor mock does not model a sub-agent failing mid-run; no headless run produces one. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-common-payload says a sub-agent event carries an identity of that sub-agent: its own session id, or the main session's together with an agent id Sloprail-Cites-User: All lgtm Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor): pin the recorded afterFileEdit edits and a JSON allow on beforeReadFile; module for the refusal Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * adr: take replay-every-sample out of batch/cursor (it lands from side/adr-replay-every-sample with batch/rules-pin) Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec: hook-common-payload/cursor discloses the attachments entries the mock does not model Sloprail-Cites-User: The mock's payloads always carry an empty attachments list; the entries a real cursor-agent adds for an attached file or rule (a type and a file_path) are not modeled, since no headless run can attach one. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(cursor): a Task call at the depth limit is answered as an unknown tool Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(cursor-mock): refuse from the structured message, not a scan of printed frames; test the unrecorded-model, Grep-keys and --yolo --add-dir cases Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(cursor-mock): --add-dir is refused with --yolo too (only --force is recorded); test it Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * cursor-mock: declare the tools the mock plays (Edit, Grep, Delete, GetDynamicTools, AwaitShell, MCP tools) in the validated schema; toolspec: prefix and open tools Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * toolspec: a prefix tool may carry a Valid rule (mcp__<server>__<tool>, both non-empty); drop the unreachable Grep-keys path Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * cursor-mock: pin the thought's null transcript path and the file tools' result frames; split replay's measure out of canon.go Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ure, resume frames, rigor tests (#158, #207) (#229) * claude-mock: a run whose result is an error exits 1 and fires no Stop; recording runs/run-failure (an unrecognised model) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * snapshots: the docs re-frozen with the recording (capture.sh) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * run failure: the core says a failed result fires no Stop (scenario.Result.Failed); run-failure cited by noninteractive-run Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: a failing SubagentStart hook still hands back the report, driven by runs/hookerrors Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: an isolated sub-agent's hand-back trailer names its branch (worktreeBranch) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: the worktree trailer lines in agent_worktree.go (file-size) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: the hand-back tests compare the mock's output with the recording's Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: a terminated background shell is really gone (background-bash-reaped-at-exit) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: a foreground sub-agent's refused tool call is denied (foreground-subagent-result) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: -r and -c, the short forms of --resume and --continue Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: a notification turn ends with its own result frame (runs/bgagent) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: hook payloads sort only within one firing of one event, firings keep their order (#207 item 3) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: only the handlers' own log lines are sorted; payload lines keep their order (review) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: scoped drops (the mock's script key only from Agent inputs, assistant bookkeeping only when empty); no capability cell may name what is dropped (#207 item 8) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: init is also dropped after a compaction (comment, reasons); a safe type assertion Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude: stream frames carry a timestamp; the replay scrubs it instead of dropping the key (#207, replay-fidelity) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: timestamp is scrubbed, not dropped, so no cell excuse for it Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a session resumed by name, path or --continue streams its SessionStart hook frames under the lookup's own id Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: resume by name, fork transcripts, no-resume with names and forks Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: the unknown resume's payload names the id, file and directory; no-resume invariant marker Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: hook payloads carry prompt_id and permission_mode, in the recorded key order Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: prompt/permission decisions live in the core, notifications continue the prompt, compaction starts one; split files; no-resume shapes tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: prompting config in the hooks module home; one capability marker Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: refuse output formats it does not produce; tests for compaction exit codes, session ids of forks/resumes, unknown resume; module-home files Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: drop the text-format no-resume test (text output is refused now) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude: deviations for what the mock refuses (--agent, text/json, stdin, partial messages, SIGTERM 143) Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: fork hooks name the fork's transcript; no-resume equals an unknown session's refusal Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: a compaction streams only the compacted session's SessionStart hook frames (runs/compact) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: refuse --include-partial-messages and a piped stdin, as the deviations say Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: the result frame carries num_turns, stop_reason, terminal_reason, permission_denials, queued_turn_count, result_index, api_error_status, is_error, session_id Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: seven recordings replay green; removed from the exception list Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: the stream ends with a result frame (not a fixed field order) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: the result frame's fields are derived from the run (turns since the last result, refused calls, result index, failure), a refused call's result carries tool_result_meta; two more recordings replay green Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: flag helpers in flags.go (file-size); the batch's noninteractive-run script prose Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * hooks: blockers that finish together are acted on last-configured first (all-hooks replays green, deterministic) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * hooks: ActedBlock back to the recorded semantics (the blocker that finished last, by Done); all-hooks stays on the exception list as on main (its order is a race); a test over built outcomes Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: every sample of a recorded run is replayed, each with its own model turns Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * recordings: all-hooks-close-first and -second, two blocking hooks finishing 75 ms apart (A last / B last); docs re-frozen by capture.sh Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: compaction payloads carry the common fields; one hook in two settings files runs once Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: the idle ceiling stops the agent's process; the default ceiling is ten minutes Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a sub-agent streams its prompt, tool calls and results with parent_tool_use_id, and a task_progress frame before each call (runs/isolated-worktree, fg-subagent-bash); replay scrubs end_time Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: refuse --input-format and --max-budget-usd rather than ignore them; --max-turns and --include-hook-events refused as unknown flags (tested) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude: background-bash-reaped-at-exit declares the refused piped stdin Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: the foreground agent's task frames exclude its progress frames; bgbash replays green and leaves the exception list Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * hooks: ActedBlock by finish time: the last finisher, of those finishing within 50 ms the last configured, grounded in the all-hooks recordings; all-hooks replays green and leaves the list Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: --max-turns ends the run with the error result, no Stop, exit 1 (recording runs/max-turns); recording runs/include-hook-events (not implemented: the flag stays refused) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * recording plugin-hooks-same-command (a plugin's copy of a hook stays separate from the project's: two runs); test proves the mock does the same Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: setup args (modelled flags passed on, refused flags a by-design skip) and prepare.sh are replayed; recorded exit 0 or 1; max-turns, plugin-hooks-same-command, plugin-hooks, transcript-retention replay green and leave the list; the three new entries are not added Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: Read calls, a recorded API failure and the run directory in calls are replayed; a successful non-Bash tool result streams without is_error (recorded: 83 results); tool-errors, tool-invalid-input, matcher, run-failure and all-hooks-close-second replay green and leave the list Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: all-hooks-close-first dropped (a race: A, B, A; recorded in the ActedBlock grounding); a refused-flag recording is asserted refused by the mock, not skipped; a run with no sample is not a recording Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * batch/claude: new runs cited by their cells; denormalize.go and stream_tool.go split (file-size) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: unit tests follow the mapped Read tool and the modelled --model flag Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: worktree hooks ignore their matcher; PostToolUseFailure matches the tool's name Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock e2e: compact-nohooks output, startup hook frames, unknown resume on stderr and stdout (RunSplit); close-first re-captured, replays accept a recorded outcome where samples disagree (Reconcile) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude: hook-two-settings-files recorded (the same hook in two settings files runs once; replays green); bgbash after-result order tested; the new runs cited by hooks-all-matching-run Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude: noninteractive-run's refused list names --agent too, as the user's words do Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a command hook's args run its program directly with them, shell is accepted and ignored (recordings hook-args, hook-shell); a test per recording; the hook-two-settings-files test proves hooks-all-matching-run too Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: --include-hook-events streams every main-thread hook's frames (PostToolUse ahead of the result frame; recording runs/include-hook-events replays green); a hook shell other than the recorded bash is refused Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: tests that a sub-agent's own PostToolUse context and a resumed session's SessionStart context are added (recording subagent-post-ctx); replay leaves a sub-agent's spend (tokens, duration, resolved model id) out of the comparison, so 6 recordings replay green and leave the exception list Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * replay: a sub-agent's spend is masked, not dropped: its keys stay compared, their values and the usage line's counts become <MASKED> Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: runs of several steps (a resume) against one session store, so recording resume-session-start-ctx replays green and backs the test that a resumed session's SessionStart context is added; the mock no longer streams that context as a frame (the real stream has none); a resumed start's measures are masked capture.sh also takes later steps as flat files (then-NN-prompt.txt, then-NN-args): the structure gate allows no deeper path under setup/. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * hook-additional-context/claude cites runs subagent-post-ctx and resume-session-start-ctx Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: e2e tests of prompt_id, SessionStart matchers by source, fork/resume context, the failure payload's tool and input; RunSplit uses procexec; files under the size limit Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: hooks of a session with a scratchpad are told scratchpad_dir (recording scratchpad-dir; after cwd, as recorded), none carries effort; replay installs a setup's env; RunSplit is a test helper again (e2etest.go is untouched) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * hook-common-payload/claude cites run scratchpad-dir Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: A10N_MOCK_NO_RESUME names the session a forking --resume resumed, not the fork's new id; tested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a settings deny rule for an exact Bash command refuses the call as recorded (runs/permission-denied): permission_denied frame, the denial in the result, no PostToolUse; any other deny rule is refused Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: an unknown flag fails on stderr in claude's words before any frame (runs/invalid-flag); --bare and --agent are refused by name (runs/bare); the cell cites the runs Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: allowManagedHooksOnly in a settings file is refused by name, not ignored Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a scenario that streams system/api_retry is refused by name (no model API to retry; a retry cannot be recorded on demand) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * ADR refused-flag-replay: a recorded run made with a flag the mock refuses replays as a check that the mock refuses it Sloprail-Cites-User: A recorded run made with a flag the mock refuses replays as a check that the mock refuses that flag; it is not an exception-list entry. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * noninteractive-run/claude declares that system/api_retry events are not modeled Sloprail-Cites-User: The claude mock does not model system/api_retry events; it refuses them until a recording drives them. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a turn that ends while a background agent works streams its result with the later turn's, after the notification turn's init frame (recording bgagent); the flag-refusal wording in the sr-agent flags test follows the new unknown-option message Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: test that a background shell killed at exit is reported after the result frame, as recorded (runs/bgbash) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a sub-agent's task_progress names an Agent call by its description and a Read or Edit by the file, as recorded (runs/meta, file-tools); test of nested foreground agents' task frames against runs/meta Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: test that a failed background shell is reported failed, against the new recording runs/bgbash-failed; the cell cites it Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: test that an unknown resume ends its session before the result frame and that frame is all of stdout (runs/resume-unknown); the unknown-resume test keeps to the recorded form Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: test that the print wait for background agents defaults to ten minutes Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: a scenario that calls the Monitor or Workflow tool is refused by name Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude cells declare what the mock refuses or does not model: --bare, --agent and other deny rules (noninteractive-run); the workflow and Monitor tools and the wall-time ceiling (print-waits-for-background-agents) Sloprail-Cites-User: All lgtm Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: the refused Monitor and Workflow tool names cite the tools reference that lists them Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude replay: the per-user temp folder claude-<uid> is the same on both sides (CI runs under another uid than the captures); the background-bash unit test no longer races a command that ends at once Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * tasks.Results holds the results of turns that end while a background agent works (core, next to AwaitAfterTurn); the deny-rule refusal lives in tool_hooks.go with the tool call; the refused-flag ADR names the code it governs Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * spec(claude): hook-common-payload cites scratchpad-dir (runs in name order) The cell lists the scratchpad-dir recording among its runs, and keeps the --agent deviation the user asked for. Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them. Sloprail-Cites-User: So those should not be limited or what's the problem? As long as capability describes it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * claude-mock: hook_additional_context records carry the text as the agent receives it (rendered system reminder, renderedRole system), as recorded; one sr:capability marker for print-waits-for-background-agents Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * ADR refused-flag-replay: scope says every mock's replay test follows it Sloprail-Cites-User: A recorded run made with a flag the mock refuses replays as a check that the mock refuses that flag; it is not an exception-list entry. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * replay masks end_time instead of dropping it (the mock's task_updated patch carries it, as runs/isolated-worktree shows); the test checks the key Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * claude-mock: runner.go under the size limit (settings load and invoker setup in runner_setup.go); flagError notes only the long-flag wording is recorded; one import block in control_records.go Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * claude-mock: --include-hook-events streams PostToolUseFailure's and a foreground sub-agent's SubagentStart/SubagentStop frames as recorded (runs/include-hook-events-more), in raw print mode Stop's too; it is refused when the run reaches an unrecorded hook (compaction, worktree, background sub-agent) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * claude replay: prepare.sh runs with a claude that accepts only plugin commands (the mock reads a local marketplace itself; CI has no claude), so plugin-hooks and plugin-hooks-same-command replay there; the unknown-resume test's hook no longer sleeps near SessionEnd's budget Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * claude-mock: --include-hook-events also refuses the subagent_start control record, a scenario-written tool_result's PostToolUse and a background task's notification UserPromptSubmit (frames not recorded); only a manual compaction refuses on SubagentStop; the early-return sub-agent writes its held start frame; worktree and other cases in TestT001_19 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * claude-mock: stream_scan.go within the size limit; the noninteractive-run cell cites runs/include-hook-events-more Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * the deny rule's refusal is the core's (toolcall.DeniedByRule), the claude settings only parse their syntax; stream_turn.go within its ceiling; gofmt Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * the deny-rule decision sits beside noninteractive-run's one marker (internal/scenario), not under a second one Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…re, replay lists only shrink (#207), layering and snapshots-current fixes (#232) * fix(rules): module-boundaries fails closed on a failed module lookup; test injects it Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): module-coverage fails closed on a failed ADR or module lookup; tests inject them Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): modules.sh load_modules and module_home_files report a failed lookup Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): module-distinct fails closed and hands its modules to jq by file; test injects the failures Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): module-leaks fails closed and hands candidates to jq by file; test injects the failures Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): concern-undeclared fails closed and hands ADRs and modules to jq by file; test injects the failures Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): module rules read through here-strings, not a printf into grep -q or awk (a pipefail SIGPIPE race); cases name their rule literally Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): module cases show the recovery and the nearest permit; mock judges read their input Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): judge rules show a real refusal, its recovery and the nearest permit (mock judges decide from input) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): the new judge-rule cases name their rule literally Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-covered refuses when its lookups fail, with a case (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-grounded refuses when a lookup fails; the catalog goes through files, not argv; with cases (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-rigor refuses when a lookup fails; the catalog, pairs and markers go through files, not argv; with cases (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): one refuses-when-a-lookup-fails case per capability rule, asserting the FileGuardChecked refusal and its reason (#204) capability-covered's case asserts through the FileGuardChecked event; capability-grounded and capability-rigor fold their separate prepare/subjects/inputs cases into one sr-checks run case each, as capability-reconciled's cases do. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): capability-grounded's case injects each prepare lookup with a marker only prepare.sh carries Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): the ADR loader, linked-adrs and confine refuse on a failed lookup; ADRs go to jq by file, not argv Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): invariant-covered, -rigor and -grounded refuse on a failed lookup Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): adr-linked, -grounded, -well-formed and -matches-sloprails refuse on a failed lookup; large values go by file Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): each rule of the invariant and ADR slice refuses on an injected lookup failure Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): the slice's cases also run through sr-checks run; invariant-covered permits when implemented and proven adr-linked reads an ADR's sections with a here-string, not printf piped to grep -q: under pipefail a grep that exits early made printf fail, which read as a missing section. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): behavioural cases for the nine rules of the invariant and ADR slice Each rule gets a case beyond its lookup-failure one: a refusal with the rule's real reason, a permit, the match's boundary, and the recovery the refusal names. The judges are mocks through SR_CHECKS_JUDGE_MOCKS (adr-conformance, adr-grounded, adr-matches-sloprails, adr-well-formed, invariant-grounded, invariant-rigor), deciding from their rendered prompts; the cited cases get the user's words from one scripted agent turn. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): the new cases name their rule literally in the jq that asserts its FileGuardChecked event rule-tests-rigorous's script floor looks for one jq selecting the owning rule by name. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(checks): invariant-covered and capability-covered error on tooling failure, not a cached fail Their spec helpers returned an empty SPEC when `ls spec/<kind>/*.yaml` failed and harnesses() listed nothing on an incomplete tree, so the rules read "no specs" / "no claude-mock/" and refused with a plain fail, which the engine caches as a terminal verdict (seen on PR #178). Since sloprail #266 a refusal carrying "error": true is never cached. - _lib/changeset.sh: refuse_error (reason + "error": true); an unset SR_TREE and a failed git grep over markers are tooling errors. - _lib/spec.sh: load_spec refuses with an error when yq/jq are missing, spec/<kind> is not in the tree, the listing fails or holds no *.yaml. Invalid YAML and every content finding stay plain refusals. - capability-covered: no *-mock/ dir in the tree is an error. - rule tests: errors-on-an-incomplete-tree for each rule (healthy -> pass, missing or empty spec dir -> error, genuine finding -> plain refusal). Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(checks): recovery and the other plain findings in the covered-rule error cases Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(checks): load_spec at top level in the pairs scripts; test the missing spec dir and mock dir A refusal inside $(...) only ends the subshell, so rigor_pairs/reconcile_pairs no longer call load_spec: their callers load $SPEC at top level (inputs-ready.sh and prepare.sh now do). spec.sh avoids a pipefail-prone grep. New cases: capability-rigor errors-on-a-missing-spec-dir (rigor and reconciled scripts), and a missing *-mock/ dir in capability-covered's case. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(checks): capability-rigor missing-spec-dir case asserts the engine's event Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(checks): no swallowed failure in the capability-rigor case Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(checks): an empty harness list is an error in reconciled, its subjects and snapshots-current Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): module-coverage tells an absent base ADR from an unreadable base Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): tooling failures in the module, capability, invariant and ADR rules are errors, not cached verdicts; the cases assert it Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): a failing subjects.sh is the engine's report, not a check's error flag Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-grounded, capability-covered and cells.sh refuse on a failed changed-files, harness or shape lookup (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): cases for every remaining capability lookup that refuses, with a skip-counting jq shim (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): no swallowed failure in the prepare case; capability-rigor's inputs-ready refusal and recovery are a case (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): load_touched_markers fails on an unreadable marker table, diff or file, and capability-rigor refuses (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-covered lists mocks itself (one refusal); changed-files matches are SIGPIPE-proof (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): cases for the marker/diff/pairs refusals, a stray z-mock, a >64KB changed list; header on the not-ready case (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): the >64KB case asserts its status instead of swallowing it (#204) Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-grounded narrowing list is built in a checked step; drop the unreachable refuse The harness list was built inside a process substitution, where a failed jq is invisible (an empty file reads as null and the judge would get no providers). It is now built first, with a refusal, and a case injects the inner jq. touched_harnesses cannot fail (load_touched already ran and refuses), so its dead `|| refuse` is replaced by a comment. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-grounded keeps its refusal on the touched-harnesses read, with a case that fires it The previous commit dropped `|| refuse` on touched_harnesses; the grounded-rule-changes judge rightly objected (it loosens a rule). The refusal stays, and an awk shim (touched_harnesses ends in an awk over the table) makes it fire in refuses-when-a-prepare-lookup-fails; without the refusal that case fails. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): capability-rigor's samples listing no longer refuses a run with no samples (#204) The listing loop ended on a false [ -f ] &&, which under pipefail failed the substitution and refused a legitimately empty run. An if keeps only a real jq failure fatal. A case runs prepare.sh over a change citing a run with no samples. Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): the late #216 fixes land as errors; capability-covered's missing-mock message is the one #225 and the other rules use Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): layering also checks a package's test imports (adr-matches-sloprails finding), with a case Sloprail-Cites-Tool: leaves out TestImports and XTestImports Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): the layering case asserts the sr-checks status instead of swallowing it Sloprail-Cites-Tool: refuses-a-test-importing-another-mock is not a rigorous sr-test case Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): snapshots-current's drifted-doc case shows the recovery Sloprail-Cites-Tool: never shows the agent fixing it Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): snapshots-current refuses when no *-mock directory exists, with a case Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): the object-id, layering and capability-grounded tooling failures are errors; yaml_str_json removed; module-leaks' BAD line is noted as a verdict Sloprail-Cites-User: A rule must refuse when it cannot work out what to check; it never passes on an empty or failed lookup. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(rules): replay-exceptions-only-shrink: every replay exception list may only shrink, and no reason moves to a weaker category Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Sloprail-Cites-User: why fucking not? do Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(rules): a replay list moved to another replay directory of its mock is compared with its old self; cases for the move Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): replay-exceptions-only-shrink sees a list renamed or named from another file, knows the reason categories, and reads the entries once into variables; the ADR names the categories Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): replay-exceptions-only-shrink keeps to the user's words: a created list is empty-only (a git mv rename is compared), no legacy exemption; the ADR is stated without time words Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * docs(rules): replay-exceptions-only-shrink's header names the three bypasses its other-file and reason-category checks close, with the output that showed each passing Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Sloprail-Cites-Tool: MISFIRE-PREFIX untriaged y weakened to y with no category: the old rule exited 0 Sloprail-Cites-Tool: MISFIRE-INIT init in zz_test.go adds a notReplaying entry: the old rule exited 0 Sloprail-Cites-Tool: MISFIRE-MOVE list moved to exceptions_test.go with a new entry: the old rule exited 0 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): capability-grounded's lookup cases use a judge mock that decides from its prompt, and harness-lacks-cited.sh has a refusal and recovery case Sloprail-Cites-Tool: so they are the always-pass mocks the criterion calls proof of nothing Sloprail-Cites-Tool: is never exercised by any case, so its refusal and recovery are untested Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): the harness-lacks case asserts capability-grounded's own outcome, refused and then passed, beside the script's reason Sloprail-Cites-Tool: test.sh asserts no event of its owning rule capability-grounded Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * test(rules): the harness-lacks case reads the outcome with one jq over the events, not a constant filter Sloprail-Cites-Tool: refuses-a-harness-lacks-deviation-with-no-citation is not a rigorous sr-test case: - test.sh runs jq Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * docs(adr): replay-exceptions-only-shrink's concern has no time word Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * docs(adr): replay-exceptions-only-shrink cites the user's words for the four reason categories Sloprail-Cites-User: A replay exception's reason must start with one of adapter:, mock gap:, untriaged: or flaky:; a reason with none of these is refused. Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): replay-exceptions-only-shrink refuses a new line naming the list in the generated test and a key with a tab, and matches any <mock>-mock/e2e path Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Sloprail-Cites-User: why fucking not? do Sloprail-Cites-User: A replay exception's reason must start with one of adapter:, mock gap:, untriaged: or flaky:; a reason with none of these is refused. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
|
All contributors have signed the CLA ✍️ ✅ |
…ness) pair; adr-well-formed asks only scope (#237, #238) (#242) * feat(rules): capability-rigor judges per (capability, harness) pair; the shared sr:capability marker touches no harness's pair Sloprail-Cites-User: A capability is judged per harness: one harness's missing tests never refuse a change that touches only another harness. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(rules): capability-covered checks only what a change touches: a capability's shape, the pairs whose cell, recording, markers or deviation ADRs changed, and the markers touched Sloprail-Cites-User: A capability is judged per harness: one harness's missing tests never refuse a change that touches only another harness. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * feat(rules): adr-well-formed asks only that a Decision names where it applies, never for detail the user did not give Sloprail-Cites-User: An ADR records the user's decision in their own words; adr-well-formed may ask only that it names where the decision applies, never for detail the user did not give. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * fix(rules): per-harness scoping keeps its fail-closed tail, back-checks a marker whose cell was touched, and the lookup cases follow the pair-scoped subjects Sloprail-Cites-User: A capability is judged per harness: one harness's missing tests never refuse a change that touches only another harness. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * test(rules): the shared-statement scenario of judges-per-harness asserts the refusal's reason Sloprail-Cites-Tool: checks54-finding-A: the shared-statement scenario asserts a refusal with no reason Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * fix(rules): capability-covered checks the harness set for every capability when a mock dir is added or deleted Per-pair scoping skipped the cell-presence check for untouched capabilities, so a new <h>-mock/ with no cells, or a deleted one whose cells remained, passed. The base and head mock-dir sets are compared now; a changed set (or an unreadable base) puts every capability in scope for cell presence. New case judges-a-changed-harness-set covers both directions. Sloprail-Cites-Tool: a new mock dir with no cell: capability-covered did not refuse Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * test(capcells): the fixture's change adds the capability file, so it touches every pair it judges The four tests ran covered.sh over an empty changeset. Under per-pair scoping an empty change touches no pair, so the bare false, the false without docs or runs, the pending cell and the supported cell without a proving test were never judged: the fixture, not the rule, was out of scope. The change now adds spec/capabilities/cap.yaml, which touches every (capability, harness) pair; the assertions are unchanged. Sloprail-Cites-Tool: cells_test.go:120: a supported cell without a proving test must be refused, got 0 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * test(subjects): capability-rigor subjects are (capability, harness) pairs; the expectations say so The capability-rigor expectations still named bare capability ids ('a ', 'b '), but its subjects are <id>/<harness>. They now expect 'a/claude ', 'b/claude '. The key lookups used the bare id too, so "b's proving test changed, b's key moves" compared two empty keys (and the two "same key" checks compared empty with empty, vacuously); they look the pair up now and compare real keys. capability-grounded keeps its bare ids. Sloprail-Cites-Tool: FAIL: cap-rigor: a run of b touches b alone: 'b/claude ' != 'b ' Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…ty list (#250) * fix(replay-exceptions-only-shrink): the one-line empty map is the empty list gofmt writes an empty map as `var notReplaying = map[string]string{}`, so the two-line form the rule demanded is not gofmt-clean, and the rule refused the empty list it exists to reach (#239 emptied claude's list). A one-line empty map is now read as the declaration and its closing brace: emptying a list to {} passes, and an entry added back to a {} list is refused, naming it. Sloprail-Cites-Tool: gofmt-writes-empty-map-q9: var notReplaying = map[string]string{} Sloprail-Cites-Tool: with one "run": "reason" per line (file-guard "replay-exceptions-only-shrink") Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * replay-exceptions-only-shrink cases: a delete( of the list in the generated test is refused, as the comment says The case's comment claims a delete( and a maps.Copy( are refused, but it only wrote the maps.Copy(. It now also adds a delete(notReplaying, ...) line and asserts the refusal. Sloprail-Cites-Tool: rigor-finding-q9: case refuses-a-generated-test-that-gains-a-line-naming-the-list: its comment says Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
…e cited docs mention (#234) * capability-rigor: judge the statement's scope, not every behaviour the cited docs mention The judge demanded an sr:proves test for every behaviour anywhere in a cell's cited doc pages, so each run surfaced new "the doc says X, no test" items, even for doc behaviour outside what the capability claims. It now judges whether the provider's tests prove the cell's statement (the capability's scope) plus the declared deviations. The docs and runs ground the statement and are read for its meaning; they are not a checklist, and doc behaviour outside the statement is not demanded. A recorded run replaying green proves what it recorded. Fail-closed: a statement clause that no test and no replay proves still refuses. Three sr-test cases (mock judge: it proves the rubric reaches the prompt and the verdict path, how the real model reads it is for sr-eval): a doc-only behaviour outside the statement is permitted; a statement clause with no test is refused; a clause proven only by a green replay is permitted. Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * capability-rigor cases: assert the rule's event in each test.sh, the refusal with its reason Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * capability-rigor cases: the permit case's failure text names no refusal Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * capability-rigor: only a run that replays green counts as proof prepare.sh reads each harness's replay exception list (notReplaying in <h>-mock/e2e/*/replay_allowlist_test.go; a list that cannot be parsed, or a harness with neither a list nor a replay test, refuses) and marks every run replays: true|false with its notReplayingReason. The template says only a replays="true" run proves a clause; a run on the exception list proves nothing. The mock judge honours the flag, and a new case (refuses-a-clause-shown-only-by-a-run-that-does-not-replay) installs a notReplaying entry for the run. The cases share one scope-lib.sh and judge.sh in capability-rigor/test-lib/, and the file-guard.yaml comment is rewrapped. Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * capability-rigor cases: the two refusals recover once a test asserts the clause Sloprail-Cites-User: We don't implement all the docs, right? We only implement the scope of the capability, correct? Um, and then based on the snapshot, kind of make sure that it replays. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * capability-rigor: list every finding in the first refusal (harness-mocks#236) The template tells the judge to finish the checklist and name every unproven clause of every subject in the one refusal, one finding each, never only the first. The mock judge names all of them, and two cases cover it: lists-every-unproven-clause-at-once (beta, gamma and delta in one refusal, fixed in one pass) and reruns-after-an-unrelated-change-raise-nothing-new (the same single finding after an unrelated change). The two refusal cases from the replay change now recover once a test asserts the clause, as rule-tests-rigorous asked. Not done: handing the judge the previous stored verdict's findings and the diff since. A verdict is a pure function of its cache key (rule hash, subject files, fingerprint); prepare's output is not in the key, the store is an internal compressed format with no reader for a script (sr-checks show --json takes about a minute per range), and mocked judges are not cached, so it cannot be proven here. Sloprail-Cites-User: capability-rigor lists every finding for a subject in its first refusal, and a re-run raises new findings only for what changed since. Sloprail-Cites-Tool: (the rubric says the fix is a test), but no case adds a test for beta Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG * capability-rigor cases: the older sandboxes have a replay exception list, and the samples-listing injection lets the list's own jq -R through Sloprail-Cites-Tool: c/claude: no replay exception list or replay test of claude-mock: cannot tell which runs replay green Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * capability-rigor cases: per-pair fixtures give each mock an empty replay exception list After the rebase onto the per-pair rules, judges-per-harness and the doc-problem scenario of subjects.sh refused for the missing list instead of the behaviour they test. Sloprail-Cites-Tool: c/claude: no replay exception list or replay test of claude-mock: cannot tell which runs replay green Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * docs-follow-recordings: the sandbox has an empty replay exception list capability-rigor's prepare now needs each harness's replay exception list (or a replay test) to tell which runs replay green; the suite's sandbox had neither, so 'a drifted page is read' got the refusal instead of the page. The sandbox gets an empty notReplaying list. Sloprail-Cites-Tool: docs-follow-fail-q9: FAIL: capability-rigor prepare: a drifted page is read Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * capability-covered: list every pending cell of every capability, touched by the change or not adr/capability-once says the rule lists every pending cell on each run, but the per-pair scoping skipped untouched capabilities and pairs before the pending case, so only the pending cells of touched pairs were listed. The pending list is now taken for every capability before the scope filters; the checks themselves stay scoped. A Go test covers an untouched capability with a pending cell, still listed when the change only adds another capability's file. Sloprail-Cites-Tool: adr-pending-q9: The ADR's Decision says `file-guard/capability-covered` lists every `pending` cell on each run, but Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* noninteractive-run/claude declares that system/api_retry events are not modeled
Sloprail-Cites-User: The claude mock does not model system/api_retry events; it refuses them until a recording drives them.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: a turn that ends while a background agent works streams its result with the later turn's, after the notification turn's init frame (recording bgagent); the flag-refusal wording in the sr-agent flags test follows the new unknown-option message
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: test that a background shell killed at exit is reported after the result frame, as recorded (runs/bgbash)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: a sub-agent's task_progress names an Agent call by its description and a Read or Edit by the file, as recorded (runs/meta, file-tools); test of nested foreground agents' task frames against runs/meta
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: test that a failed background shell is reported failed, against the new recording runs/bgbash-failed; the cell cites it
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: test that an unknown resume ends its session before the result frame and that frame is all of stdout (runs/resume-unknown); the unknown-resume test keeps to the recorded form
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: test that the print wait for background agents defaults to ten minutes
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: a scenario that calls the Monitor or Workflow tool is refused by name
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude cells declare what the mock refuses or does not model: --bare, --agent and other deny rules (noninteractive-run); the workflow and Monitor tools and the wall-time ceiling (print-waits-for-background-agents)
Sloprail-Cites-User: All lgtm
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: the refused Monitor and Workflow tool names cite the tools reference that lists them
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude replay: the per-user temp folder claude-<uid> is the same on both sides (CI runs under another uid than the captures); the background-bash unit test no longer races a command that ends at once
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* tasks.Results holds the results of turns that end while a background agent works (core, next to AwaitAfterTurn); the deny-rule refusal lives in tool_hooks.go with the tool call; the refused-flag ADR names the code it governs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec(claude): hook-common-payload cites scratchpad-dir (runs in name order)
The cell lists the scratchpad-dir recording among its runs, and keeps the --agent deviation the user asked for.
Sloprail-Cites-User: The claude mock refuses --agent, text/json output, piped stdin and --include-partial-messages, and does not model the SIGTERM exit 143, until a recording drives them.
Sloprail-Cites-User: So those should not be limited or what's the problem? As long as capability describes it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: hook_additional_context records carry the text as the agent receives it (rendered system reminder, renderedRole system), as recorded; one sr:capability marker for print-waits-for-background-agents
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* ADR refused-flag-replay: scope says every mock's replay test follows it
Sloprail-Cites-User: A recorded run made with a flag the mock refuses replays as a check that the mock refuses that flag; it is not an exception-list entry.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* replay masks end_time instead of dropping it (the mock's task_updated patch carries it, as runs/isolated-worktree shows); the test checks the key
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: runner.go under the size limit (settings load and invoker setup in runner_setup.go); flagError notes only the long-flag wording is recorded; one import block in control_records.go
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: --include-hook-events streams PostToolUseFailure's and a foreground sub-agent's SubagentStart/SubagentStop frames as recorded (runs/include-hook-events-more), in raw print mode Stop's too; it is refused when the run reaches an unrecorded hook (compaction, worktree, background sub-agent)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude replay: prepare.sh runs with a claude that accepts only plugin commands (the mock reads a local marketplace itself; CI has no claude), so plugin-hooks and plugin-hooks-same-command replay there; the unknown-resume test's hook no longer sleeps near SessionEnd's budget
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor: replay every recording through the mock's own replay command (#158)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor replay: review fixes (named gaps in the reasons, temp dir cleanup, empty hook log not green, print-script exit 2, multi-sample test)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor session-start-hook: cite runs/session-resume; drop proves markers from internals unit tests
The no-start-hook-on-resume claim rests on runs/session-resume, now in the cell's runs. The decision_test.go markers for pretooluse-refusal and hook-timeout sat on unit tests of Interpret; both cells keep their e2e proves.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: background-agent discloses the subagentStart/Stop doc and recording conflict; print-waits quotes the subagentStop sentence (part of #161)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(spec): hooks-all-matching-run/cursor cites the combined-refusal run and says the refusals are merged
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* spec: drop the cell note (notes field removed in #176)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(spec): stop-block-cap/cursor cites the hooks doc's loop_limit default and discloses it conflicts with the print-mode recording (part of #161)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(spec): empty-tool-result-placeholder/cursor cites the exit-0 empty result
background-bash-start records ls exiting 0 with empty stdout: stream empty strings, hook {"output":"","exitCode":0}, no transcript tool result. Stays unsupported, with the sharper reason.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(spec): manual-compaction/cursor cites the recorded auto compactions and says what they show
The cell cited only the headless /compress run. runs/compaction-transcript-continuity
records two real compactions with a preCompact payload (trigger auto, no hook after,
nothing stopping them); the hooks doc calls preCompact observational. Stays
unsupported: the request, the hook after and the stop are all absent.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(spec): manual-compaction/cursor says no compaction hook fires on /compress
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(cursor): transcript user record opens with the empty timestamp element; transcript cells say what the recordings show (part of #161)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* spec: tidy the transcript-file deviation wording (part of #161)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* test(cursor): pin the mock's transcript_path at the first preToolUse and beforeShellExecution (part of #161)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* spec: disclose the doc's null-if-disabled transcript_path against the recordings (part of #161)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* spec(session-transcript-file): codex's recorded file-before-start behaviour is the mock's too, not a deviation (part of #161)
The entry had no kind and described behaviour the mock matches, so it declared nothing; the codex run still cites it.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: task-stream-frames and hook-common-payload say what the recordings show for sub-agents
Drop the stale 'mock has no sub-agents' deviation of hook-common-payload/cursor for a
harness-lacks one the subagent-lifecycle-hooks recording grounds; prove the sub-agent half
of task-stream-frames/cursor and the sub-agent payload identity with e2e tests.
Part of #161.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: hook-common-payload discloses the sub-agent's own transcript path; task-stream-frames proves the background command's frames
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: disclose the subagentStart/Stop doc conflict; assert the recording has no sub-agent frame of its own
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: attribute subagentStart/Stop doc fields per event; assert neither fires
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: sub-agent payload test asserts Shell cwd is empty and no other event carries cwd
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: stream tests assert the recorded Task result and notification order on the recording too
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: hold the removal of the stale 'mock has no sub-agents' deviation until the user's words arrive
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: the task stream test asserts the recorded interleaved order of a turn's calls, for recording and mock
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: disclose the Task call's missing postToolUse as a doc/recording conflict and pin it
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor: drop the mock-not-modeled claims the recordings and tests prove false (the mock runs sub-agents)
Removes hook-common-payload's 'mock has no sub-agents' and task-stream-frames' 'mock runs no sub-agent / Task answered as an error'; the true no-await-tool clause stays.
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor-mock: time a file tool's call; postToolUse reports a positive duration as recorded (part of #161)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* posttooluse-payload/cursor: drop the false 'duration of 0 for a file tool' clause
runs/file-tools records positive durations (45.586, 1.695, 0.759, 27.441 ms).
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor-mock: file-size split; postToolUse test no longer claims distinct tool_use_ids (the recording repeats one across a turn's calls)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor-mock: move Edit beside the file tools; toolexec.go back under the size limit
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(cursor-mock): fire beforeReadFile on each successful Read, as recorded (file-tools/cursor)
The mock carried the beforeReadFile docs marker but never fired the hook, and
the replay compare stripped it from the recording. It now fires after the
Read's preToolUse and before postToolUse with file_path, content and
attachments; the replay compares it; the cell drops its deviation.
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* test(cursor-mock): a read of an existing file fires beforeReadFile between preToolUse and postToolUse, as recorded
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* refactor(cursor-mock): move the tool Result types to result.go, under the file-size limit
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* spec(cursor): the mock fires beforeReadFile, so two deviations no longer say it does not
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* test(cursor-mock): a beforeReadFile hook is matched on the tool name
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* snapshots(cursor): runs/before-read-refusal, beforeReadFile hooks blocking reads
Six reads: exit 2, a JSON deny, invalid JSON, a crash (fails open), a failClosed
crash, and JSON with an unknown permission. Recorded with capture.sh at the
pinned cursor-agent 2026.09.28.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(cursor-mock): a beforeReadFile hook that refuses blocks the read, as recorded (file-tools/cursor)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* spec(cursor): beforeReadFile refuses a read as recorded, so pretooluse-refusal's deviation no longer lists it as unmodeled
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* test(cursor-mock): the beforeReadFile refusal test also proves pretooluse-refusal/cursor
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* spec(cursor): pretooluse-refusal cites the beforeReadFile refusal run and doc
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* test(cursor-mock): a blocked read gives the agent the failure hook's text, as recorded
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* fix(cursor-mock): the invalid-response wording is a beforeReadFile's alone, as recorded; the doc-only matcher test claims no cell
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* test(cursor-mock): a failed command's duration is the time it took
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UnEPsFQscGE2mKoG8Aq7Mx
* cursor-mock: every tool_call frame carries toolCallId and hookAdditionalContexts (the after-tool hooks' context, on the completed frame)
Recorded in every tool_call frame of runs/*; the contexts in runs/additional-context.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: Task frames carry the harness's args and the preToolUse payload is the call as the model made it; five recordings replay green
subagent-lifecycle-hooks, foreground-subagent-result, subagent-stop-block-loop, subagent-transcripts and subagent-worktree-isolation leave notReplaying.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* test(cursor-mock): the background Task test proves the receipt is at once, the recorded stream order and the payload fields (background-agent/cursor)
The capability-rigor judge found the test only asserted a completed taskToolCall
frame exists. It now also asserts the stream order of the Task call's frames,
the end notice and the result against runs/background-agent, that the receipt
precedes the agent's own command and the end notice follows it, and the
sub-agent's afterShellExecution output and sandbox and the sessionStart and
sessionEnd fields against the recording.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record a beforeReadFile hook timing out (failClosed or not) and a run with a second root added; replay tests
runs/before-read-timeout (cursor-agent 2026.09.28): a beforeReadFile hook that outruns its 1 s timeout is killed and the read goes through; the failClosed one blocks the read with a postToolUseFailure saying it failed closed and timed out after 1000ms. runs/multiroot-workspace: a run with --add-dir still tells every hook one workspace_roots entry. The mock accepts --add-dir (ignored, as recorded). Both runs are cited by their cells and replayed by e2e tests.
Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* adr: a mock replay replays every captured sample of a recorded run
Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record --add-dir against reads inside it and outside every root; the mock ignores --add-dir because the recordings show no difference
runs/add-dir-access and runs/no-add-dir-access (cursor-agent 2026.09.28): the same three Reads (a file of a sibling directory, a file outside every root, a file of the project), with and without --add-dir for the sibling. All reads succeed in both, with the same hooks, and hook payloads name one workspace root. The replay now maps <TMP> (the directory the workspace sits in) so such paths replay.
Sloprail-Cites-User: Record the cursor runs for a timing-out beforeReadFile hook (with and without failClosed) and a multiroot workspace_roots run.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* adr: replay-every-sample states the decision only
Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: --add-dir is accepted only with --force, the mode its recordings cover
The recordings of --add-dir (runs/add-dir-access, no-add-dir-access, multiroot-workspace) are all of -p --force. Without --force the mock refuses the flag, naming it, until a recording covers that.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: split decision.go and toolhost.go back under the 150-line limit (hooks/combine.go, runner/toolhost_after.go)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* module: toolhost_after.go is in internal/toolcall's home beside toolhost.go
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record --print alongside -p and test it; the depth-limit test asserts the recorded Task hooks and the call left in the transcript
runs/print-long-form (cursor-agent 2026.09.28): -p --print is accepted and the run is as with -p. TestDepthLimitStopsNesting now also asserts the two recorded Task preToolUse hooks and that every recorded agent said it had no Task tool, and that the mock makes the Task call at the limit (it is in the transcript) with no hook seeing it.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record matchers on Task and on beforeReadFile; the replay plays a recorded Task call
runs/hook-matchers-task-read (cursor-agent 2026.09.28): a preToolUse hook matched on Task runs for the Task call and one matched on Shell does not; a beforeReadFile hook matched on Read runs for a read and one matched on Shell does not; the postToolUse hook matched on Task never runs (no postToolUse fires for a Task call). The replay helper plays a recorded Task call as a Task whose sub-agent only replies.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-common-payload/cursor says a sub-agent is told from the main agent only by its own ids
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* adr: take replay-every-sample out of the batch (kept on side/adr-replay-every-sample)
The adr-well-formed and adr-grounded judges refuse the user-worded decision sentence: it has a gap clause (not checkable, not present tense) and names no replay command or check. Rewording it needs the user's words, so the ADR waits on its own branch.
Sloprail-Cites-Tool: The only Decision bullet fails 'Checkable' and 'Present tense, current state only': the clause "a mock that replays only the newest sample is a gap to close, not a convention" is a plan or description of existing code, not a rule
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: pretooluse-refusal/cursor discloses the agent_message doc and recording conflict
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: session-start-hook/cursor says its payload carries no start-kind field
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* test(cursor-mock): the symlinked-cwd run proves the sessionStart hook fires once with no transcript path
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-matcher-filter/cursor drops the matcher run, to be cited with its capture
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-matcher-filter/cursor cites runs/hook-matchers-task-read, the capture of matchers on Task and beforeReadFile
Sloprail-Cites-Tool: captured runs/hook-matchers-task-read/samples/20261005-213709
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record Grep and Delete calls with matchers on each; the mock models both, as recorded
runs/hook-matchers-grep-delete (cursor-agent 2026.09.28): a preToolUse or postToolUse hook matched on Grep or Delete runs for that tool and not for the other, with the recorded hook inputs and outputs, grepToolCall and deleteToolCall frames with their results. The mock runs both tools; a Grep with no match and a Delete of a file it cannot read fail rather than guess (unrecorded).
Sloprail-Cites-Tool: captured runs/hook-matchers-grep-delete/samples/20261005-215835
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: pretooluse-refusal/cursor lists the agent_message conflict after the existing deviations
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record an MCP tool call with matchers on MCP:<tool>; the mock calls the project's stdio MCP server, as recorded
runs/hook-matchers-mcp (cursor-agent 2026.09.28, --approve-mcps, a local stdio MCP server written in sh): hooks name the tool MCP:echo; a matcher on it runs the hook and one on MCP:other does not; the agent first reads the tool schema in a getMcpToolsToolCall no hook sees, then makes the mcpToolCall. The mock starts the server of .cursor/mcp.json for the call and refuses MCP calls without --approve-mcps (unrecorded).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record a Task call with an invalid model; the mock answers it as recorded, with no sub-agent
runs/foreground-subagent-failure (cursor-agent 2026.09.28): the Task call is not started, its preToolUse hooks fire, and it completes with "Invalid model selection ... could not be resolved to a valid subagent model" and the account's model list, with no after-tool hook. The mock answers the same for any model but default and lists only default.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: a background sub-agent launched by a sub-agent outlives it and its end is announced on the run's stream, as recorded; the nested background test asserts the recorded hooks, order and notification
runs/nested-subagents-background: the launching sub-agent reports STARTED while the background one goes on, its command's hooks and the sessionEnd come after, and the main stream carries a task_notification naming it. The mock owned the background sub-agent by its launcher, so it was never announced.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* test(cursor-mock): deny beats ask on beforeShellExecution and refusal messages are joined on beforeShellExecution and beforeReadFile
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: a blocked read's completed frame carries no args, as recorded; models inherit and default are the valid ones for a Task; before-read-refusal replays green
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor replay: setup prepare.sh and args are installed and passed (a flag the mock does not model is refused), MCP, Grep and Delete calls are mapped, a hook script's own log tag joins its event's group
prepare.sh runs in the repository under the run's home as the capture runs it; setup/args go on the mock's command line, only the flags it models; a recorded CallDynamicTool is a mcp__<server>__<tool> call with the model's arguments (its lookup left to the mock); Grep and Delete are unified tools; the model catalogue after 'Allowed model slugs:' is not behaviour. A closed.sh-style log line ran concurrently with its event's payload, so it is grouped with that event: before-read-refusal no longer depends on timing.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor replay: the new runs enter the exception list with their gaps, six greens leave it
session-resume-unknown replays green with args; no-add-dir-access, multiroot-workspace, print-long-form, hook-matchers-mcp, hook-matchers-grep-delete and foreground-subagent-failure are listed with the gap each shows.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: a Read's result names the file resolved, Grep, Delete and MCP frames carry their toolCallId in their args, as recorded; the print-flag and multiroot runs read a file, so they replay
The new runs print-long-form and multiroot-workspace are recorded with a Read instead of a shell command; no-add-dir-access and hook-matchers-grep-delete replay green and leave the exception list.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: assistant text is shown as one frame when the next call starts or at the end of the turn, a refused Task completes with only its error; MCP calls carry the model's description and their own hook id; hook-matchers-mcp and foreground-subagent-failure replay green
Recorded (runs/foreground-subagent-failure): a call refused before it started shows no frame at its call, so the text before it goes out with the rest at the end. The replay adapter reads the model-written mcpToolCall description from the stream. Both runs leave the exception list.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-matcher-filter and pretooluse-refusal back to main, to be corrected again in a commit that carries the citation
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: the beforeReadFile deviations of hook-matcher-filter and pretooluse-refusal match the recordings (the mock fires beforeReadFile); the agent_message doc conflict and the matcher runs are cited
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: the MCP server is run through internal/procexec; files split under the size limit and kept in their modules
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: the user's ~/.cursor/hooks.json is a hook source beside the project's, a hook in both runs once per source, as recorded
runs/hooks-all-matching-run-same-hook-two-sources. The e2e helper installs a run's user-hooks.json in its home and reads hook logs whose concurrent writes ran together.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hooks-all-matching-run/cursor says the mock reads the user source beside the project source, as recorded
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: assistant and tool_call frames carry model_call_id and timestamp_ms (the text at the end of the turn neither), as recorded; tests for them and for the common payload fields
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: replay_test.go back under the file-size limit; out.go belongs to the scenario module with frames.go
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: session-transcript-file keeps its codex section as on main; the kind-less deviation removal is handed to batch/codex as a patch
Sloprail-Cites-Tool: index 08642bc7..49cd9654 100644
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* test(cursor-mock): the transcript file exists when the first beforeShellExecution names it, as recorded
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-exit-code-semantics and hook-common-payload say where cursor differs from their statements (exit 0 with invalid JSON blocks; transcript_path null with transcripts on)
Sloprail-Cites-Tool: Cursor cell lists no deviation for exit 0 with invalid JSON on a permission hook
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: frame printing joins frames.go (no module change); the tests prove the sub-agent hooks that never fire, the notification detail as all the sub-agent said, and the transcript path on later hooks; hook-exit-code-semantics back to main for a cited re-correction
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-exit-code-semantics/cursor lists beforeReadFile among the events the mock fires and says an exit 0 with invalid JSON blocks, as the recordings show
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor replay: a hook log line of several objects run together is read as the objects (hooks of one event append to one log side by side)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record workspaceOpen firing in print mode; the mock fires it before the session start with the recorded payload
runs/workspace-open (cursor-agent 2026.09.28): the hook fires once, first, with only the event name, the Cursor version, the workspace roots and the user email (null), no session, model or transcript path.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: the cursor cells say what the recordings show of workspaceOpen, multiroot workspaces and model_id/model_params
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-common-payload/cursor drops the model_id/model_params deviation: it is a mock gap (the thinking hook is not fired), not a harness difference
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* afterAgentThought: a script's thinking block is a thought of the turn (core), the shared loop tells a Thinker host before the turn's message and calls, cursor fires the hook with the model's fields and its own generation
Optional: a host that does not implement turnloop.Thinker (claude, codex) never hears of a thinking block. A sub-agent and a turn that a finished background task starts are model requests of their own; the result frame keeps the run's first.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor replay: the thoughts a recording holds are fed to the mock, response by response, and no longer dropped
The unified Call and Agent carry an optional Thinking; the adapter puts the i-th afterAgentThought of a conversation on its i-th response when the recording holds one per response, and refuses a recording that holds them for some only (no guessing, nothing dropped). The generation's number and name are masked, a thought and the preToolUse beside it are concurrent. nested-subagents replays green.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: a step's text and its call are one transcript record, as recorded; tests for the transcript's records, the foreground/background contrast; the cells say workspaceOpen names no session and a background Task is modeled
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor replay: a thought is put on the response its generation names, so a run whose model thought in some responses only is replayed (symlinked-cwd)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: the cursor cells list afterAgentThought and workspaceOpen among the events the mock fires, and cite the run that replays with thoughts
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: record matchers on afterAgentThought, workspaceOpen and sessionEnd with a thinking model; the mock tests a thought hook's matcher against AgentThought, names the run's model in the init frame and sessionEnd, and refuses models it has no record of
runs/hook-matchers-thought (cursor-agent --model cursor-grok-4.5-high): the afterAgentThought hook matched on AgentThought and the one with no matcher run for each thought, the one matched on nothing does not; workspaceOpen and sessionEnd run every hook whatever its matcher. The sessionEnd payload and the init frame name the model the run was started with.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-matcher-filter/cursor cites the thought matcher run and discloses that no postToolUse follows a Task call
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: a Grep with any parameter but the pattern fails; unrecorded sub-agent models are refused as not modeled; --add-dir is gated on --force alone; the replay masks the clock and the service ids instead of dropping them, and tool_call frames say when a call began and ended
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-matcher-filter/cursor drops runs/nested-subagents, which has no matchers
Sloprail-Cites-User: A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me; adding a new one still needs my words.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor: StrReplace, GetDynamicTools and AwaitShell are modeled from the recordings that use them; file-tools and nested-subagents-depth replay and leave the exception list, the other four are listed with what is left
A StrReplace is an edit whose hooks see the whole file it makes and whose result carries the diff; a catalogue search finds nothing (the mock has no catalogue); a wait with no task named lasts as long as it was told. A hidden read before a write carries the write call's id.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: a catalogue search is answered only where a recording shows it finding nothing, any other is refused as not modeled
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: a refusal of something not modeled fails the run, a sub-agent's too; nested-subagents-depth is listed again, as its recording does not hold what the sub-agent's search answered
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: files back under the size limit and in their modules; a test that the result agentId is the sub-agent's own id and not the args'
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor-mock: the background Task test checks it returned at once; think.go joins the hooks module and refusal.go the scenario module
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* adr: each mock's replay test replays every captured sample of a recorded run
Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* cursor replay: the exception list is accepted as it stands and may only shrink
Sloprail-Cites-User: Accept the cursor mock's initial replay exception list as it stands; from then on it may only shrink.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: foreground-subagent-result/cursor says the mock does not model a sub-agent failing mid-run
Sloprail-Cites-User: The cursor mock does not model a sub-agent failing mid-run; no headless run produces one.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-common-payload says a sub-agent event carries an identity of that sub-agent: its own session id, or the main session's together with an agent id
Sloprail-Cites-User: All lgtm
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* test(cursor): pin the recorded afterFileEdit edits and a JSON allow on beforeReadFile; module for the refusal
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* adr: take replay-every-sample out of batch/cursor (it lands from side/adr-replay-every-sample with batch/rules-pin)
Sloprail-Cites-User: Each mock's replay test replays every captured sample of a recorded run; a run counts as replaying green only when every sample does.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* spec: hook-common-payload/cursor discloses the attachments entries the mock does not model
Sloprail-Cites-User: The mock's payloads always carry an empty attachments list; the entries a real cursor-agent adds for an attached file or rule (a type and a file_path) are not modeled, since no headless run can attach one.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* test(cursor): a Task call at the depth limit is answered as an unknown tool
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* fix(cursor-mock): refuse from the structured message, not a scan of printed frames; test the unrecorded-model, Grep-keys and --yolo --add-dir cases
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* fix(cursor-mock): --add-dir is refused with --yolo too (only --force is recorded); test it
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F8q8xpGsstSzjknnXFEfFG
* claude-mock: --include-hook-events also refuses the subagent_start control record, a scenario-written tool_result's PostToolUse and a background task's notification UserPromptSubmit (frames not recorded); only a manual compaction refuses on SubagentStop; the early-return sub-agent writes its held start frame; worktree and other cases in TestT001_19
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: declare the tools the mock plays (Edit, Grep, Delete, GetDynamicTools, AwaitShell, MCP tools) in the validated schema; toolspec: prefix and open tools
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: stream_scan.go within the size limit; the noninteractive-run cell cites runs/include-hook-events-more
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a Shell call's frames carry what Cursor reports of it (parsed command by a shell parser, limits and flags, ids, description); five runs replay green
hook-matchers, hooks-together, session-end-hook-output, stop-block-cap-always and plugin-hooks leave notReplaying; the rest of the hooks half have their remaining gaps named.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a shell command's own hook names no transcript; replay places a thought by request, not by the response number alone
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a Task call needs only its prompt; one without it is never started, its completed frame has no args (agent-input-validation, -description replay green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude replay: several texts before a call, and the answers that end a turn the harness goes on from (a notification, a Stop block), are steps of the script; a Stop block streams its feedback frame, the error notice once, and the cap's informational and notification frames, the override counting a turn; a background agent is announced with its launch and has no prompt frame. cap, stops and hook-exit-codes leave the exception list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude replay: a tool the adapter has no mapping for goes on under its own name (the mock runs or refuses it); a wakeup's wall-clock time, seconds-to-go and scheduledFor are masked. schedule-wakeup-limits leaves the exception list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor replay: hooks that name a call by an id of their own are matched to the call by its input, a file call that failed has no args, no hook log is no hook ran, the flaky hand-written comparison of additional-context goes
A call's hooks name it by an id other than its call id in 17 of the recordings: the adapter puts the id on the call whose command, prompt, pattern or path the hook saw, and refuses an id no call matches; the mock names the call's hooks by it (hook_tool_use_id, mock-only). A completed Read or Edit that errored carries no args. A recording with no payloads.jsonl whose hooks were configured, and whose stream compares green, proves the hooks never ran. additional-context, hook-exit-codes, hook-fail-closed, hook-matchers-after-events, hook-timeout, hook-timeout-early-output, hooks-all-matching-run-same-hook-two-sources, pretool-refusal, pretool-refusal-combined, pretool-refusal-file-tools, session-start-continue-false, stop-block-continuation, tool-failure and user-prompt-submit-hook leave notReplaying, with symlinked-cwd, noninteractive-force-write and foreground-subagent-failure, which the shell frames and the hook ids turn green too.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: a run that took no model turn ends with a result frame that has no terminal_reason or api_error_status (prompt-blocked, prompt-blocked-json, prompt-blocked-suppressed replay green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: a child's environment names the harness's messaging socket and secret, Bash its executable, a SessionStart hook its env file; --name names a session it starts; a named session's prompts carry its title; a replay runs an earlier run of the recording (prepare.sh) and the resume, continue, fork and no-persistence flags (nested-session-env, subprocess-session-env, resume-*, no-session-persistence replay green)
Sloprail-Cites-User: '''A 'mock-not-modeled' deviation that a recording proves false may be removed or corrected to match the recordings without asking me
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude replay: a /compact step replays (the compaction, its summary text, the preserved messages, an unwritten logical parent and the summarizer's own output are read from the recording); /compact is the harness's own command: no UserPromptSubmit, no Stop, its result says local_command; the summary streams as a synthetic user frame; SessionStart:compact names the model; a first-run --session-id is the replay's own. compact and compact-nohooks leave the exception list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: a background shell a sub-agent launched is announced owned_by_subagent, as recorded (runs/fg-subagent-bash); the entry leaves the exception list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a foreground command's hooks name it by an id of its own, its tool_input carries the model's timeout (none for a background one); the shell sees CURSOR_REQUEST_ID and CURSOR_RIPGREP_PATH, hooks CURSOR_RIPGREP_PATH; replay does not compare what Cursor leaves unsettled (the first beforeShellExecution's transcript path); a refused command reports the time it took (shell-exit-status, symlinked-cwd, noninteractive-force-write, noninteractive-no-force, nested-session-env, subprocess-session-env and, as a result, 12 hooks-half entries replay green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor: shell syntax the recordings show is modeled (double-quoted words, < >> 2>&1, command substitution), the rest is refused by name; stop-hook-payload re-recorded and closed
runs/shell-syntax records a double-quoted word (string), < (fd 0), >> (fd 1), 2>&1 (operator >&, number target, all-quiet) and $(...) (command_substitution, its commands listed after the outer). Variables, && and ||, here-documents, globs, subshells, loops, assignments, backticks and background jobs have no recording: the validator refuses them (Param.Unmodeled). runs/stop-hook-payload is re-recorded with a prompt that avoids the catalogue search.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: a Stop hook that blocks streams its feedback to the agent as a synthetic user message, and its error notice once per run; a sub-agent re-run after a blocked SubagentStop streams no second prompt; a replay makes the answers a Stop hook sent the model on from (stops, hook-exit-codes, cap-sub replay green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a turn's closing text waits for the run's end (a finished background shell's turn follows); a background shell's pid and id are masked, its result in the order Cursor words it; its executionTime is measured; commands naming the harness's ripgrep are refused as not modeled (background-bash-start, bg-bash-reaped-at-exit, task-notifications-bg, task-notifications-inturn replay green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude replay: a run, earlier or later, starts in a directory of the repository (cwd, symlink) with the project files of that directory (symlinked-cwd, forkresume replay green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: allowlist after merging sprint/cursor-a (the union of both halves' removals)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: --tools names the tools a run has; an Edit's omitted replace_all is filled in as recorded, the stream naming the input as sent; a replay maps Write, Edit and Glob and plays the inputs as the model sent them (file-tools replays green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor replay: the harness's own skill files a run read are laid out from the recorded read, a read's content_length counts UTF-16 units; schedule-wakeup-ask closes
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a sub-agent's conversation takes the id a replayed reply quotes (mock-only agent_id), a background Task reports its measured duration; background-agent, print-waits-for-background-agents replay green; shell ids masked as numbers (task_id, taskId)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* fix the merge: core.Agent carries the agent's own id (the cursor adapter reads it), a test closed
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude replay: a sample recorded when a setup file read otherwise than it does now (the model's Read of hook.sh shows it) is replayed with the file as the model read it (hookmix replays green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a response's Task calls start after its other calls; consecutive text-only responses of one turn are one answer (nested-subagents-background replays green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude replay: adapter.go and denormalize.go within the size limit
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a background shell's file has the header and footer Cursor writes, a wait on a named shell is answered, a response's Task goes after its other calls; core: steps and exit statuses in the unified recording
task-stream-frames stays listed with what is left: the calls' preToolUse order and the Read hooks a wait fires.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: a command a foreground sub-agent's end ended still gives the parent its follow-up turn; a process listing is compared as a listing, not by which processes (foreground-subagent-bash-ends-with-response replays green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* claude-mock: a foreground Bash command of the main agent that runs 3 seconds or more is a task (task_started when it crosses that, task_notification when it ends), as recorded in the new runs bash-long-foreground and bash-short-foreground (hook-timeout replays green)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor replay: a run of several steps is replayed step by step over one home, exit statuses are compared (core), a later step may resume the session or be refused; session-fork closes
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* cursor-mock: record the catalogue search a sub-agent at the depth limit made (runs/catalogue-search-task: matches []), so nested-subagents-depth replays green
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* feat(replay): the codex adapter gives the mock a run's -c agents.max_depth, -C and --ephemeral, and installs a project layer's hooks
hooks-all-matching-run-same-hook-two-files, nested-subagents-nowait, compaction-transcript-continuity and hook-command-subdir
replay green and leave notReplaying. The mock makes a relative -C absolute, as the harness does.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* feat(replay): a sub-agent's own sub-agents are replayed (nested-subagents, nested-subagents-limit)
The loader attaches sub-agents recursively, taking a sub-agent's receipts from its own rollout (its spawns are not in the
run's stream), and the scenario gives each its own script, named across the whole tree.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* feat(replay): a refused run, a run's own command line, a resume of an unknown session and a stopped compaction replay
The adapter reads the recorded command line (--json, --skip-git-repo-check, --dangerously-bypass-*) and gives the mock the same
flags, skips the git setup for a run recorded outside a repository, and expects the recorded exit status. A run the harness
refused (no rollout, non-zero exit) replays as that refusal; a turn aborted with no interrupt is a stopped compaction.
noninteractive-run-git-check-refused, noninteractive-run-no-git-check, session-resume-unknown, manual-compaction-auto-blocked and
manual-compaction-auto-post-stopped leave notReplaying.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* feat(replay): a run's env file and its prepare.sh (with codex plugin and sed stubs) are replayed: nested-session-env, plugin-hooks
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* feat(replay): the model's apply_patch calls are replayed (file-tools, file-tools-failure)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* feat(replay): a hook's own non-JSON lines are kept in the compared log (bg-bash-reaped-at-exit)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah
* feat(codex-mock): write_stdin polls a command left running; a resumed script cell is replayed (task-stream-frames, subagent-stop-block-loop-cap)
The mock keeps a yielded command's output and item, and …
…41d413 (#264) Verdicts are keyed by their input only (sloprail #290, #298): the store moves to schema v2026-10-07, which the old pin refuses. The old pin's store gc also crashes on dictionary training (fixed on main). Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…e recording's gates, not timing (#263) * fix(claude-mock): the order of a parent's and its sub-agents' frames comes from the recording's gates, not timing Replays of bgagent-concurrent-limit and bgagent-nested-launcher swapped a parent's text or hook with a sub-agent's frame under load. The gates that order agents' steps had gaps, each closed here from the recording's own times: - a step waits for the calls of sub-agents below its sub-agents (the inner agent's tool_use was ahead of the main agent's answer), via ChildCalls.Via; - a step waits for a sub-agent's call the sample shows carried out ahead of it (the shell's task_started ahead of the answer), via ChildCalls.Executed; - a step of an agent waits for the calls of the agents above its parent to have finished (main's PostToolUse ahead of the inner agent's PreToolUse), via Gate.AncestorDone; - a sub-agent launched by a sub-agent's call is started after that call's PostToolUse when the payloads show it, as the main agent's already was. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * refactor: split the gate and spawn-log files under the 150-line limit file-size refused gates.go, gate.go and spawn_log.go (176-178 lines). Moved by responsibility, no code changed: the gate walk helpers (gates_walk.go), the hold functions (gate_hold.go), the carried-out-call progress (progress_exec.go) and a spawned agent's progress (spawn_progress.go). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
… their finish order (#262) An event's hooks run together, so the log holds them in finish order, which the recording does not promise. Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Only the main session runs sr-checks run; sloprail/sloprail#294 ships the gate (off by default). Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah Sloprail-Cites-User: you can have it in this plugin disabled by default and then enabled explicitly in our tool repos Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…es; adr/replay-every-sample (#253) * mocks: a flaky: entry is replayed three times by every mock's generated replay test, and fails when none is green claude, codex and cursor generated_replay_test.go run a flaky: entry through replayUntilGreen (flakyRuns = 3) and fail it when no run is green, instead of passing it on one green run or skipping it. Each package has a flaky_replay_test.go for replayUntilGreen. Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * replay-exceptions-only-shrink: the replay test package is pinned whole, the replay files are read from the syntax tree, and the generated tests are pinned to canonical copies tools/replaycheck reads each mock's notReplaying list and generated replay test from the syntax tree and type information instead of their text, and compares replayUntilGreen and TestGeneratedReplay with the canonical copies in the rule's folder (one per mock: claude, codex, cursor), so a flaky: entry is still run three times and failed when never green. A mock's replay_allowlist_test.go and generated_replay_test.go are no longer deleted, renamed or moved: that would end the replay without removing an entry. A reason that starts with none of adapter:, mock gap:, untriaged: or flaky: is still refused. Cases for the package pin, the text a grep would miss and the map outside its file; the move cases go, as a move is now refused. Sloprail-Cites-User: The replay exception list may only shrink; adding an entry fails CI Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Sloprail-Cites-User: A replay exception's reason must start with one of adapter:, mock gap:, untriaged: or flaky:; a reason with none of these is refused. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * adr: a mock's replay replays every captured sample of a recorded run (linked to the replay-exceptions-only-shrink rule that pins the replay tests) The decision is the user's own words; the pin on generated_replay_test.go is what keeps a mock's replay test from being weakened to fewer samples. Sloprail-Cites-User: A mock's replay command replays every captured sample of a recorded run, and "replays green" means all of them do; a mock that replays only the newest sample is a gap to close, not a convention. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * replay-exceptions-only-shrink: the timed replays may be skipped on macOS and run three times on Linux; any other -skip of a replay is refused The workflows and a mock's Makefile are judged too: a go test -skip there may name only TestGeneratedReplay/all-hooks-close-(first|second) (directly, or through TIMED_REPLAYS defined as exactly that), and a workflow that skips them must also run them on their own with -count=3. Any other skip, a widened TIMED_REPLAYS, or a skip with no run is refused. Cases for both directions. Sloprail-Cites-User: Fail, never skip Sloprail-Cites-User: A replay exception's reason may not move to a weaker category (flaky < untriaged < triaged), and a flaky entry runs 3 times and fails if never green; it is never skipped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * replay-exceptions-only-shrink: the timed-replay skip is the user's decision and requires Linux, three runs and every PR; adr/replay-every-sample drops its link to the pin rule adr/replay-exceptions-only-shrink states the timed-pair exception in the user's words (macOS e2e jobs skip it; the Linux timed-replays job runs it three times on every PR). The rule now also requires the workflow to be triggered by pull_request and the job that runs the timed pair to be on ubuntu, with cases for both. adr/replay-every-sample goes back to file-guard/adr-conformance: the pin rule does not check which samples a replay covers. Sloprail-Cites-User: Timing-sensitive replays whose recorded gaps a CI runner's timers cannot keep may run on Linux only, three times each, as long as they run on every PR. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah * adr/replay-exceptions-only-shrink: the skip bullet states the policy in the user's words adr-grounded refused the bullet for deciding more than the cited words (it named the tests, the platform and the job). It now says only what the user decided: every other replay skip is refused, and timing-sensitive replays may run on Linux only, three times each, on every PR. The names stay in the rule's script. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah Sloprail-Cites-User: Timing-sensitive replays whose recorded gaps a CI runner's timers cannot keep may run on Linux only, three times each, as long as they run on every PR. Sloprail-Cites-User: Yes, ban other skips --------- Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Comment-only PR: the diff from the last harness-mocks commit of the 2026-10-03 capabilities-matrix demo (9a26dad, #138; sloprail-community demos/20261003-capabilities-matrix) to current main. Base branch review/demo-baseline is pinned at 9a26dad. Do not merge.
Compare: 9a26dad...main
🤖 Generated with Claude Code
https://claude.ai/code/session_01RakB4JKMC2VyTt7wPgFSah