Files
stack/agents/dewey/work/queue-51/evidence.md
T
jason.woltjeandClaude Opus 5.5 717e3404f4 feat(conversation): release a stopped cohort scope on start and recover, no kill after a proven stop (#1537)
Row 51, owner Dewey, candidate c34039dc (5 files). A claim recorded
stopped on a cohort proof whose scope is still listed is released on
start (classify's stopped branch and the free path's session and seat
heads) and on a confirmed recover, after every recover check. A release
that ended unavailable or still listed is retried; the proof check stays
in #releaseScope. close({ killEngine: true }) no longer signals the
recorded PID after a proven stop. Tests R7-R12 kill relok, closeany,
noprooffkind and nopush; R9 pins the recover order.

Reviews: Filbert approve (27070, 27074), Darkwing changes then approve
(27071, 27075). Gate green on 032b5408 plus the candidate.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-10 02:06:20 -05:00

233 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Row 51 (#1537): cohort release follow-ups: crash window, retry and close
Dewey, 2026-10-10. Brief: `docs/plans/2026-10-10_cohort-release-follow-ups.md`,
section "Cohort release follow-ups: engine exit, crash window and close"
(rev 296). Base bbb2167f. The candidate touches four files in
`packages/conversation/` (`src/controller.mjs`, `tests/cohort.test.mjs`,
`tests/ctrl-child.mjs`, the README) and this file. `cohort.mjs` and
`shim.mjs` are unchanged. Engine exit is row 52 (decision 79); the EOF path
is as row 50 left it.
## Round 2
Darkwing asked for changes (comment 27071): the recover release ran before
the pins, target and confirmation checks, so a refused `recover` could
still retry the release, record it and push it. Sage ruled: move the call,
keep the text. Filbert approved round 1 (comment 27070). Round 2:
- `#recover` retries the release after its last check, the confirmation,
and before the acquire. A refused `recover` changes nothing.
- R9 sends two refused recovers before the confirmed one: a confirmation
never issued, and one confirmed for `force-stop`. Each is refused
`confirmation`; `evidence.releases` is still the one `unavailable` entry
and the unit is still active. This kills `recconf` (round 1's order) and
`recbefore` (the release above the stop-proof check). A recover with no
confirmation field is refused `malformed` by the schema before
`#recover`, so it isn't one of the cases (K6 covers it).
- The cheap review notes, all inside the same four package files:
- Darkwing note 1, `closeproof`: R10 now calls `close({ killEngine: true })`
after the verifier is withdrawn and checks no signal reaches the engine
PID.
- Darkwing note 3, `nolookup`: R11 now restarts on the released claim; the
start launches and `evidence.releases` is empty.
- Filbert N1, `startnoawait`: R7's launcher records
`evidence.releases.length` when the launch begins; it is 1.
- Not taken: Darkwing note 2 and Filbert N2, `noseat`. It needs a second
session on the same seat, a new fixture shape. Listed under follow-ups.
- README: the scope-release bullet says which `recover` retries, that a
refused one releases nothing, and the R10/R11 additions.
## Claim-protocol rules touched
The change writes no claim record. Every release still goes through
`#releaseScope`, which acts only on a record with state `stopped`, a
`cohortProof` and a shim. `classify` and the claim store are unchanged.
- CHAT-00, alternate launcher and crash restart (line 179): "prior-cohort
proof before start; uncertainty retains claim". Unchanged. Start releases
only a claim already proven `stopped` on a cohort proof; an `uncertain`
or `stopping` pair still goes to `#serveOrphan`.
- W3 (acquire refuses `unsafe-replacement` while stopped without proof):
unchanged. A claim without a cohort proof is never released.
- W5 and W15 (crash between the keys; restart finishes under the same claim
ID): unchanged. The release runs after `classify` has finished the pair,
on the record it returned (R8).
- W7 (boot ID differs: stopped on a boot proof): unchanged, and that claim
is never released (R12). A boot proof says the cohort died with the boot,
not that this scope is ours to end.
- W8 (resume after a proven stop): the release runs before the acquire, on
the prior session head and, when it names another claim, the seat head.
A failed release doesn't stop the resume: the proof already shows every
member ended, so a listed scope is an idle shim, not a running cohort.
- W14 (stopped on a no-unit observation): no release; not a cohort proof.
- K6 (recover needs its own exact confirmation): unchanged. The recover
release runs inside a confirmed `recover`, after its `stop-proof`, pins,
target and confirmation checks; a refused `recover` releases nothing
(R9, round 2). No new route to `stopped`.
- K17 and K18 (eligibility): unchanged. On `start` the release runs before
the pins; on `recover`, after them and before the acquire. It changes
neither the session leaf nor the branch.
No CHAT-01, CHAT-01C or slice 1 rule changes.
## The change
`src/controller.mjs`
- `#releaseListed(why, claim)`. For a claim whose scope may still be
listed. If this controller already has a result for the claim and it is
`released` with `unit: "absent"`, it returns that. Any other earlier
result is dropped from `scopeReleases` and tried again. A unit that
`units.lookup` reads `absent` gets no request. Otherwise it calls
`#releaseScope`, which keeps the proof check.
- `start()`:
- The `stopped` branch (classify finished a pair whose keys disagreed,
or found a boot proof) calls it on the classified claim.
- The free path calls it on the session head and on the seat head when
that names another claim, before the acquire. Both keys read `stopped`
with a proof after a crash between `#claimFinish` and the release, so
`held()` is false and classify returns free.
- `#recover` calls it on the controller's claim after every check, the
confirmation last, and before the acquire (round 2). This retries a
release from this controller's force stop that ended `unavailable` or
`still listed`.
- `close()`: see below.
`tests/ctrl-child.mjs`: `dieAtState`, which dies at a barrier only when
its patch sets a given state. R8 uses it to die at `between-keys` on the
`stopped` finish and not the earlier transitions.
`README.md`: the scope-release bullet (R1, R6–R12) and the cohort test row.
## Close: no signal after a proven stop
The brief allows either: don't signal, or check the PID is still the same
engine first. I chose not to signal.
- The binding reaches `stopped` only in `#escalate`, after the verifier
accepted a proof that every member of the cohort, the engine included,
terminated. There is no engine left to kill.
- A check followed by a kill still races PID reuse between the two. Not
signalling has no race.
- The skip keys on binding state `stopped`, not on `#stopped(...)`. A proof
that stops verifying later (R10's withdrawn verifier) doesn't bring the
signal back: the engine was proven dead when the binding moved.
- After the skip, close doesn't wait on `exec.proc.exited` either; nothing
was signalled.
## Tests (`cohort.test.mjs`, needs a systemd user manager)
- R1 (extended): the release entry pushed to the client matches
`evidence.releases` (Darkwing note 1, `nopush`).
- R6 (extended): `releaseCohort` while the engine runs returns
`unavailable`, the shim still answers and the unit is active (Filbert
N1, `relok`).
- R7: a child controller proves a force stop and is SIGKILLed at the
`scope-release` barrier. Both keys read `stopped` with a cohort proof;
the unit is active and the shim answers. A new controller's `start()`
launches, and `evidence.releases` is one entry (`start`, `released`,
`absent`, the unit, the dead controller's claim ID). The unit is gone and
the shim PID dead. The launcher reads `evidence.releases.length` when the
launch begins: 1, so the release finished first (round 2, Filbert N1).
- R8: the same, dying between the keys on the `stopped` finish. `start()`
returns `{ launched: false, classified: { state: "stopped", proofKind:
"cohortProof" } }` and the same single release.
- R9: the force stop's release finds the shim's socket moved and records
`unavailable`; the unit stays. With the socket back, a `recover` with a
confirmation never issued and one with a confirmation confirmed for
`force-stop` are refused `confirmation`, and the release list and unit
are unchanged (round 2). A confirmed `recover` returns
`recovery-eligible`, and the second entry is (`recover`, `released`,
`absent`). The unit is gone.
- R10: the force stop is held at `scope-release`, the verifier is then
withdrawn, and `close({ killEngine: true })` releases nothing and sends
no signal to the engine PID (round 2, Darkwing note 1); the unit is
active. The held
stop then releases on the proof it verified (Filbert N3, `closeany`).
- R11: after a proven stop and its release, `close({ killEngine: true })`
sends no `process.kill` to the recorded engine PID (or its group). A new
controller then starts on the released claim: it launches, and
`evidence.releases` is empty, since the unit reads absent (round 2,
Darkwing note 3).
- R12: a child controller is SIGKILLed with its engine running. A new
controller on a different boot ID gets `stopped` on a boot proof and
releases nothing; the unit is active and the shim answers (Filbert N3,
`noprooffkind`).
Not tested: the retry after `still listed`. `releaseCohort` reads the unit
through `systemdUnits` directly, so a lingering listing can't be faked
without changing `cohort.mjs`. The path is the same as R9's: the retry is
by claim ID, not by outcome. After `still listed` the shim has already
exited, so a retry gets `unavailable` from `hello`, which is recorded.
## Mutation check
Round 2, scratch copies of the final working tree, the full conversation
suite per mutant (`~/dewey-scratch/r51/mut-tools/run.sh`, definitions in
`mutants.py`, logs in `out/`; round 1's in `out-r1/`), 06:17:00Z to
06:36:00Z. After each run the runner lists any shim or fake engine left
under the mutant's TMPDIR. No mutant left one, and no `mosaic-chat-*` unit
was listed afterwards.
| Mutant | Change | Result | Killed by |
|---|---|---|---|
| base | none | 177/177 | (baseline) |
| recconf | the recover release before the pins, target and confirmation checks, round 1's order (Sage's ruling) | 176/1 | R9 |
| recbefore | the recover release above the stop-proof check (Darkwing) | 176/1 | R9 |
| startnoawait | the free-path release isn't awaited (Filbert N1) | 176/1 | R7 |
| noseat | the free path checks only the session head (Darkwing note 2, Filbert N2) | 177/0 | survives: no test has a seat head naming another claim |
| nolookup | `#releaseListed` sends a request for a unit already absent (Darkwing note 3) | 176/1 | R11 |
| nodelete | `#releaseListed` never drops an earlier result (Darkwing) | 176/1 | R9 |
| closeproof | close skips the signal only while `#stopped()` still accepts the stop (Darkwing note 1) | 176/1 | R10 |
| norecover | recover releases nothing | 176/1 | R9 |
| nostart | start releases nothing | 175/2 | R7, R8 |
| nostartfree | no release on the free path | 176/1 | R7 |
| nostartstopped | no release on classify's `stopped` branch | 176/1 | R8 |
| noretry | an earlier result is never retried | 176/1 | R9 |
| closekill | close signals the engine after a proven stop | 175/2 | R10, R11 |
| nopush | the release entry isn't pushed | 176/1 | R1 |
| closeany | close drops only its `#stopped(...)` check (Filbert's form) | 176/1 | R10 |
| noprooffkind | `#releaseScope` drops its `cohortProof` check | 176/1 | R12 |
| relok | `releaseCohort` ignores a refused `release` | 176/1 | R6 |
Darkwing's `nodedupe` (the seat head's same-claim skip dropped) is
equivalent, as their review says: the per-claim memo returns the first
result. Not rerun.
Under noprooffkind the shim itself refused the release, because the engine
in R12 still runs: the start recorded `unavailable` ("the engine cgroup is
not empty"). The shim's `populated 0` check is a second guard; the test
fails on the attempt.
## Gate
Round 2, sequential, on a detached worktree of b7e9efb7 (no change under
`packages/` or `scripts/` since bbb2167f) with the candidate applied
(`git diff HEAD -- packages/conversation`, sha256
`a0aedb21…4c077a102`) and `node_modules` linked, 06:36:16Z to 06:40:31Z,
output in `~/dewey-scratch/r51/gate/out/` (round 1's in `out-r1/`):
- conversation 177/0, webui 22/0, control-board 124/0
- test-auth 15/0, test-conductor 17/0, test-config 24/0, test-discord 66/0,
test-extension-package 18/0, test-foundation 44/0, test-release 14/0,
test-task 98/0
- test-queue 148/0 (node) and 27/0 (shell). Its canonical-root check skips
in a worktree, as in every worktree gate.
After the gate no `mosaic-chat-*` unit was listed and no shim or fake
engine was running. The worktree is removed and `core.hooksPath` is unset.
## Follow-ups (bounded, not fixed)
- `close({ killEngine: true })` on an `uncertain` binding whose engine
exited on its own (EOF) still signals the recorded PID. Row 52 moves that
binding to `stopped` through the engine-exit proof, which removes most of
the window; a binding that stays `uncertain` keeps it.
- The retry after `still listed` (above) has no test.
- The free path's seat-head release (a different session's controller that
died in the crash window on the same seat) has no test; `noseat`
survives. It needs a fixture with two sessions on one seat.
- `#releaseListed` reads the unit once before each retry. A unit that is
listed only because systemd hasn't collected it yet gets one more request,
which ends `unavailable` and is recorded. Harmless; it costs one `hello`.