Files
stack/agents/dewey/work/queue-50/evidence.md
T
jason.woltjeandClaude Opus 5.5 a9cc522a0c conversation: release a cohort scope after a proven stop (row 50, #1536)
releaseCohort sends `release` only to a scope shim that answers hello
with the recorded invocation ID, then waits for systemd to drop the
unit. The controller releases once per claim, after a proven force stop
and on close of a proven-stopped binding; every uncertain path keeps the
scope as evidence. The harness's killShims refuses units outside
^mosaic-chat-, R5 checks liveShims positively, and K19's wait on
proc.exited is bounded.

Dewey's candidate, manifest 375594fc (8 files), approved by Filbert
(comment 27053) and Darkwing (comment 27055). Normal engine exit still
leaves the scope; the follow-up is #1537.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-10-10 00:14:00 -05:00

8.2 KiB
Raw Blame History

Row 50 (#1536): scope release after a force stop

Dewey, 2026-10-10. Brief: docs/plans/2026-10-10_cohort-scope-release.md. Base 16a8d038. The candidate touches seven files in packages/conversation/ (two in src/, four in tests/, the README) and this file. shim.mjs is unchanged. The force-stop phases, the claim protocol and the pgroup fallback are untouched.

A normal engine exit

The brief asks whether a normal engine exit leaves the shim. It does.

Probe (~/dewey-scratch/r50/probe-exit.mjs, output in ~/dewey-scratch/r50/logs/probe-exit.json): a ScopeLauncher scope running /bin/sleep 0.3, read 2010 ms after launch.

  • The engine is gone; the shim is alive and its socket is there.
  • events reads populated 0, frozen 0; hello reports engineExit {code: 0, signal: null}.
  • systemctl --user show reads the unit loaded, active, with the same invocation ID.
  • After release, the shim and socket are gone and the unit reads not-found.

So a scope whose engine exits on its own stays up with an idle shim until something sends release. In the controller the engine's EOF ends the binding uncertain, with no stop recorded.

I didn't add a release at exit. Two reasons:

  • A collected scope is an absent observation, not proof that the cohort ended (claim.mjs classify). The release has to follow a proof, not replace one.
  • A later force stop reads the scope through the shim. With the scope collected, its hello fails, the stop returns unavailable and the claim stays uncertain for good.

The path that does end it: a force stop on that binding finds the cohort empty, proves it (members: []), records stopped and releases (R4).

The change

src/cohort.mjs

  • releaseCohort({ kind, unitName, invocationId, shimSocket, waitMs }). It returns unavailable for a cohort that isn't a scope, a shim that doesn't answer hello, a hello whose invocation ID isn't the recorded one (or no recorded ID), and a refused release. Otherwise it polls systemdUnits.lookup for up to waitMs and returns { outcome: "released", unit: "absent" | "still listed" }.
  • The header comment says the scope outlives the stop until releaseCohort ends it.

src/controller.mjs, force-stop and close paths only

  • #releaseScope(why, claim) acts only on a claim record with state stopped, a cohortProof and a shim. It sends one request per claim ID (scopeReleases), so the force stop and close share it. It never rejects: a thrown error becomes an unavailable result. Each result is an entry in evidence.releases (kind, claim, unitName, why, outcome, unit or reason, at) and is pushed as evidence.
  • The force stop calls it last, after the binding, the stopped event and the stop evidence, for the claim it proved (taken before the pause). A new crash barrier scope-release sits before it. Every fail path returns before it.
  • close() calls it when the binding is stopped and #stopped() holds for its stop. This covers a stop whose release didn't run. The comment now says close is never a claim release.

README.md: the shim row, a "Scope release" bullet after "Force stop" (R1–R6, the uncertain case, the normal exit, the crash window), and the cohort row of the test table.

Harness notes from #1536 comment 27045

  • N1: R5 launches a mosaic-chat- scope and asserts that liveShims returns exactly [{ pid, socket, unit }] for it, that reap leaves shimsGone empty, that the shim process exits, and that the unit is gone.
  • N2: killShims skips any shim whose unit doesn't match ^mosaic-chat- and returns those as refused. It signals neither the scope nor the PID. R5 launches a shim in an r5-not-mosaic-<id> scope: killShims returns it, and the shim and unit are still alive afterwards. That shim is ended through its own kill and release ops.
  • N3: K19's finally waits on proc.exited through within(p, 10000), which returns null on timeout and clears its timer. A hang fails the test at 10 s instead of holding the file.

claim.test.mjs W5 and W13: comments only. Their proven stops now release their scopes; the early reap stays for anything else they left.

fake-pi.mjs: an exit op, so R4 can end the fake engine normally after its reply is written.

Tests (cohort.test.mjs, needs a systemd user manager)

  • R1: after a proven force stop, evidence.releases is one entry (force-stop, released, the unit, the claim ID), the unit is absent, the shim PID is dead and hello fails.
  • R2: a gate holds the scope-release barrier. After the force stop the unit is still active and nothing is released. close() releases it (close, released, absent). After the barrier opens, there is still exactly one entry.
  • R3: two stops that end uncertain. One launcher runs the real force stop and then reports it unavailable. One verifier rejects the cohort proof. In both, after close(), nothing is released, the unit is active, and the shim answers hello with the recorded invocation ID.
  • R4: the normal exit above, through the controller. The binding is uncertain, the shim reports exit code 0, the unit is active and nothing is released. The force stop then proves members: [], records stopped, and releases. The unit is gone.
  • R5: N1 and N2, above.
  • R6: releaseCohort releases nothing for another invocation ID, a null ID, a pgroup cohort or a fake cohort; the shim still answers. The correct call returns { outcome: "released", unit: "absent" }.

Mutation check

Scratch copies of the final working tree, the full conversation suite per mutant (~/dewey-scratch/r50/mut-tools/run.sh, definitions in mutants.py, logs in out/). After each run the runner lists any shim left under the mutant's TMPDIR. No mutant left one.

Mutant Change Result Killed by
base none 171/171 (baseline)
norelstop the force stop doesn't release 169 pass, 2 fail R1, R4
norelclose close doesn't release 170/1 R2
norel neither releases 168/3 R1, R2, R4
relfail the uncertain path releases, and #releaseScope checks only the shim 170/1 R3
closeany close releases in any binding state, and the same weak guard 170/1 R3
nomemo no once-per-claim memo 170/1 R2
nohello releaseCohort skips the invocation check 170/1 R6
nokind releaseCohort skips the kind check 170/1 R6
n2 killShims signals any unit 170/1 R5
n3hang K19 waits on a promise that never settles 170/1 K19, at the 10 s bound

relfail and closeany weaken the guard as well, because with it intact the extra call is refused, and the mutant would test the guard and not the call site.

Gate

Sequential, on a detached worktree of 16a8d038 with the candidate applied (git diff 16a8d038 -- packages/conversation, sha256 ec5ce9ae…e0e3f26a) and node_modules linked, 04:42:13Z to 04:46:41Z, output in ~/dewey-scratch/r50/gate/out/:

  • conversation 171/0, webui 22/0, control-board 124/0
  • test-auth 15/0, test-conductor 17/0, test-config 24/0, test-discord 66/0, test-extension-package 18/0, test-foundation 44/0, test-release 14/0, test-task 98/0
  • test-queue 148/0 (node). Its canonical-root check skips in a worktree, as in every worktree gate.

After the gate no mosaic-chat-* unit was listed and no shim.mjs was running. The worktree is removed and core.hooksPath is unset.

A first run (out-r1/) left out the node_modules link: conversation failed 114 on engine-pin-mismatch and discord failed its pi binary check. Both are the missing link, not the candidate.

Follow-ups (bounded, not fixed)

  • Crash window: a controller killed after the claim records stopped and before the release leaves the scope and its shim. A restart classifies the pair stopped and doesn't release it. The fix belongs in the start/recovery path, which this row doesn't own. The README states the gap.
  • Pre-existing: close({ killEngine: true }) after a proven stop still SIGKILLs the recorded engine PID, which by then may belong to another process.
  • Not checked: whether W5's traced barrier list now includes scope-release. That depends on whether the release reaches the barrier before the probe closes. W5 passes either way.