releaseCohort sends `release` only to a scope shim that answers hello with the recorded invocation ID, then waits for systemd to drop the unit. The controller releases once per claim, after a proven force stop and on close of a proven-stopped binding; every uncertain path keeps the scope as evidence. The harness's killShims refuses units outside ^mosaic-chat-, R5 checks liveShims positively, and K19's wait on proc.exited is bounded. Dewey's candidate, manifest 375594fc (8 files), approved by Filbert (comment 27053) and Darkwing (comment 27055). Normal engine exit still leaves the scope; the follow-up is #1537. Co-Authored-By: Claude Opus 5.5 <[email protected]>
8.2 KiB
Row 50 (#1536): scope release after a force stop
Dewey, 2026-10-10. Brief: docs/plans/2026-10-10_cohort-scope-release.md.
Base 16a8d038. The candidate touches seven files in packages/conversation/
(two in src/, four in tests/, the README) and this file. shim.mjs is
unchanged. The force-stop phases, the claim protocol and the pgroup fallback
are untouched.
A normal engine exit
The brief asks whether a normal engine exit leaves the shim. It does.
Probe (~/dewey-scratch/r50/probe-exit.mjs, output in
~/dewey-scratch/r50/logs/probe-exit.json): a ScopeLauncher scope running
/bin/sleep 0.3, read 2010 ms after launch.
- The engine is gone; the shim is alive and its socket is there.
eventsreadspopulated 0,frozen 0;helloreportsengineExit {code: 0, signal: null}.systemctl --user showreads the unitloaded,active, with the same invocation ID.- After
release, the shim and socket are gone and the unit readsnot-found.
So a scope whose engine exits on its own stays up with an idle shim until
something sends release. In the controller the engine's EOF ends the
binding uncertain, with no stop recorded.
I didn't add a release at exit. Two reasons:
- A collected scope is an absent observation, not proof that the cohort
ended (
claim.mjsclassify). The release has to follow a proof, not replace one. - A later force stop reads the scope through the shim. With the scope
collected, its
hellofails, the stop returnsunavailableand the claim staysuncertainfor good.
The path that does end it: a force stop on that binding finds the cohort
empty, proves it (members: []), records stopped and releases (R4).
The change
src/cohort.mjs
releaseCohort({ kind, unitName, invocationId, shimSocket, waitMs }). It returnsunavailablefor a cohort that isn't a scope, a shim that doesn't answerhello, ahellowhose invocation ID isn't the recorded one (or no recorded ID), and a refusedrelease. Otherwise it pollssystemdUnits.lookupfor up towaitMsand returns{ outcome: "released", unit: "absent" | "still listed" }.- The header comment says the scope outlives the stop until
releaseCohortends it.
src/controller.mjs, force-stop and close paths only
#releaseScope(why, claim)acts only on a claim record with statestopped, acohortProofand a shim. It sends one request per claim ID (scopeReleases), so the force stop and close share it. It never rejects: a thrown error becomes anunavailableresult. Each result is an entry inevidence.releases(kind,claim,unitName,why,outcome,unitorreason,at) and is pushed as evidence.- The force stop calls it last, after the binding, the
stoppedevent and the stop evidence, for the claim it proved (taken before the pause). A new crash barrierscope-releasesits before it. Everyfailpath returns before it. close()calls it when the binding isstoppedand#stopped()holds for its stop. This covers a stop whose release didn't run. The comment now says close is never a claim release.
README.md: the shim row, a "Scope release" bullet after "Force stop"
(R1–R6, the uncertain case, the normal exit, the crash window), and the
cohort row of the test table.
Harness notes from #1536 comment 27045
- N1: R5 launches a
mosaic-chat-scope and asserts thatliveShimsreturns exactly[{ pid, socket, unit }]for it, thatreapleavesshimsGoneempty, that the shim process exits, and that the unit is gone. - N2:
killShimsskips any shim whose unit doesn't match^mosaic-chat-and returns those asrefused. It signals neither the scope nor the PID. R5 launches a shim in anr5-not-mosaic-<id>scope:killShimsreturns it, and the shim and unit are still alive afterwards. That shim is ended through its ownkillandreleaseops. - N3: K19's
finallywaits onproc.exitedthroughwithin(p, 10000), which returnsnullon timeout and clears its timer. A hang fails the test at 10 s instead of holding the file.
claim.test.mjs W5 and W13: comments only. Their proven stops now release
their scopes; the early reap stays for anything else they left.
fake-pi.mjs: an exit op, so R4 can end the fake engine normally after
its reply is written.
Tests (cohort.test.mjs, needs a systemd user manager)
- R1: after a proven force stop,
evidence.releasesis one entry (force-stop,released, the unit, the claim ID), the unit is absent, the shim PID is dead andhellofails. - R2: a gate holds the
scope-releasebarrier. After the force stop the unit is still active and nothing is released.close()releases it (close,released,absent). After the barrier opens, there is still exactly one entry. - R3: two stops that end
uncertain. One launcher runs the real force stop and then reports itunavailable. One verifier rejects the cohort proof. In both, afterclose(), nothing is released, the unit is active, and the shim answershellowith the recorded invocation ID. - R4: the normal exit above, through the controller. The binding is
uncertain, the shim reports exit code 0, the unit is active and nothing is released. The force stop then provesmembers: [], recordsstopped, and releases. The unit is gone. - R5: N1 and N2, above.
- R6:
releaseCohortreleases nothing for another invocation ID, a null ID, apgroupcohort or afakecohort; the shim still answers. The correct call returns{ outcome: "released", unit: "absent" }.
Mutation check
Scratch copies of the final working tree, the full conversation suite per
mutant (~/dewey-scratch/r50/mut-tools/run.sh, definitions in
mutants.py, logs in out/). After each run the runner lists any shim
left under the mutant's TMPDIR. No mutant left one.
| Mutant | Change | Result | Killed by |
|---|---|---|---|
| base | none | 171/171 | (baseline) |
| norelstop | the force stop doesn't release | 169 pass, 2 fail | R1, R4 |
| norelclose | close doesn't release | 170/1 | R2 |
| norel | neither releases | 168/3 | R1, R2, R4 |
| relfail | the uncertain path releases, and #releaseScope checks only the shim |
170/1 | R3 |
| closeany | close releases in any binding state, and the same weak guard | 170/1 | R3 |
| nomemo | no once-per-claim memo | 170/1 | R2 |
| nohello | releaseCohort skips the invocation check |
170/1 | R6 |
| nokind | releaseCohort skips the kind check |
170/1 | R6 |
| n2 | killShims signals any unit |
170/1 | R5 |
| n3hang | K19 waits on a promise that never settles | 170/1 | K19, at the 10 s bound |
relfail and closeany weaken the guard as well, because with it intact the extra call is refused, and the mutant would test the guard and not the call site.
Gate
Sequential, on a detached worktree of 16a8d038 with the candidate applied
(git diff 16a8d038 -- packages/conversation, sha256
ec5ce9ae…e0e3f26a) and node_modules linked, 04:42:13Z to 04:46:41Z,
output in ~/dewey-scratch/r50/gate/out/:
- conversation 171/0, webui 22/0, control-board 124/0
- test-auth 15/0, test-conductor 17/0, test-config 24/0, test-discord 66/0, test-extension-package 18/0, test-foundation 44/0, test-release 14/0, test-task 98/0
- test-queue 148/0 (node). Its canonical-root check skips in a worktree, as in every worktree gate.
After the gate no mosaic-chat-* unit was listed and no shim.mjs was
running. The worktree is removed and core.hooksPath is unset.
A first run (out-r1/) left out the node_modules link: conversation
failed 114 on engine-pin-mismatch and discord failed its pi binary
check. Both are the missing link, not the candidate.
Follow-ups (bounded, not fixed)
- Crash window: a controller killed after the claim records
stoppedand before the release leaves the scope and its shim. A restart classifies the pairstoppedand doesn't release it. The fix belongs in the start/recovery path, which this row doesn't own. The README states the gap. - Pre-existing:
close({ killEngine: true })after a proven stop still SIGKILLs the recorded engine PID, which by then may belong to another process. - Not checked: whether W5's traced barrier list now includes
scope-release. That depends on whether the release reaches the barrier before the probe closes. W5 passes either way.