# CHAT-03 brief: live adapters and mediated terminal (#1507, row 5) Author: Dewey, 2026-09-26/27. Final text for pinning: R3 (`2c5be6b4`, frozen as `BRIEF-r3-2c5be6b4.md`) with the one edit lead decision 32 orders. It is not a review round (lead decision 27). This is the brief only. It includes no source, no contract edits and no seat changes. R1 (`5dd447f7`) is frozen as `BRIEF-r1-5dd447f7.md`, and R2 (`5c5b45a2`) as `BRIEF-r2-5c5b45a2.md`. §0 lists what changed. Jason chartered CHAT-03 on 2026-09-26 (`docs/plans/2026-09-26_lead-decisions.md` item 22). Plan row: `docs/plans/2026-09-13_webui-session-chat.md` line 215. It depends on CHAT-01 (`28d4e98a`) and implements part of the CHAT-01C companion (`b023841c`). The layout follows `docs/plans/BRIEF-TEMPLATE.md`, and the depth follows the CHAT-02 brief (`agents/dewey/work/chat-02/BRIEF.md`, sha256 `636b0fac…`). The six carry-forwards Sage set are mapped to their evidence under "Carry-forwards". Line numbers and hashes below are at `f2b9e622`. Between `40a02d2b` (the R2 base) and `f2b9e622`, no contract, source or pinned Pi file changed. Queue A1 landed (`34a72af9`) with `docs/plans/BRIEF-TEMPLATE.md`, lead decisions gained items 26 and 27, and the goals review (`docs/plans/2026-09-27_goals-review.md`) was added. Lead decision 27 made R3 the last review round, and its answer to anything pinned Pi can't prove is to refuse or report unknown, never new machinery to prove it. Sage rescoped CHAT-03 against Gate E from R3's section map (lead decision 30), ruled on Rocko's R3 finding (lead decision 31) and ordered this edit (lead decision 32). ### 0. Changes since R3 (lead decision 32) Filbert approved R3 on the sections Gate E keeps, with no blocking finding (`agents/filbert/work/chat-03-brief-review-r3-2026-09-27.md`, `48447592`). Rocko closed his R2 finding and raised one blocking finding in the seal (`agents/rocko/work/chat-03-r3-adversarial-2026-09-27.md`, `19e3fcff`), which lead decision 31 resolves. This edit does the five things lead decision 32 lists, and nothing else. | Item | What changed | |---|---| | 1. Lead decision 30 cuts, removed and not reworded | Removed: C-1, C-2 and C-4; increment I2 and §7's Pi dialogs (P1, P2, P4–P6); increment I1b (C-5 moves to CHAT-04, and the interim rule stays); the §2 idle drift check, W10, W18, W19, mutant 25, the drift choice and the `session-drift` name. H5–H8 stay for Claude in I4. References to removed items were adjusted where they stood: What ships, E5, carry-forwards 1 and 4, §11, Limits 5 and 11, Out of scope and Gate. Two sentences were kept by moving them: the claim's reach, into the live-session guard bullet, and P3's disabled Pi dialog, which is now §7's only Pi rule. | | 2. Lead decision 31: no explicit extensions | The seal is `--no-extensions`, `--no-prompt-templates`, `--no-themes` and no `--extension`. The registry, the input-silent review and the tree hashes are removed, not replaced with pinning. N5 and N15 are fake-only. N24 and mutant 37 refuse any `--extension`. Limit 11 and §11 say explicit extensions return with goal in CHAT-06. | | 3. Filbert n1: the Pi pin | Pi is pinned by the `package-lock.json` integrity of `@earendil-works/pi-coding-agent` 0.85.1, checked against npm's installed record. The listed `dist/core` and `dist/extensions` files are now labelled as citation sources, since `pi` runs `dist/bundle`. The built-in `llama.cpp` is pinned with the package. | | 4. Filbert n2: the clear's own `queue_update` | O5 counts the empty `queue_update` Pi's clear emits before its response as the controller's own. N10's fake emits it, and N25 proves that an ordinary Interrupt reconciles. | | 5. Rocko's R3 note | A build note under What ships: the `no-turn` cleanup lifts only its own fence. H10 covers the force-stop, overlap and revocation branches. | Filbert's n3 (an `aborted` with no stop in progress) isn't in lead decision 32, so it isn't applied. ### 0a. Changes from R2 to R3 (historical) R3 answers Rocko's review of R2 (`agents/rocko/work/chat-03-r2-adversarial-2026-09-26.md`, `07b938fb`), Filbert's review of R2 (`agents/filbert/work/chat-03-brief-review-r2-2026-09-26.md`, `d149e8cc`), Sage's ruling on deviation V-1 (`docs/plans/2026-09-26_lead-decisions.md` item 25) and Sage's rule for all three R2 blocking findings: a receipt settles only on positive evidence tied to its own prompt. Rocko closed R1 findings 1–7, and Filbert closed B2–B6 and his R1 notes. The full disposition table is in `REVIEW-REQUEST.md`. | Finding | What changed | |---|---| | Rocko R2 finding 1 (blocking): a settled slot, an idle engine and empty clears don't prove an interrupted turn | §3 splits the old slot rule. Rule 4 settles the item's receipt from its own evidence, and routes a late ack through the same table. Rule 5 classifies the stop: Interrupted, Completed first, Failed on its own, No run, Unknown. A normal completion stays `finished`. Only Interrupted gives `turnState: interrupted` and reconciles; `input-reconciled` isn't borrowed. The rest leave the stop `uncertain`, admission closed and force stop available, until contract item C-5 (new, separate, reviewed). Rule 6 (reopening) requires Interrupted until C-5. Interrupt with nothing running refuses `no-turn`. N14 and N15 cover Rocko's completion and handled-ack schedules, and N3 now covers his no-run preflight error. N16–N18 add `no-turn`, Failed on its own and Unknown. N3, N4 and N9 assert the receipt and the stop. A receipt never moves back from `working`, and rule 8 states the receipt order. Mutants 30–33. Limit 9. | | Rocko R2 note 2: N11 wording | The `ack-without-start` note says Pi's ordering is local to Pi, and that the ordered output chain says nothing about which prompt a settle belongs to. With F1, the outcome is now `delivery-unknown`, never `failed` (next row). The `pi-agent-core` run-failure lines (349–364) are cited, and its two files are pinned by hash. N11 adds the unparseable-line case. | | Filbert R2 F1 (blocking): an extension run can overlap a Mosaic prompt's preflight, so order proves nothing | §3 R3-1 drops the "at once or not at all" premise and states the verified window (lines 860–949 and their awaits), the acked loser, the swallowed throw, the spurious settle and the multi-pair run. A binding needs a sealed engine: `--no-extensions` and a registry of hash-pinned, input-silent extensions, otherwise `unsealed-engine`. The seal is named as the basis for attributing by order. Overlap signals O1–O6 close admission and make the binding `uncertain` (`run-overlap`). R2's `failed` row for a run that fails before its user message becomes `delivery-unknown` / `ack-without-start`, and `working` requires that no overlap signal has been seen. A misattribution to `working` that a later signal exposes is shown as outcome unknown. Goal can't bind in CHAT-03 (Limit 11, a CHAT-06 item). N8, N10, N13 and N19–N21, N24; mutants 34–37. | | Filbert R2 F2 (blocking): an empty `clear_queue` doesn't prove that nothing was removed | Rule 3 narrowed to Pi's steer and follow-up queues, with the seal as the basis. C-1's premise doesn't survive, so C-1 is withdrawn and restated for CHAT-06 ("Contracts implemented"), and I1b is C-5 only. Limit 7 names the agent-level and `nextTurn` gaps. N22 and N23. | | Filbert R2 notes n1–n4 | n1: the ordering note cites `rpc-mode.js` 28–29 and `output-guard.js` 71, and makes no transport claim. n2: recovery reads the file before `get_entries` (§2). n3: the drift baseline is taken after load; W18 adds the truncated-line, empty-file and migration cases; "before the first assistant message" applies to new files only. n4: H17 cites CHAT-01C 92–102; §1 cites `rpc.md` 56–65 and 78–104; `agent.js`, `output-guard.js` and the goal extension are pinned. | | Rocko R2 remark on W20 | A marker on either key counts, and completing a pair copies it. W20 runs with the marker on one key only. | | Sage, lead decision 25: V-1 accepted with limits | V-1 is accepted. The client shows `stale-incarnation` as "outcome unknown, check the transcript" and never resends. H23 proves that no request pending across a restart reaches the engine. V-1 closes only when CHAT-04's durable receipts restore the contract behavior, and CHAT-04's brief must list it. Gate, carry-forward 4, §11 and Limit 10 updated. | | Consistency passes | Two passes over the R3 draft found 22 defects of wording and cross-reference. A third pass over the finished draft found 17, including three where a fixture's expected result didn't follow from the rules (N6, N8, N20). All are fixed. | | Base | Line numbers and hashes move to `f2b9e622`. No contract, source or pinned Pi file changed since `40a02d2b`. Queue A1 is now committed, so the overlap table names its commit. | ### 0b. Changes from R1 to R2 (historical) The rule numbers in this table are R2's. R2's rule 6 is R3's rule 7. R2 answered three reviews of R1: - Rocko's adversarial review (`agents/rocko/work/chat-03-r1-adversarial-2026-09-26.md`, `89752c2b`) and his addendum narrowing finding 3 (`…-addendum-2026-09-26.md`, `e5b6bd00`); - Filbert's review (`agents/filbert/work/chat-03-brief-review-2026-09-26.md`, `ec00544e`). Sage's rule for R2: where a guarantee can't be proven on pinned Pi, the brief refuses or reports uncertainty instead of claiming it. The full R2 disposition table is in `REVIEW-REQUEST-r2-c75ad86f.md`. | Finding | What changed | |---|---| | Filbert B1, Rocko 3 and addendum: R3-1 premise | §3 R3-1 rewritten from the source. A mediated prompt can't enter Pi's queues: without `streamingBehavior` it throws while `isStreaming` (`agent-session.js` lines 860–863, 617, 773, 348). `clear_queue` returns only external input. The residual risks are a fence during preflight and an ack with no run. No text-identity machinery. External input queued after the last clear is stated as a limit (rule 6, Limit 7). N1–N13 rebuilt. | | Filbert B2, Rocko 6: no entry IDs on the stream | §2: the drift check compares disk entries with the engine's own `get_entries` list (`rpc.md` lines 717–745), not with stream IDs. It runs at idle. §3 advertises replay as unavailable. W10, W18 and W19. | | Filbert B3: schema gaps | C-1 is repurposed as a `turnProof` field for external items that `clear_queue` removed. Until it lands, a non-empty clear leaves the stop `uncertain`. C-4 is an optional event type for unknown native events, with an interim rule. | | Filbert B4 | §2 live-session guard (`live-session-refused`, symlinks resolved), G1–G3, mutant 14. CHAT-07 lifts it. | | Filbert B5: order and gate | Each increment starts when its "Needs first" column is met. Gate adds C-1 adoption or carry, a B1-not-passed trigger and owner, and what happens if C-2 is declined. Increments renamed I1–I4 so they don't collide with CHAT-03D. | | Filbert B6, Rocko 1: claim gaps | §2 rewritten: one claim ID across both keys, atomic `link()` publication, the more conservative key wins, restart rules for `reserved` and half-done pairs, a scope named from the claim ID, and a no-unit path to `stopped` that a spawn marker closes. W4, W5, W12–W17, W20. The restart dedup is recorded as deviation V-1 for Sage. | | Rocko 2: pending prompts, partial writes | §1 pending dispatch slot and three write outcomes. A partial or failed write poisons the pipe. H3, H18–H20. | | Rocko 4: cohort proof | §6 rewritten: shim-held scope, `engine` child cgroup, invocation-ID epoch, freeze, enumerate, `cgroup.kill`, unavailable distinct from empty, the K13 migration test, and fixture-grade verification (CHAT-01 lines 330–333). K1–K16. | | Rocko 5: dedup across restart | §1 incarnation token and `stale-incarnation`. H21, H22. Deviation V-1. | | Rocko 7 (non-blocking) | §6 single-use eligibility with a new reservation. K8, K17, K18. | | Filbert's notes N1–N12 (non-blocking; not the N fixtures) | Skills are live (§4, S2). Refusal names `generation` and `controller` reused. Citations corrected. B1 floating sources fixed by hash. Smoke isolated. Suite count by glob. | ## CHAT-03: live adapters and mediated terminal ### Problem CHAT-02 made every repository Pi conversation readable. Nothing can drive a conversation yet except the old paths: the Pi TUI in a tmux pane, and board Reply or agent-send, which paste into that pane. Those paths have no controller, no generation fence and no stop proof. At 2026-09-26T20:10:57Z a leftover `/` in Pi's composer was concatenated in front of a board reply. Pi read the result as plain text that time, but the same path could run a reply as a slash command, because `tools/tmux/send-message.sh` (lines 45–54) pastes onto whatever the composer holds and then presses Enter (`docs/plans/DEFERRED.md`, "Board send can turn into a Pi slash command"). CHAT-01 defines the records and commands for one controller per conversation. It defers the writer-claim record (`chat-01/README.md:342`), and CHAT-01C defers R3-1 (`chat-01c/README.md:217`). Lead decisions item 8 moves both here. The Claude catalogue still refuses `unsupported-harness` (`packages/conversation/src/reader.mjs`: constant at line 24, refusals at lines 51 and 65). Jason ruled that CHAT-03 owns it after B1. Claude on this host reports 2.1.283, but the plan's evidence named 2.1.269 (plan line 159). The binary changes under any adapter that does not pin it. ### Owner and reviewer - **Source author: ``.** The plan has Darkwing name the backend author (line 215; lead decisions item 22). Darkwing is on queue A1/A2, so the author is named after this brief is approved. The brief does not pick one. - Brief author: Dewey. - Filbert reviews this brief, then the exact candidate of each increment. - Rocko does an adversarial pass on §5 (control races) and §6 (stop and recovery), in this brief and later in the code of each increment. Rocko also reviews the B1 packet (§8), because it is evidence that a false pass would turn into a wrong adapter. - No one reviews their own work. If Darkwing names Filbert or Rocko as author, Sage names a replacement for that review. - Tracking: #1507. ### Files owned The source author may create or change only these paths. | Path | Contents | |---|---| | `packages/conversation/src/**` | New modules for the controller, claim, supervisor shim, Pi adapter, events, transport and mediated terminal. The Claude parser and adapter come in I4. `reader.mjs` changes only to lift the Claude refusal for the pinned version. | | `packages/conversation/tests/**` | Tests, fake engines, fixture extensions and redacted recordings | | `packages/conversation/README.md`, `packages/conversation/package.json` | Documentation and the new refusal names. No new dependencies. | | `agents//work/chat-03/**` | Review packets, evidence and the B1 packet | Module names inside `src/` are the author's call (plan lines 150–151). Any path outside this table needs a brief amendment. Excluded, and why: - **`tools/tmux/**`.** Plan line 244 excludes it, and §4 shows the mediated path needs no change there. - **`packages/control-board/**`.** No board edit is needed (§4, S6). Tests import `replyToRow` read-only by relative path, the same way the queue-as-data plan imports `scan.mjs`. - **`packages/seat/**`, `scripts/agent-host-dev.sh`, `agents/*/launch.sh`.** No seat migrates, so no seat or launcher changes. The launcher handoff belongs to CHAT-07. - **`packages/webui/**`.** The chat UI is CHAT-05. - **`docs/plans/chat-0*/**`.** Contracts change only through the C items below. - **`roles/**`, `contracts/**`, root files, auth and `~/.mosaic`.** No overlap with queue-as-data or the ledger: | Owner | Paths | Overlap | |---|---|---| | A1, Darkwing, committed `34a72af9` (`agents/darkwing/work/queue-a1/build-manifest.sha256`) | `packages/queue/**`, `scripts/queue-commit.sh`, `scripts/test-queue.sh`, `scripts/git-hooks/pre-commit`, `docs/plans/BRIEF-TEMPLATE.md` | none | | A2, Darkwing (Filbert's queue-as-data plan `282fabbb…` §8.1 and lines 130 and 776–777; lead decisions item 20) | `scripts/mosaic` dispatch, `docs/plans/queue.json`, `docs/plans/QUEUE.md`, `scripts/test-{darkwing,rocko}-launch.mjs` | none. The mediated terminal runs as `node packages/conversation/src/.mjs`. It is not a `scripts/mosaic` verb. | | Piece B, after A | `AGENTS.md`, `agents/*/CONTEXT.md` | none | | Ledger | `packages/ledger/**` | none | The author may import the process-identity helpers in `packages/discord/src/journal.mjs` read-only, as A1 does. That changes neither package. ### Contracts implemented These are unchanged since the commits named. The sha256 values are at `f2b9e622`, and `git status` shows no local edits. | Path | sha256 | Commit | |---|---|---| | `docs/plans/chat-00/README.md` | `991663e607404c2f022716bd45b40d5b3092d53c3f9f61bbafd455c0659b832f` | `370823b3` | | `docs/plans/chat-00/sources.json` | `1a07ae88de45fd3219eca10acae13637ca75416cb1c598aa9080825bd7598af9` | `370823b3` | | `docs/plans/chat-00/fixtures.json` | `ad9d4fe94131b2c4bb30691ad846698c5b0818f0935d20a38de73fa3f267bd80` | `370823b3` | | `docs/plans/chat-00/check.mjs` | `568f004f5a0cfce64d65ff202c420ab491e716919a93989377a7bb9e7b8be622` | `370823b3` | | `docs/plans/chat-01/README.md` | `61aba7d60f380ff8a135a04a2e851f11c250795ead11cca2e594566849483163` | `28d4e98a` | | `docs/plans/chat-01/contracts.schema.json` | `38382e08c97635f8864e6953b6cf44ec324040397a1aa8fc3f86e279abe36db1` | `28d4e98a` | | `docs/plans/chat-01/fixtures.json` | `403c8ae91963a379de17805380bc1425bc4e96c1ef34094385b703d17cf5ce7d` | `28d4e98a` | | `docs/plans/chat-01/check.mjs` | `2e164e4bfa61963bb5dd8639e26e67407e64d19278cc56ed8a4f77e430f1dee5` | `28d4e98a` | | `docs/plans/chat-01c/README.md` | `63d9f9edffc2aa1aa3bc99364b864e8c6f234eedc8ad2f9da231531970d54250` | `b023841c` | | `docs/plans/chat-01c/contracts.schema.json` | `da132f0a02f29281344cc5350396af53c893d7895308bc430f9bf3ca3b3785b9` | `b023841c` | | `docs/plans/chat-01c/fixtures.json` | `00639a219b00b67c5e0ab904163a6d945481f6b0df45a058b32e00f1f1326611` | `b023841c` | | `docs/plans/chat-01c/check.mjs` | `5971ed5f11a32f0dff5720db03d8d4bc4c007a5d2654ba50726bb35ba753eba6` | `b023841c` | Design inputs, which CHAT-03 does not implement in full: `docs/plans/foundation-v1-candidate/RUNTIME.md` §§3–5 (`b1a2b4d0…`) and the plan itself (`48142829…`). What CHAT-03 takes from each contract: - **CHAT-01.** The records `binding`, `connection`, `clientRequest`, `request`, `receipt`, `event`, `stop`, `cohortProof`, `effectReport`, `turnProof`, `nativeDecision`, `approval` and `confirmation`. The commands `observe`, `prompt`, `takeover`, `acquire-recovery-control`, `approval`, `interrupt`, `force-stop`, `recover`, `issue-confirmation` and `answer-confirmation`. The draft, upload and queue-edit commands belong to CHAT-04. - **CHAT-01C.** Only the confirmation restoration rules (lines 92–104). The private-state pages, upload ranges and `recover-refused-draft` need CHAT-04's durable store. The CHAT-01 text at line 342 stays as published. Lead decisions item 8 records the move. **Contract changes, each a separate reviewed item.** Each one lands and is approved before the code that needs it. Sage assigns the author; Filbert and Dewey review, as they did for CHAT-01C. - **C-3: Claude record fields.** Needed only if B1 shows that CHAT-01 lacks a field, for example the capability-negotiation result on `binding`. Otherwise C-3 is void. - **C-5: interrupt outcomes other than an interrupted turn (R3).** `turnProof.turnState` allows `interrupted`, `unknown` and `input-reconciled`, and `reconcile-interrupt` passes only with `interrupted` (`check.mjs` line 315). A turn that completed or failed before the abort took effect, or an Interrupt that found no run, has no honest value. C-5 adds values for those (for example `completed-before-interrupt`, `failed-before-interrupt`, `no-run-at-interrupt`), allowed by `reconcile-interrupt` without recording `turn-interrupted`. Until C-5 lands, those stops stay `uncertain` (§3 rule 5). C-5 moves to CHAT-04 with its own contract review (lead decision 30), so CHAT-03 ships that interim rule only. **Deviation V-1, accepted with limits (lead decision 25).** CHAT-01 line 145 says an exact retry after reconnect returns the existing receipt. CHAT-01C line 212 says a reconnect or restart retry returns the original result, for recovery requests. CHAT-03 has no durable receipts, so an exact retry across a controller restart refuses `stale-incarnation` (§1). That is safe, because nothing is replayed, but it departs from the contract. It is recorded here the way the CHAT-02 `newer` deviation was. Sage accepted it for CHAT-03 with three limits, which the brief carries: - the client shows `stale-incarnation` as "outcome unknown, check the transcript" and never resends on its own (§1); - a fixture proves that no retry after a restart reaches the engine (H21, H23); - V-1 closes only when CHAT-04's durable receipts restore the contract behavior. CHAT-04's brief lists that as a required item. Not contract changes: - The writer-claim storage record (§2) is internal and never crosses the wire. It stores the CHAT-01 `binding` fields it needs. If one is missing, that becomes a C item, not an edit. - New refusal and reason names (`busy`, `stale-incarnation`, `engine-pin-mismatch`, `live-session-refused`, `foreign-host`, `transport-unknown`, `handled-without-run`, `ack-without-start`, `interrupted`, `no-turn`, `unsealed-engine`, `run-overlap`) are bounded draft IDs, like the existing names CHAT-01 line 363 describes. Where CHAT-01 already has a name, it is reused: `generation` (`check.mjs:205`) and `controller` (`check.mjs:230`). The package README lists them. ### What ships Three increments: I1, I3 and I4. Each is its own candidate and review round, as with the queue-as-data A1/A2 split. An increment starts when everything in its "Needs first" column is met. Lead decision 30 cut I2 (Pi dialogs) and moved I1b (C-5 adoption) to CHAT-04. | Increment | Content | Needs first | |---|---|---| | I1 | Controller, transport, writer claim, Pi adapter on a fake Pi engine, incremental events, mediated terminal, interrupt, force stop and recovery. Fixtures in §2–§6 and §9, except the approval races H5–H8, which run in I4; the §10 checks. | Brief approved; author named | | I3 | The B1 evidence packet for Claude (§8) | I1 approved. Jason's go for any run that calls a model or reads Claude auth. | | I4 | Claude adapter, catalogue and history (§7, §8), and H5–H8 on Claude permission requests | I3 approved, so B1 passed; C-3 if it exists | **Build note (Rocko, R3).** The `no-turn` refusal lifts only the fence its own Interrupt set. It never reopens admission that a concurrent force stop, overlap signal or revocation closed. H10 covers those branches. **If B1 does not pass.** Sage records B1 as not passed when any of these happens: - Jason declines the go for the I3 recordings; - Filbert or Rocko refuses the B1 packet, and no path to a pass exists on the pinned version; - Jason says to close CHAT-03 without Claude. CHAT-03 then closes without I4. Claude keeps refusing `unsupported-harness`, and CHAT-06's both-harness gate stays blocked with CHAT-03 as its owner. The outcome is recorded, not a silent narrowing. #### 1. Controller and transport - One controller process per execution owns the engine's stdin. Pi runs in `--mode rpc` with no TUI, so a mediated engine has no native composer or terminal left to fence. That is how CHAT-03 meets plan lines 133–134: a losing native TUI can't stay writable because none exists. - Engine output is split on LF only (`rpc.md` lines 30–38). Node `readline` is not used, because it also splits on U+2028 and U+2029. The ledger refused a live line for that class of bug (lead decisions item 18). Fixture E2 covers it. - Clients connect over a Unix socket in a 0700 directory. Every connection starts as an observer. The browser does not reach the socket in CHAT-03; the WebUI side is CHAT-05. - **Actor.** There is one actor, `local-operator`, as in CHAT-02. The socket directory's mode is the only boundary, and every process with the same uid passes it. That includes agent seats, which run as the same user. So CHAT-03 cannot stop a same-uid agent from taking control of another conversation. See Limit 1. This blocks any live use until a local owner-channel design passes review (plan lines 181–185; B4). - **No broker queue.** CHAT-04 owns durable queues (plan line 216). In CHAT-03, a prompt sent while the engine is busy is refused with `busy`, and the text stays in the client. The controller never sends `streamingBehavior`, `steer` or `follow_up`, so mediated input never enters Pi's native queues (`rpc.md` lines 56–65 for `streamingBehavior`, 78–104 for `steer` and `follow_up`). Tradeoff: there are no queued follow-ups (Q6) until CHAT-04, but I1 has no native-queue ambiguity of its own making. - Admission and dispatch both recheck, under one dispatch lock: binding state, open admission, the controlling connection and its generation, the execution incarnation, the controller incarnation token and the text policy. CHAT-01 line 133 requires serialized dispatch, lines 153–160 the dispatch recheck of controller, fence and policy, and lines 181–183 the text policy. The single lock is this brief's way of meeting them. - **Pending dispatch slot.** There is one slot. Under the dispatch lock, the controller reserves it before writing a prompt. The engine counts as busy while the slot is held, and not only from `agent_start`. The slot is released only by reconciliation: an error response; or an ack, then the `agent_settled` of its run with no overlap signal (§3); or an ack, then a `get_state` showing no run with no `agent_start` or `agent_settled` in between (handled without a run, §3); or a refusal at the dispatch recheck before the write. On an overlap signal the slot stays held, the binding goes to `uncertain`, and only force stop moves it on. A second prompt sent while the slot is held refuses `busy`. This keeps "at most one unstarted item" true, which R3-1 depends on. - **Write outcomes.** The controller records three outcomes: - **written:** the whole line was accepted by the pipe; - **acknowledged:** Pi's `prompt` response arrived; - **unknown:** the write was partial, returned an error (EPIPE), the controller died before recording the result, or the line was written but no response came within the bound. A write accepted by the pipe is not native consumption. An unknown outcome always comes before `working`, since no ack has arrived. After an unknown outcome, the pipe counts as poisoned, because a partial line would join the next write. The controller never writes to it again. The receipt becomes `delivery-unknown` with reason `transport-unknown`, admission closes, and the binding goes to `uncertain`. Only force stop moves it on. Nothing is retried automatically and nothing is called unsent. - **Dedup key and incarnation token.** The key is the actor, conversation and client request ID (CHAT-01 line 144). The index lives only as long as one controller incarnation. Each controller start mints a random incarnation token, and `observe` returns it. Every command carries it. A command with an old token refuses `stale-incarnation`, whatever its request ID or text. So after a restart, an old retry is refused, even though the empty index can't recognize its ID. The client library marks requests pending under the old token as outcome-unknown, and shows a `stale-incarnation` refusal as "outcome unknown, check the transcript". It never resubmits them under the new token. Durable receipts are CHAT-04, which must restore the contract behavior (deviation V-1, lead decision 25). - **Text only.** Images and files are CHAT-04. #### 2. Writer-claim record (D1) CHAT-01 lines 336–343 allow at most one non-stopped binding per conversation or native session identity, and they defer the record. RUNTIME.md §3 item 3 adds one claim per (agent, project, workspace). The record: - **Keys and claim ID.** Each claim has two keys, and both must be held: the seat tuple (seat, project, workspace) and the native session identity (the Pi header ID or the Claude session UUID). A claim ID, random and minted at acquisition, ties them together. Every revision on either key names it. Keys are always taken seat first, then session, and released in the reverse order. A contender reads both keys before publishing. If the session key turns out held after it has published on the seat key, it follows that revision with `stopped` and a no-unit proof reference (it never spawned). A key counts as held when its highest revision is anything other than `stopped` with a proof reference. The pair is free only when both keys are free. - **Pair state.** Neither key's chain is authoritative alone. The pair's state is the more conservative of the two, in this order: `uncertain`, `stopping`, `active`, `reserved`, `stopped`. Two keys naming different claim IDs are both held, and acquisition refuses `unsafe-replacement` until each is resolved under its own claim ID. - **Atomic publication.** Each key is a directory under the claim root, and each revision is a numbered file in it. To publish, the controller writes the complete revision to a temporary file in the same directory, fsyncs it, then calls `link()` to the next revision name and fsyncs the directory. `link()` fails if the name exists, which makes publication exclusive, and a revision name only ever points at complete contents. A losing `link()` means another writer published first: the controller re-reads and re-decides. Revisions are never rewritten. - **Unreadable revision.** If the highest revision exists but won't parse, for example after disk damage, the key stays held as `uncertain`, and acquisition refuses `unsafe-replacement`. The controller never skips a damaged revision to reuse an older `stopped` one. - **Fields.** A revision stores: - the key, the claim ID and the binding ID; - the harness, conversation, and branch and leaf at launch; - the config pins: engine version, engine pin (the Pi package integrity of §3, or the Claude binary sha256 of §8) and launch argv digest; - the host identity (`/etc/machine-id`) and boot ID; - the owning controller: pid, process start time and incarnation token; - the intended scope unit name, derived from the claim ID, recorded before anything is spawned; - once known, the scope's systemd invocation ID and the engine pid and start time; - the controller generation; - the state: `reserved`, `active`, `stopping`, `uncertain` or `stopped`; - the proof reference that allowed `stopped`. There are three kinds: a cohortProof (§6), a boot proof (§6), and a no-unit observation, which is valid only when no spawn marker was published. - **Lifecycle.** Each step publishes the new state on the seat key, then on the session key: 1. reserve both keys, recording the intended unit name; 2. publish a spawn marker on both keys (still `reserved`), then start the engine inside that scope (§6), then record the invocation ID and engine identity; 3. publish `active` on both keys. Release publishes `stopping`, then `stopped` with the proof, on each key. - **Owner check.** A controller acts on a claim only if it owns the claim, or if the recorded owner is proven gone: a different boot on the same host, or no process with the recorded pid and start time. A paused or slow owner that still exists is never repaired, reclassified or orphaned by a second controller. The second controller refuses `already-active`. - **Restart and half-done pairs.** A controller that finds a claim whose owner is proven gone does not launch. It completes or classifies the pair under the same claim ID: - **Host differs from the recorded machine ID.** The pair is treated as `uncertain` and held, as the queue lock treats a foreign host (queue-as-data plan 8.4), and acquisition refuses `foreign-host`. Nothing is written, and a copied claim root never becomes `stopped`. - **Same host, different boot ID.** Every process of the recorded boot is gone. The supervisor issues a boot proof (§6), and both keys move to `stopped`. Tool calls with no recorded end become `uncertain` effects. - **Same boot, `reserved`, keys possibly half-done.** The intended unit name is looked up. - No spawn marker and no unit: nothing was started under that claim. The reservation moves to `stopped` with a no-unit observation. - Spawn marker and no unit: the engine may have run and exited, and a collected scope is an absent observation (§6), not proof that every process ended. The pair stays `uncertain`. Only a boot proof moves it on in CHAT-03. Any other rule is a reviewed change. A marker on either key counts, and completing a half-done pair copies the marker to the other key; it never drops it. - A unit exists: the pair is `uncertain`, and only force stop (§6) moves it on. - **Same boot, engine or unit alive.** This is an orphan: the stdin pipe is gone, so no controller can attach. The state is `uncertain`, and only a confirmed force stop of that recorded scope moves it on. - **Same boot, keys disagree** (for example one `stopped`, one `stopping`). The pair stays held. The restart continues the unfinished transition for that claim ID and never starts a new one. - **Location and live-session guard.** The claim root and every session path are constructor arguments. CHAT-03 has no production default. At construction and again at bind time, the controller resolves real paths. It refuses `live-session-refused` for any session file or claim root that is not inside the explicit fixture root it was given. It also refuses any path under the repository's `.pi/state/`, `~/.pi`, `~/.claude`, the configured data root, or a path named in any seat registration. This is what makes CHAT-03 fixture-only in code and not just by convention. The guard comes out only at cutover (CHAT-07), when the live location is reviewed as a data-map change. The claim excludes only controllers that use it: a `pi --session` run outside Mosaic isn't prevented, and the guard is what keeps CHAT-03 off real sessions. The idle session drift check moves to CHAT-07 (lead decision 30). - **What never releases a claim.** Disconnect never changes a claim. `agent_settled`, EOF, SIGTERM, idle and an abort acknowledgement never release one. CHAT-00 line 65 says settled must not release a writer claim. CHAT-01 line 325 says SIGTERM, EOF, an abort acknowledgement and idle don't prove death, so none of them can support a `stopped` revision either. - **No writes to sessions.** The controller never writes a session file and never uses `SessionManager.open` (CHAT-00 line 67). | # | Fixture | Expected | |---|---|---| | W1 | Two processes acquire the same pair at once | Exactly one claim. The other refuses `already-active`. | | W2 | Acquire while a claim is `active` or `reserved` | `already-active` | | W3 | Acquire while `stopping`, `uncertain`, or `stopped` without proof | `unsafe-replacement` | | W4 | Same session with a different seat tuple, and the reverse | Both refuse. A loser that already published on the seat key follows it with `stopped` (no-unit). It never spawns, and the winner's revisions are untouched. | | W5 | SIGKILL between every publication barrier of acquire, transition and release: after the temp write, after its fsync, after `link()`, after the directory fsync, and between the two keys | After restart: never two holders and never a lost claim. Every visible revision is complete. A half-done pair is completed or classified under its claim ID, as in "Restart". | | W6 | Controller killed mid-turn and restarted while the engine is alive | `uncertain`. No launch, and prompts refuse. | | W7 | Recorded boot ID differs, same machine ID | `stopped` with a boot proof. Open tool calls become `uncertain`. | | W8 | Resume after a proven stop with the same pins | New claim ID, generation +1, same conversation, branch and leaf | | W9 | Resume with a changed binary, argv digest, branch or leaf | Refused. The claim is unchanged. | | W11 | Controller writes to session files | None. The CHAT-02 F17 check runs over the fixture session directory. The fake engine's own appends are recorded separately and excluded. | | W12 | A live owner paused with SIGSTOP; a second controller starts | The second refuses `already-active`. The paused owner's revisions are unchanged after it resumes. | | W13 | Crash after the engine spawns but before `active` is published | Restart finds the reservation and the live unit: `uncertain`, force stop only. No second spawn. | | W14 | Crash after reservation, before the spawn marker | No unit and no marker: `stopped` with a no-unit observation. The pair is free. | | W20 | Crash after the spawn marker; the engine exits and the scope is collected before restart. Run twice: marker on both keys, and marker on the seat key only. | No unit, but a marker: `uncertain` in both runs. The session key gains the marker when the pair is completed. No launch until a boot proof. | | W15 | Crash between the two keys during release | The pair stays held. Restart finishes the release under the same claim ID. | | W16 | A highest revision that won't parse | Held as `uncertain`, and acquisition refuses. The older `stopped` revision is not reused. | | W17 | A claim root copied from a fixture "other host" (different machine ID) | `foreign-host`. Nothing is promoted. | | G1 | Session path or claim root under `.pi/state/`, `~/.claude`, the data root, or named in a registration | `live-session-refused` at construction | | G2 | A symlink inside the fixture root pointing at a live session file | `live-session-refused` at bind (real-path check) | | G3 | A fixture path that is swapped for a live path after construction | Refused at bind | #### 3. Pi adapter, incremental events and R3-1 **The Pi pin (lead decision 32).** Pi 0.85.1 is pinned by the `package-lock.json` integrity of `@earendil-works/pi-coding-agent`, `sha512-FGRN+OHbWaefBPGaTggAdLjrIHW+s2PzLyglz/5dfLzb9of7uuXMXYC0fJIeZTw+shS32o2cuQ9jF7YSDuL/oQ==`. `pi` runs `dist/bundle/cli.js` (the package's `bin`) and its chunks, not the `dist/core` and `dist/extensions` files listed below. At launch the controller checks that `package-lock.json` and npm's installed record (`node_modules/.package-lock.json`) both name 0.85.1 with that integrity, and otherwise refuses `engine-pin-mismatch`. That ties the install to the package through npm's record. It isn't a hash of the files on disk. The files below are the sources this brief cites for line numbers, hashed at `f2b9e622`, under `node_modules/@earendil-works/pi-coding-agent/`. Filbert's R3 review found that the bundle matches them on every point the brief relies on. - `docs/rpc.md` (`15fcd26bee72777b373fd5f2edd77091a01cadd4de95e48b08422ced0552a28d`) - `dist/modes/rpc/rpc-mode.js` (`e7e4724aa55c5aac73cf36793653b26736200e5c59d58373990fc31028f86477`) - `dist/modes/rpc/rpc-types.d.ts` (`e968e5be01dc7ad9615f938ae867ef136fa495f13dcf169942e9f781a299d9eb`) - `dist/core/agent-session.js` (`fb8a3981c20c8c0bbd42231b1c99a10335fb3858b659056b341954de9cfa467f`) - `dist/core/session-manager.js` (`ccace64949db25379a43971ecea750c1b7ec6344e1bc31b9d5fe596ac2f1c9f3`) - `dist/core/output-guard.js` (`e860db94650c57e07582c300983671737bf9e796682193b498f75e3dd72e9024`) - `dist/core/resource-loader.js` (`8e8a1bc1c5bc9e955f6a2314dd1db02071be56d48b1fea7b8b2cacb4fc9a0628`) - `dist/cli/args.js` (`bfb311d2c5d919fa4015d6aaa3c5a71a90b90011320e5e40f56b12e448c44dfc`) - `dist/main.js` (`f0b7e5a8419af8d149ffe367af2992c76ce70b73484c15492bd50787d4f4962a`) - `dist/extensions/index.js` (`f980647d447657237cb12b189cec903dc093c55f7a5942a57930c620b01420cc`) and `dist/extensions/llama/index.js` (`446b17f49d6197de5aaa6548f78da5934e4e5dfc83acdeed7119a2971ce5e8c1`) - `docs/usage.md` (`588896ba21944ff002d637444edc22698fd24959c59fe25f95010b9707b47d92`) - `node_modules/@earendil-works/pi-agent-core/dist/agent.js` (`d84351e451b9fef40fe2532c446aca90d26a4be9038b2d77d3d45dd6eab21d41`) and `dist/agent-loop.js` (`6732a1c65c09577d2ffcb716b48e4f4673e57e3e333f10ebfce5132d82e4d7a2`), version 0.85.1 - The goal extension, `extensions/goal/index.ts` (`5ccf78ce7e285ce290add9798b34f0e4e4b30fc8a154c006e34d2494e51a06ae`). `scripts/sync-dev-extensions.sh` copies it into `.pi/extensions/goal/`, which is untracked (`.pi/.gitignore` line 1), and the five Pi repository seats load that copy through `scripts/agent-host-dev.sh` line 137. The copy has the same hash. CHAT-03 cites goal as the example of a turn-starting extension and never loads it. The plan requires reading the Pi docs completely before Pi implementation (lines 163–164). The author does that before any I1 code. - **The fake engine.** A fixture process that speaks the pinned protocol: - prompt acceptance and refusal; - message and tool events; - `agent_end` arriving before `agent_settled`, and retries; - `clear_queue` and `abort`; - dialogs; - scripted pause points, so a test can land a race at an exact step; - the overlap model of N10: preflight awaits between the busy check and the run start, an acked loser whose throw is swallowed and which settles with no `agent_start`, several `agent_start` … `agent_end` pairs in one run, agent-level messages that `clear_queue` doesn't return, and `nextTurn` messages that survive a clear. Tests use these to simulate an unsealed extension without loading one. Its behavior is checked against the pinned files, not guessed. - **Real-binary smoke.** The pinned `pi` starts in `--mode rpc` in a scratch directory, with a scratch `HOME` and a scratch Pi agent directory, so the default `~/.pi` auth can't be found, and no model call. The setup checks that no auth file is reachable before starting the binary. It answers `get_state`, `get_commands`, and `clear_queue` and `abort` while idle. The recorded exchange shows that the fake's framing matches the binary. If the binary won't start without auth, the author records that, and the packet says plainly that I1 rests on the fake alone. - **Events.** Native events map onto CHAT-01 `event` records: message-start, text and thinking deltas, tool-start, tool-update and tool-end, message-end with `updateMode: replace`, and run-settled. - Every event carries the execution incarnation and a sequence number for that execution. - An unknown native event gets no client event. It is recorded in controller evidence and counted, and the terminal shows the count. It is never dropped silently and never passed through raw. - `agent_settled` maps to run-settled, never to cohort termination. - **Joining history and the stream.** CHAT-01 lines 105–112 require an atomic cut between a history page and the stream; otherwise streaming is advertised as unavailable. Pinned Pi emits `message_end` before it persists the entry (`agent-session.js` lines 386–398), so a page read just after the event can miss that message. Unless the author shows a cut from the pinned source, I1 advertises replay as unavailable. A client gets the page, then live events from the moment it subscribes, with a reconcile marker at the seam. Stream events carry no entry ID, so the seam can't be deduplicated by ID; the marker tells the client to reconcile by re-reading the page once the run settles. Fixture E4 covers both cases. **R3-1: dispatched input when a fence lands.** R1 assumed a Mosaic prompt could wait in Pi's queue. It can't. R2 replaced that with a second premise, that Pi runs a Mosaic prompt at once or not at all. Filbert's R2 review (`agents/filbert/work/chat-03-brief-review-r2-2026-09-26.md`, `d149e8cc`, F1) showed that premise is false too. R3 builds only on what the pinned source shows (`dist/core/agent-session.js`, hash above): - `isStreaming` is `_isAgentRunActive` (lines 616–617). `_runAgentPrompt` sets it at line 773. Its `finally` clears it and then emits `agent_settled`, running extension handlers before the RPC event (lines 780–784 and 347–351). - `prompt()` without `streamingBehavior` throws while `isStreaming` is true (lines 860–863). The controller never sends `streamingBehavior`, so a Mosaic prompt never enters Pi's steer or follow-up queues. - The check at line 860 isn't atomic with the run start at line 949. Between them, `prompt()` always awaits `emitBeforeAgentStart` (line 915). It also awaits `_checkCompaction` (line 895) once a turn exists, and `emitInput` (line 843) when an input handler exists. In that window another run can start: an extension's `sendCustomMessage` with `triggerTurn` (lines 1120–1121), or an extension prompt through `sendUserMessage` (lines 1161–1187) that passed its own line-860 check first. - The prompt that loses is still acked (`preflightResult(true)`, line 948). Its `agent.prompt` throws (`pi-agent-core` `dist/agent.js` line 228). The RPC handler swallows the throw because preflight succeeded (`rpc-mode.js` lines 314–317). Its `finally` then emits `agent_settled` and clears `isStreaming` while the other run goes on. Filbert confirmed this on the installed method with a stub receiver. No real engine was run. - One prompt's run can hold several `agent_start` … `agent_end` pairs before its single `agent_settled`, because retries and compaction continue it (`_handlePostAgentRun`, lines 787–810). In this brief, "a run" means one prompt's run, from its first `agent_start` to its `agent_settled`. - Events carry no prompt or run ID. The only native evidence tied to a Mosaic prompt is its own response: the ack or the error. So `agent_settled`, an idle `get_state`, and the first user `message_start` after an ack aren't tied to a Mosaic prompt whenever other code can start a turn. The goal extension that repository seats load does that from every `agent_settled`. Its handler calls `sendUserMessage` (goal lines 468–483 and 189), and Pi dispatches the call without awaiting it (lines 2020–2021). `clear_queue` doesn't see all external input either (Filbert F2). `clearQueue()` returns only `_steeringMessages` and `_followUpMessages` (lines 1195–1203). An extension's `sendMessage` while streaming goes straight to `agent.steer` or `agent.followUp` (lines 1112–1118). `clearAllQueues` drops it without returning it, and the pending count in `get_state` (line 1206) misses it too. `nextTurn` messages (line 1110) survive clear and abort, and attach to the next prompt (lines 910–913). **The rule R3 follows (Sage).** A receipt settles only on positive evidence tied to its own prompt. Idle, settled or empty settle nothing on their own. Where pinned Pi can't tell the cases apart, the controller refuses or reports unknown. **Sealed engine, a binding precondition.** Pinned Pi gives no evidence that ties a run to a prompt. So CHAT-03 removes every other source of turns and input, and doesn't guess. A binding needs a sealed engine: - The controller builds the launch argv. Pi starts with `--no-extensions`, `--no-prompt-templates` and `--no-themes`, and with no `--extension` argument (lead decision 31; `dist/cli/args.js` lines 135–140, 167 and 170; `docs/usage.md` lines 224 and 233–236). With `--no-extensions`, Pi loads only the paths given on the command line and ignores settings and packages (`dist/core/resource-loader.js` lines 316–318), so no explicit extension loads. Skills stay allowed: they expand text and start no turn (§4). - Pi also always loads its built-in extensions, whatever the flags (`dist/main.js` line 439). Pi 0.85.1 has one, `llama.cpp` (`dist/extensions/index.js`). It registers a provider and a `/llama` command, and no event handler or input call (`llama/index.js` lines 37 and 163). It is part of the pinned package (the Pi pin above), so it is pinned with it. - Any `--extension` argument, or an argv without the three `--no-*` flags, refuses the binding with `unsealed-engine`. Explicit extensions aren't pinned or reviewed for binding: Rocko's R3 review showed that an extension can import code outside any hashed tree. They come back with goal in CHAT-06 (lead decision 31). The goal extension starts turns, so a session that loads it can't bind in CHAT-03 (Limit 11). Under lead decision 30, a seat bound to the Console runs without goal, and goal continuation is a CHAT-06 item. Under the seal, the Mosaic prompt in the slot is the only thing that can start a run, so a run that follows its ack is its run. That exclusion is the evidence basis for attributing by order, and the brief names it as the basis. The seal rests on the launch argv and the pinned Pi build, so the controller also watches for signs that it failed. **Overlap signals.** Each of these means the seal failed, or Pi behaved outside the pinned model: - O1: an `agent_start` while no slot is held, before the slot's ack, or after the slot's run has settled; - O2: an `agent_settled` while no slot is held, while the slot is held but before its ack, or a second `agent_settled` for one ack; - O3: an `agent_settled` while an `agent_start` has no matching `agent_end`; - O4: an `agent_settled` after an ack, with no `agent_start` and no failure message between them; - O5: a non-empty `clear_queue`, a non-zero pending count, or a `queue_update` the controller didn't cause. Mosaic never queues, so under the seal these are always empty. Pi's own clear emits a `queue_update` before the clear's response (`agent-session.js` line 1201, `rpc-mode.js` line 334). A `queue_update` read after the controller writes `clear_queue` and before that response, with empty `steering` and `followUp`, is the clear's own, and O5 counts it as caused (N25); - O6: a second user `message_start` in one run. On any overlap signal, admission closes, the binding goes to `uncertain` with reason `run-overlap`, and force stop is the way on. The slot's item becomes `delivery-unknown` with reason `run-overlap` if it hasn't reached `working`. If it has, it stays `working`, shown as "outcome unknown". A stop in progress gets the Unknown outcome. Because a receipt never moves back from `working`, a seal failure can misattribute a run to the slot: another run's user `message_start` arrives before any signal, the item goes to `working`, and the signal that follows can only mark it "outcome unknown" (N8). Some violations give no signal at all: custom messages queued straight into the agent, and `nextTurn` messages (Limit 7). A spurious settle delayed until after the item's receipt has settled signals only once the receipt is final (N21, Limit 11). The rule: 1. **Order.** Interrupt closes admission, sends `clear_queue`, then `abort`. `abort` would run anything still queued (`rpc.md` line 158). 2. **Clear timeout or failure.** `abort` is not sent, because it would run whatever is queued. The turnProof records `nativeQueue: unknown`, the stop outcome is Unknown, the stop is `uncertain`, and admission stays closed. Force stop stays available. It needs no cooperation from the engine (§6). 3. **What the clear returned.** An empty clear gives `nativeQueue: cleared`, which here means that Pi's steer and follow-up queues returned nothing. It doesn't prove that no external input was removed, because agent-level custom messages are dropped without being returned (F2). The seal is what keeps those out, and the evidence names the seal as that basis, not the clear. A non-empty clear is overlap signal O5. The stop stays `uncertain`, `reconciled` isn't published, admission stays closed (CHAT-01 lines 311–312), and force stop is the way on. The removed items go to controller evidence as digests and byte counts. They are never attributed to a request and never resent. 4. **Receipt settlement.** The Mosaic item's receipt settles from evidence tied to its own prompt, never from `clear_queue` and never from the stop. The prompt's own response is tied to it by ID. A run is tied to it only through the seal, and only while no overlap signal has been seen. A late ack that arrives after the fence goes through the same table as an earlier one. An ack alone never means a run. | Evidence for the item | Receipt | |---|---| | Error response to the prompt, before or after the fence | `failed`, native error kept | | Refused at the dispatch recheck before any write (fence set) | `dispatch-refused`; no engine bytes | | Ack; then `get_state` says not streaming, with no `agent_start` or `agent_settled` between the ack and that reply | `delivery-unknown`, reason `handled-without-run` | | Ack; then `agent_start` and a user `message_start`, with no overlap signal | `working`. At the run's `agent_settled`, if the observation from the ack to the settle is complete and ordered and holds no overlap signal, the `stopReason` of the last assistant `message_end` decides: `stop`, `length` or `toolUse` gives `finished`; `error` gives `failed`; `aborted` gives `failed` with reason `interrupted`, and evidence links the stop ID. Otherwise (no final assistant `message_end`, another `stopReason` such as `pending` or `deferred`, or a gap), the receipt stays `working`, shown as "outcome unknown", and the binding goes to `uncertain`. | | Ack; then a run that ends in a failure message before any user `message_start`; then `agent_settled` | `delivery-unknown`, reason `ack-without-start` | | Any overlap signal before `working` | `delivery-unknown`, reason `run-overlap` | | Write outcome unknown, a missing or unparseable line, or no ack, `agent_start` or `get_state` reply within the bound | Before `working`: `delivery-unknown`, reason `transport-unknown`. Once `working`: it stays `working`, shown as "outcome unknown", because a receipt never moves back from `working`. Either way the pipe is poisoned (§1) and the binding goes to `uncertain`. | | Ack; then `agent_settled` after any sequence no row above matches | `delivery-unknown`, reason `ack-without-start` | The transport row and the overlap row take precedence over the rest. Otherwise the rows are exclusive. A line lost without a trace can't be detected, and the table doesn't claim to detect it. An unparseable line is detected. A preflight still in flight when the fence lands is awaited, bounded, and then classified by this table. If the classification shows a run that is still active, the controller repeats clear, then abort (rule 1), at most three times. If the run is still going after the third, `nativeQueue: unknown`, the stop is `uncertain`, and force stop is the way on. **Why `ack-without-start` isn't `failed`.** A run that fails before its user message emits a failure assistant message from Pi's run-failure handler (`pi-agent-core` `dist/agent.js` lines 349–364). Pi persists a run's user message only in the `message_end` handler (`agent-session.js` lines 386–398), so the session gains no user entry. That ordering is inside Pi. Output reaches the controller through one ordered promise chain (`rpc-mode.js` lines 28–29, `output-guard.js` line 71), so a received `agent_settled` means that every earlier event from the process was written first. It doesn't mean that the settle belongs to this prompt (F1). It also doesn't mean that input or before-agent-start extensions had no effect. The controller has no positive evidence that the item had no effect, so the receipt is `delivery-unknown`, and a `failed` that invites a resend is never shown. This answers Rocko's R2 note 2 and Filbert's F1 row 4. A `delivery-unknown` item is never relabelled unsent and never resent. The client shows it as "outcome unknown". The actor may copy the text into a new request, and the client doesn't present that as a safe resend. 5. **The stop outcome.** The Interrupt stop is classified separately from the receipt. A run is active at a moment if its first `agent_start` was read before that moment and its `agent_settled` wasn't. A run is in scope if it was active when the fence was set, or if its first `agent_start` was read before the last `abort` was written. The Interrupted, Completed first and Failed on its own rows also need all of this: exactly one run in scope, its `agent_settled` read, and an observation from the fence to that settle that is complete and ordered and holds no overlap signal. Unknown's conditions are checked first. Otherwise the first row that matches wins. | Stop outcome | Evidence | CHAT-01 reconciliation | |---|---|---| | Interrupted | The run was active when an `abort` was written, and its last assistant `message_end` has `stopReason: aborted`. What caused the abort isn't claimed: an extension's abort in the same window still interrupted the turn. | `turnState: interrupted`. `reconciled` may follow when the other rule 6 conditions hold (`check.mjs` line 315). | | Completed first | The run's last `stopReason` is `stop`, `length` or `toolUse`: it finished before the abort took effect | None honest. The receipt stays `finished`, and its effects stay the run's own. The completion is never relabelled as an interruption. | | Failed on its own | The run's last `stopReason` is `error` | None honest | | No run | No run is in scope. The slot held a written item that settled `failed` or `handled-without-run`, and the observation from the fence to the last `abort` has no gap and no overlap signal. | None honest | | Unknown | Anything else: no `abort` written (rule 2), any overlap signal, more than one run in scope, no final assistant message, any other `stopReason`, a transport gap, a missing settle, or a run whose first `agent_start` was read only after the last abort | None | CHAT-01 lets `reconcile-interrupt` pass only with `turnState: interrupted` (`check.mjs` line 315). `input-reconciled` belongs to revocation, and CHAT-03 doesn't borrow it. So until C-5 gives the other outcomes an honest value, every outcome except Interrupted leaves the stop `uncertain`: no `reconciled`, prompts refuse (CHAT-01 lines 311–312), and force stop is the way on. The stop advances to `uncertain` (`check.mjs` line 332), and no turnProof is used for reconciliation. Controller evidence records the outcome, so the client can say "the turn had already finished" instead of "interrupted". The author confirms from `pi-agent-core` that every run ended by an abort emits a final assistant `message_end` with `stopReason: aborted` (`dist/agent.js` line 357, `dist/agent-loop.js` line 124). Any path that doesn't is classified Unknown. **Nothing to interrupt.** Interrupt sets the fence, takes the dispatch lock, and then checks the slot. The dispatch recheck reads the fence under the same lock immediately before the write (§1), and a write already under way counts as written. A slot that was reserved but not written therefore refuses at the recheck: the item becomes `dispatch-refused`, the slot is released, and no bytes reach the engine (H9). If, after that, no written item holds the slot and no run is active, Interrupt refuses `no-turn`. The fence is lifted, admission reopens, and no stop record is created. The refusal has one recorded effect, the `dispatch-refused` item if there was one, and names it. This narrowing happens on the controller side, like `busy`. CHAT-01's checker creates a stop for every interrupt it admits (`check.mjs` line 244). The narrowing leaves only the race with a turn that finishes during the clear and abort exchange. If the controller refuses just before it reads a new run's `agent_start`, the run keeps going, and the actor can interrupt again or use force stop. 6. **Reopening.** `reconciled` is published and admission reopens only when all of these hold: - the stop outcome is Interrupted (rule 5). Once C-5 is adopted, Completed first, Failed on its own and No run qualify too; - if the slot held a Mosaic item, its receipt is settled by rule 4 as `finished`, `failed` or `delivery-unknown` (reason `handled-without-run` or `ack-without-start`). A `transport-unknown` or `run-overlap` receipt, or a poisoned pipe, keeps the stop `uncertain`; - the in-scope run's `agent_settled` was read (No run has none), and a `get_state` sent after the last `abort` shows not streaming; - a `clear_queue` sent after that returns empty; - every clear in the sequence was empty (rule 3), and no overlap signal was seen. Under the seal, a non-empty post-settle clear is overlap signal O5: the stop stays `uncertain` and admission stays closed. 7. **What the proof covers.** The proof covers Pi's steer and follow-up queues as of the last clear, and the turnProof says so with its `observedAt`. Under the seal no extension queues input. If the seal fails, three things stay invisible: input queued after the last clear, custom messages queued straight into the agent, and `nextTurn` messages that attach to the next Mosaic prompt. The first can show up later as an overlap signal; the other two can't (Limit 7). 8. **Late native events.** A user `message_start` that arrives after the fence is recorded as an event. CHAT-01 says only that a receipt is monotonically revised (README line 36) and has no transition table, so the brief fixes the order: `admitted` < `dispatched` < `acknowledged` < `working` < `finished` or `failed`. `dispatch-refused` is reachable only before `dispatched`, and `delivery-unknown` only before `working`. A receipt never moves to a state claiming the item wasn't sent. Outside a fence, the slot's item becomes `working` at the first user `message_start` after its ack and `agent_start`, if no overlap signal has been seen. The attribution rests on the seal, not on order alone. The adapter doesn't compare text, because skills and input handlers can change it. The README says that `working` is attributed through the seal. The fake engine models the overlap (N10). Fixtures that break the seal don't load a real extension, since the binding would refuse it. The fake simulates the extension's effect instead. | # | Fixture | Expected | |---|---|---| | N1 | During an active Mosaic run, Interrupt; after the fence is set and before the clear, the fake simulates an unsealed extension that queues a follow-up | `clear_queue` goes before `abort`, and the follow-up doesn't run. The `queue_update` and the non-empty clear are O5: `run-overlap`, stop outcome Unknown, stop `uncertain`, prompts refuse, and evidence holds the item's digest. The Mosaic receipt stays `working`, shown as outcome unknown. Queued before the fence, the `queue_update` is O5 at once: the binding goes `uncertain`, and CHAT-01 refuses the Interrupt as `fenced` (`check.mjs` line 243). | | N2 | N1 with `abort` sent first (mutant) | The fake runs the external item, so the test fails. This is the ordering guard. | | N3 | Fence while the Mosaic prompt is in preflight; preflight then errors, and no run exists | Receipt `failed` with the native error. Stop outcome No run: no `turnState: interrupted`, stop `uncertain`, prompts refuse until C-5, force stop available. | | N4 | Fence while the Mosaic prompt is in preflight; the ack arrives after the first `abort`, and a run starts | Classified by rule 4 as a run. Clear, then abort again. The run ends `aborted`: receipt `failed`, reason `interrupted`; stop outcome Interrupted; `reconciled` after a post-settle empty clear. | | N5 | Mosaic prompt acked but handled by an input handler the fake simulates, with no run | `get_state` shows no run, and no `agent_start` or `agent_settled` arrived. Receipt `delivery-unknown`, reason `handled-without-run`. Nothing is resent. | | N6 | During an active Mosaic run, Interrupt; the fake simulates an unsealed extension that queues between `clear_queue` and `abort` | The `queue_update` is O5. `abort` continues the queued item inside the same run (`agent-session.js` lines 787–810), so its user `message_start` is O6. `run-overlap`, stop outcome Unknown, stop `uncertain`, admission closed, force stop is the way on. | | N7 | `clear_queue` times out during an active run | No `abort` sent. Stop outcome Unknown. `nativeQueue: unknown`, stop `uncertain`, admission closed. A confirmed force stop still ends the cohort (K1). | | N8 | Filbert's order one: during Mosaic preflight (paused at the line-915 await), the fake simulates an extension prompt that starts a run first. The wire shows the Mosaic ack, the other run's `agent_start` and user `message_start`, then the losing Mosaic prompt's `agent_settled` | The other run's user `message_start` arrives before any signal, so the receipt goes to `working`. The losing settle is O3: the receipt stays `working`, shown as outcome unknown, never `finished` or `failed`. Binding `uncertain`, admission closed, nothing resent. The fixture pins the misattribution window (Limit 11). | | N9 | The item's run started before the fence and ends `aborted` | Receipt `failed`, reason `interrupted`, not `delivery-unknown`. Stop outcome Interrupted (`turnState: interrupted`). | | N10 | Fake conformance | The fake throws on a prompt while streaming and acks before running. It emits `agent_settled` from a `finally`, and it can hold several `agent_start` … `agent_end` pairs in one run. It pauses at the preflight awaits (lines 843, 895, 915) between the line-860 check and the run start. A colliding prompt is acked, its throw is swallowed, it settles with no `agent_start`, and it leaves `isStreaming` false while the other run goes on. `clearQueue` emits an empty `queue_update` before its response, doesn't return agent-level custom messages, and leaves `nextTurn` messages in place. A fake that queues a Mosaic prompt, or can't produce the overlap, fails the test. | | N11 | Ack; then the run ends in a failure message before any user `message_start`; `agent_settled` arrives; the stream is complete and ordered | Receipt `delivery-unknown`, reason `ack-without-start`, never `failed`. The session file gains no user entry. The client shows outcome unknown and no resend offer. The same schedule with the failure line unparseable gives `delivery-unknown` / `transport-unknown`. | | N12 | During an active run, the fake simulates an extension that queues after the final empty clear | The stop records the clear's `observedAt`. The queued item's run starts with no slot held: O1, `run-overlap`, binding `uncertain`, admission closed. It is not part of the stopped turn's proof. | | N13 | The fake emits an `agent_start` with no slot held (a simulated turn-starting extension) | O1: binding `uncertain`, `run-overlap`, admission closed. A prompt sent afterwards refuses with zero engine bytes. | | N14 | The run completes normally (`stopReason: stop`) while `clear_queue` is in flight; `abort` then reaches an idle engine; all clears empty | Receipt `finished`, effects kept. Stop outcome Completed first. No `turnState: interrupted`, no `reconciled`, stop `uncertain`, prompts refuse until C-5, force stop available. | | N15 | Fence while the Mosaic prompt is in preflight; after the first clear and abort, an input handler the fake simulates handles it and Pi acks with no run | Receipt `delivery-unknown`, reason `handled-without-run`. No second clear or abort is sent. Stop outcome No run: `uncertain` until C-5. | | N16 | Interrupt with no slot held and no visible run | Refused `no-turn`. No stop record, no engine bytes. Admission is open afterwards, and the next prompt is admitted. | | N17 | The run fails on its own (`stopReason: error`) during the clear and abort exchange | Receipt `failed`. Stop outcome Failed on its own: `uncertain` until C-5. | | N18 | The run settles with no final assistant `message_end`, or a line is lost during the exchange | Receipt `working`, shown as outcome unknown. A lost line after `working` never moves it back to `delivery-unknown`. A lost line before `working` gives `delivery-unknown` / `transport-unknown`. Stop outcome Unknown: `uncertain`. | | N19 | Filbert's order two: the Mosaic prompt wins, and a simulated extension prompt that lost emits its `agent_settled` before the Mosaic run's user `message_start` | O3 (or O4 if it lands before the Mosaic `agent_start`). Receipt `delivery-unknown`, reason `run-overlap`, never `failed ack-without-start`. Binding `uncertain`. | | N20 | A simulated extension timer calls `sendCustomMessage` with `triggerTurn` during Mosaic preflight | As N8. `triggerTurn` goes to `_runAgentPrompt` with no preflight (`agent-session.js` lines 1120–1121), so the extension's run always starts first. Never `finished` or `failed`. The same call while the Mosaic run streams is queued straight into the agent (lines 1112–1118) and gives no signal (Limit 7, N22). | | N21 | The losing prompt's `agent_settled` is delayed until after the Mosaic receipt settled `finished` | O2 on arrival: binding `uncertain`, `run-overlap`, admission closed. The receipt stays `finished`, since receipts are monotonic, and evidence records the overlap against it (Limit 11). | | N22 | During an active run, the fake simulates an extension `sendMessage` queued straight into the agent; Interrupt | The clear returns empty and the item is gone. Evidence records agent-level queues as unobservable, with the seal as the basis, and never as "nothing removed". This fixture pins the Limit 7 gap: no signal fires. | | N23 | The fake simulates a `nextTurn` message queued before an Interrupt | It survives clear and abort and attaches to the next Mosaic prompt. No signal fires. The fixture pins the Limit 7 gap. | | N24 | Launch with any `--extension` argument (a local path, an npm or git source), or an argv missing `--no-extensions`, `--no-prompt-templates` or `--no-themes` | `unsealed-engine` at bind, for each case. No engine is started. | | N25 | An ordinary Interrupt of an active Mosaic run: the fake emits Pi's empty `queue_update` before each `clear_queue` response, and the run ends `aborted` | No overlap signal. Stop outcome Interrupted, `reconciled` after the post-settle empty clear, admission reopens. A non-empty `queue_update` in the same window is O5. | #### 4. The slash path (carry-forward 3) There are two separate hazards: text left in a composer gets concatenated with a new message, and Pi interprets certain prefixes. - **Concatenation.** `send-message.sh` pastes onto whatever the composer holds. The mediated path has no engine-side composer. Each prompt is one JSONL record, and its `message` field holds the whole text. Nothing left over can prefix it. The mediated terminal's own composer is a local buffer. It is empty at start, cleared after each submit and on control transfer, and it submits only while its connection is the controller. - **Interpretation.** RPC still acts on a leading `/`. Extension commands run immediately, even while streaming, and skill and template commands expand (`rpc.md` lines 67–69). - Repository seats load one extension command, `/goal` (`extensions/goal/index.ts:230`, loaded through the `.pi/extensions` path at `scripts/agent-host-dev.sh:137`). Prompt templates are off (line 136). Skills are live: `--no-skills` stops discovery only, and the ten explicit `--skill` paths still load (`agent-host-dev.sh` lines 77–83 and 138; Pi `docs/skills.md` line 42). So `/skill:` expands on repository seats. - A CHAT-03 binding can't load goal (§3, sealed engine), so `/goal` has no command behind it there. The only built-in command, `/llama`, acts only in TUI mode. Admission doesn't rely on either fact. - Admission refuses `text-policy` when the first non-whitespace character is `/` (CHAT-01 lines 181–183). Dispatch checks again. The rule holds whatever extensions are loaded, and S1 tests it. - CHAT-01 lines 368–370 leave `!`, `@` and slashes on later lines open under B3. I1 settles them from the pinned source. The author lists every prefix that Pi 0.85.1 RPC `prompt` interprets, and the list becomes a fixture. Admission refuses each listed prefix, and any prefix the author can't classify. - **Board and agent-send for a mediated seat.** A mediated engine has no tmux pane. The board already refuses a reply to a registration without a tmux session before the tool runs (`packages/control-board/src/serve.mjs:84`; test `serve.test.mjs:810`). CHAT-03 relies on that and changes neither the board nor `tools/tmux`. No real seat registers as mediated in CHAT-03, so the refusal first applies to a real seat at cutover. - **Seats still on tmux keep the hazard.** The DEFERRED entry stays open. It closes when a seat moves to the mediated path (CHAT-07), or when a CHAT-03I charter covers `tools/tmux` (CHAT-01 lines 232–238). Neither happens in CHAT-03. | # | Fixture | Expected | |---|---|---| | S1 | `/goal x`, and the same with leading spaces or a tab | `text-policy` at admission. Zero bytes reach the fake engine. | | S2 | Each prefix on the author's list, including `/skill:ms-unslop` | Refused. The test reads the same list the code uses. | | S3 | `/goal` on the second line | The fixture pins the result from the pinned source: admitted if Pi doesn't interpret later lines, refused otherwise | | S4 | The terminal composer holds a `/` from an abandoned edit, then control transfers and returns | Composer cleared. The next submit sends only the new text, and the fake engine records the exact bytes. | | S5 | An observer terminal gets a paste and then Enter, as `send-message.sh` does | Not admitted: `controller`. Zero engine bytes. | | S6 | A mediated-shaped registration (no tmux) passed to `replyToRow` with a recording `exec` | 409, "no tmux session". `exec` is never called. | | S7 | A prompt containing ESC, bracketed-paste markers or U+2028 | Sent as one JSON string. The fake engine receives the exact text, and no record splits. | #### 5. Control races (Rocko) Each race uses the fake engine's pause points to land at an exact step. Every race asserts the engine bytes, the receipts and the events. H5–H8 run in I4 on Claude permission requests. Pi dialogs are cut (lead decision 30). | # | Race | Expected | |---|---|---| | H1 | Two takeovers with the same expected generation | One wins, and the generation goes up by 1. The other refuses `generation`. | | H2 | The old controller's prompt arrives after a takeover commits | Refused. Zero engine bytes. | | H3 | A takeover commits while a prompt holds the dispatch lock | The prompt either completed its write before the commit, and the receipt keeps the old actor, or it is refused. If the write outcome is unknown, the §1 rule applies: `delivery-unknown`, pipe poisoned, `uncertain`. Never dispatched under both controllers. | | H4 | Self-takeover | Refused (CHAT-01 line 269) | | H5 | The old controller answers an approval after a takeover | Refused. The decision is answered once, through the new projection. | | H6 | Duplicate approval answers with the same content | One native response | | H7 | Conflicting approval answers with the same request ID | The second is refused | | H8 | An approval answer after a native timeout or cancel | Refused. The state comes from native evidence. | | H9 | Interrupt racing a prompt's dispatch | Fence set before the write: the item is `dispatch-refused` with no engine bytes, and with no run active the Interrupt then refuses `no-turn` (no stop record). Fence set after the write began: §3 rules 4–6. | | H10 | Interrupt and force stop at the same time; again with an overlap signal or a revocation closing admission before the Interrupt finds no slot and refuses `no-turn` | One stop chain. Force stop supersedes (CHAT-01 lines 319–322). The `no-turn` cleanup lifts only its own fence: admission stays closed under the surviving reason. | | H11 | The controller disconnects mid-turn | Work continues and the claim is unchanged. Control stays with the disconnected connection until an observer takes over explicitly. Nothing happens automatically. | | H12 | Exact retry of a prompt after reconnecting to the same controller incarnation | The same receipt. No second dispatch. | | H13 | Retry with the same request ID and different text | Refused | | H14 | Late stdout from the old engine after a replacement | Dropped by execution incarnation and counted in the evidence. Never rendered. | | H15 | A revoked connection sends a command | Refused, with the revocation fence of CHAT-01 lines 271–280 | | H16 | A second controller process for the same session | Refused as in W1 and W2. The first controller is untouched. | | H17 | A confirmation reused, answered from another connection, or answered after the stop context changed | Refused. Confirmations are single-use (CHAT-01 lines 264–268; CHAT-01C lines 92–102). | | H18 | Two prompts sent before any native output from the first | The second refuses `busy`. One engine write. | | H19 | A large prompt line under backpressure; the pipe closes (EPIPE) mid-line, or the controller dies mid-write | `delivery-unknown`, reason `transport-unknown`. Pipe poisoned, no later write, `uncertain`. No retry. | | H20 | The line is written completely, but the native acknowledgement is lost when the controller dies | After restart: the claim is an orphan (W6). The request's outcome is unknown, and nothing is resent. | | H21 | Crash after native dispatch, before the client gets its receipt; restart; the client reconnects and retries the exact request with the old token | `stale-incarnation`. No second engine write. | | H22 | After H21 and a valid recovery, the client sends a new request with the new token | Admitted normally | | H23 | Requests pending when the controller restarts; the client library reconnects and gets the new token | The library resends none of them. The fake engine records zero bytes for them. Each shows "outcome unknown, check the transcript". | #### 6. Stop, cohort proof and recovery (Rocko) CHAT-01 lines 330–333: a self-posted hash is not trust, and a real producer/verifier and complete cohort containment remain B3/B4. CHAT-03 builds the producer and its observation procedure. The verifier stays the fixture's trusted digest registry, as in CHAT-01. So no CHAT-03 proof is live authority. Where the procedure below can't be carried out, the stop ends at `uncertain`, not `stopped`. - **Producer.** A supervisor shim in `packages/conversation/src/` starts with `systemd-run --user --scope`, delegated, under the intended unit name recorded in the claim (§2). Inside the scope it moves itself into a `supervisor` child cgroup and execs the engine in an `engine` child cgroup. So the engine is contained from its first instruction, with no window before it could fork. The cohort is the `engine` cgroup. The shim is not a member. It holds the scope open, so systemd can't garbage-collect the cgroup before emptiness is read. - **Identity and epoch.** The cohort reference is the host (`/etc/machine-id`), the boot ID, the unit name and the scope's systemd invocation ID. The invocation ID is the membership epoch. A unit with the right name but a different invocation ID is a different cohort. The supervisor sends it no signals and reports evidence unavailable. - **Containment against migration.** A same-uid process can move a pid between cgroups in the user's delegated tree, so a member could leave the scope and survive a "complete" kill. The candidate defence: run the engine in a cgroup namespace rooted at `engine`. This host mounts cgroup2 with `nsdelegate` and allows unprivileged user namespaces, and with `nsdelegate` a namespaced process can't migrate pids outside its namespace root. The author must show this with K13. If K13 can't be made to refuse the escape, real cohorts never reach `stopped` in CHAT-03: K1 then expects `uncertain`, and the packet says so. A same-uid process outside the cohort that moves members out is Limit 4 and B2, not something CHAT-03 claims to stop. - **Unavailable is not empty.** Emptiness is a readable `engine/cgroup.events` showing `populated 0`, for a scope whose invocation ID matches, read by the live shim. A missing or unreadable path, a gone shim, or an invocation-ID mismatch is evidence unavailable, and the stop stays `uncertain`. The shim stays a member of the scope in its own leaf, so systemd doesn't collect the scope while the shim reads `engine`. A collected scope or an absent `engine` cgroup is an absent observation, never an empty one. - **Proof.** `stopped` needs a cohortProof with every CHAT-01 field (schema `cohortProof`): stop, authority, conversation, execution, cohort reference and membership epoch, `membershipComplete`, members with boot, pid and start time and a death time each, observation time, and verification digest. Membership is complete as of the freeze in the kill phase below, when no member can fork. A process that exited before the freeze isn't listed; its effects belong to the effect report, not the death proof. `membershipComplete: true` is written only when the freeze, enumeration and `populated 0` all succeeded. SIGTERM, EOF, an abort acknowledgement, `agent_settled` and idle never promote a stop to `stopped` (line 325). - **Effects.** A tool-start with no tool-end at stop time is an `uncertain` effect. Killing never counts as rollback (line 334). Network and remote effects stay `uncertain`. - **Interrupt (Q20).** As in §3 rules 1–8: close admission, `clear_queue`, `abort`, settle the item's receipt (rule 4), classify the stop outcome (rule 5), a post-settle empty clear, the turnProof, `reconciled`, and reopen admission under current control (lines 307–313). Only the Interrupted outcome reconciles until C-5. If the turn completed or failed first, or no run was active, or any overlap signal was seen (a non-empty clear included), or the outcome is Unknown, or the pipe is poisoned, the stop stays `uncertain`, `reconciled` isn't published, and prompts refuse, as CHAT-01 lines 311–312 require before that transition. Force stop is the way on. With nothing running, Interrupt refuses `no-turn`. If it hangs or the clear times out, force stop stays available. - **Force stop (Q8).** Needs an answered confirmation. The stop ID, the answered confirmation and the current phase go into the `stopping` revision before any signal. Then: 1. check the scope's invocation ID against the claim; 2. **TERM phase:** SIGTERM to every member of `engine`, then a bounded grace period; 3. **kill phase:** freeze `engine` (`cgroup.freeze`, waiting for `frozen 1`), enumerate every member with pid and start time, write `cgroup.kill`, and wait for `populated 0`; 4. build the proof. It acts only on the selected execution's `engine` cgroup, and other executions survive. If a kernel feature the author relies on (`cgroup.freeze`, `cgroup.kill`, delegation) isn't available, the stop ends at `uncertain`. - **Controller death during force stop.** A restart (owner proven gone, §2) never assumes TERM or kill happened. It checks the invocation ID. If that matches, it re-runs the escalation from the TERM phase for the same stop record and the same scope. That continues the recorded stop; it doesn't reuse the confirmation for a new one. If identity can't be checked, it sends no signals, and the stop is `uncertain`. - **Boot proof.** For a claim whose recorded boot differs on the same host, the supervisor issues a cohortProof whose evidence is the boot change. It goes through the same verifier. A different machine ID is `foreign-host`, never a boot proof. - **Recover (Q15).** Recover reports eligibility only. It needs: - a stopped proof; - effects reconciled or explicitly uncertain; - an exact confirmation; - the same pins (lines 336–343). Recover publishes a new `reserved` claim, under a new claim ID, on both keys. The eligibility record names that claim, the stop, the pins, the leaf and the controller incarnation token, and it is single-use. So no one can take the pair between eligibility and launch. Launching is a separate library call made by a launcher. It revalidates the record against the claim and the current leaf, consumes it once, and spawns. No client command launches, and a browser or controller crash can't start an engine. - **Pi resume.** Resume reuses the same session file and leaf. The author confirms from the pinned docs how Pi selects the leaf. Before admission opens, the adapter checks `get_state` and `get_tree` and refuses if the engine loaded any other leaf (RUNTIME.md §3 item 4). A wrongly loaded engine stays under the claim and its scope until a force stop proves it stopped. | # | Fixture | Expected | |---|---|---| | K1 | Force stop an engine whose tool child calls `setsid` | The child is killed. `stopped` with a proof accepted by the fixture verifier if K13 refuses the escape; otherwise `uncertain`, recorded. | | K2 | K1 on the process-group fallback (no scope) | `uncertain`, never `stopped` | | K3 | SIGTERM acknowledged while a member is still alive | Stays `stopping` until the kill phase. Never `stopped` from TERM alone. | | K4 | Two fake engines; force stop one | The other survives, shown by independent observation: its own cgroup is populated and it still answers `get_state` | | K5 | Stop during a tool call | The effect is `uncertain` and shown | | K6 | Recover without proof, without confirmation, or with changed pins | Refused | | K7 | Recover after proof, then launch | New claim ID and incarnation, generation +1, same leaf. The cancelled prompt is not replayed. | | K8 | The engine loads a different leaf on resume | Refused before admission opens. The engine stays claimed and contained until force stop proves it stopped. | | K9 | An interrupt that never settles | Stays `uncertain`. Force stop stays available, and ordinary takeover is refused while fenced. | | K10 | Controller killed between the TERM and kill phases | Restart checks the invocation ID and re-runs from TERM for the same stop. Nothing is recorded as killed that wasn't observed. | | K11 | Controller killed after the confirmation is recorded, before TERM | Same as K10 | | K12 | A member that forks in a loop during enumeration and kill | The freeze stops forking. The enumeration is complete, and `populated 0` follows `cgroup.kill`. | | K13 | A member writes its own pid to another cgroup's `cgroup.procs` | Refused by the namespace, and the kill is complete. If not refused, `stopped` is unavailable for real cohorts, and K1 expects `uncertain`. | | K14 | A unit with the recorded name but a different invocation ID | Evidence unavailable. No signals, `uncertain`. | | K15 | `engine` cgroup path missing or unreadable, or the shim gone | Evidence unavailable, not empty. `uncertain`. | | K16 | Boot proof requested for a claim from a different machine ID | `foreign-host`. No proof. | | K17 | Two launcher calls with one eligibility record | One launch. The other refuses, and no second engine starts. | | K18 | The leaf changes after eligibility, before launch | Launch refused. The reservation stays until released with proof. | #### 7. Claude permission requests (I4) - Pi native dialogs are cut from CHAT-03 (lead decision 30). CHAT-01 lines 239–242 and 292 block non-permission forms until CHAT-03D exists, so a Pi `extension_ui_request` is shown disabled, with a reason, and is never answered. - **Claude, in I4 after B1.** `can_use_tool` maps to allow-once and deny. Choices that widen permissions are disabled, and there is no invented cancel (lines 294–295). Pending permission requests are re-read from initialize on reconnect, but only if B1 shows that the pinned version supports it. | # | Fixture | Expected | |---|---|---| | P3 | A Pi `confirm`, `select`, `input` or `editor` dialog | Disabled, with a reason. No response is sent. | | P7 | A Claude permission request (I4, recorded fixture) | Allow-once and deny only. Widening choices disabled. | #### 8. Claude: B1 and the catalogue (carry-forward 2) **What B1 is.** `docs/plans/chat-00/README.md` lines 193–195: prove the exact Claude CLI and protocol combination, capability negotiation, transcript branch selection and the image and file input shapes. SDK and source examples narrow the uncertainty, but they don't certify the installed binary. CHAT-00 line 66 names the gap for history: Claude's persisted branch format and leaf selection. **What proves it (I3).** A B1 packet under `agents//work/chat-03/b1/`, approved by Filbert and Rocko, with five parts: 1. **Pin.** The exact version string and the sha256 of the binary the adapter will run. Today `claude --version` reports 2.1.283; the plan saw 2.1.269. The adapter checks both at launch and otherwise refuses with `engine-pin-mismatch`. It runs the engine with auto-update off, and the author cites the setting from the pinned docs. 2. **Sources.** The official protocol sources for that exact version, by hash. These are the C-* entries in `chat-00/sources.json`, re-pinned to that version. C-HEADLESS and C-REF are floating documents (`sources.json` lines 106–119) and can't be tied to a version. The packet fixes them as fetched copies, each with its sha256 and fetch date, and says plainly that they aren't version-bound. 3. **Recordings.** Made from that binary, in a scratch config directory with no repository credentials: - initialize and its capability answer, including whether `interrupt_cancel_queued_v1` and `pending_permission_requests` exist; - a text prompt and an image prompt; - a `can_use_tool` allow and a deny; - an interrupt. They are redacted, hash-pinned and used as the I4 fixtures. The plan rules out model calls for discovery (line 166). Any recording that calls a model or reads Claude auth therefore needs Jason's go first, because it spends money and touches credentials. Without that go, I3 stops after parts 1 and 2, and B1 stays open. 4. **Branch selection.** Two persisted transcripts from that binary, one of them with a fork (an edit or a rewind). The packet states the leaf rule and shows that `--resume` continues the leaf the parser picks. 5. **Negative.** The parser refuses a transcript from another version with `unsupported-harness`. **The catalogue after B1 (I4).** - Rocko's launcher runs Claude with the default config directory (`agents/rocko/launch.sh:86`, no `CLAUDE_CONFIG_DIR`). So its transcripts sit in the same `~/.claude/projects//` directory as every other Claude session started in this checkout, including T3 threads and Jason's own sessions. - The catalogue therefore never lists that directory. It opens only session IDs the repository launcher recorded for that seat (`.pi/state/rocko/session-id` and `launches/*/session-id`). Each ID must resolve to exactly one file under the approved root, opened with the CHAT-02 safe-open rules. - A launcher receipt is a hint, not membership. CHAT-01 lines 62–64 require an approved project and workspace mapping, and say OS readability, cwd and seat labels don't establish it. The access mapping stays under B4 review. - The CHAT-02 guarantees still apply: read-only, no writes (F17), cursors, byte caps, continuation parts and inert rendering. | # | Fixture | Expected | |---|---|---| | L1 | A Claude transcript at the pinned version | Default leaf shown and other branches readable, as in F12 | | L2 | A transcript from another version | `unsupported-harness` | | L3 | A decoy session file in the same directory, named in no receipt | Never listed, never opened | | L4 | A receipt naming a file outside the root, a symlink or another project | Refused, as in F7–F9 | | L5 | Sidechain or subagent entries | Kept separate, not merged into the main branch | | L6 | No-write check | F17 over the Claude directory | | L7 | The binary's sha256 or version differs at launch | `engine-pin-mismatch`. No launch. | The I4 adapter maps Claude stream-json onto CHAT-01 events. The §5 and §9 fixtures then run again against the recordings. #### 9. Return flow and events | # | Fixture | Expected | |---|---|---| | E1 | The plan §6 return flow on the fake engine: send, acknowledge, user, toolCall, toolResult, new final answer | Shown once in the same conversation, with no refresh, no duplication, and the draft and reading position kept | | E2 | U+2028 and U+2029 inside JSON strings, and CRLF | Each parsed as one record | | E3 | Multipart final, two blocks, null request correlation, duplicate delivery | The CHAT-01 cases at lines 110–112 | | E4 | A page read just after a `message_end` but before its entry is persisted, then a subscription; again with a gap or a new epoch | Replay is advertised unavailable. The client shows a reconcile marker at the seam and re-reads the page after `agent_settled`. After that, each message appears exactly once. A gap or new epoch also reconciles. If the author shows an atomic cut from the source, overlap is deduplicated instead, and the fixture pins that case. | | E5 | An unknown native event | No client event. Controller evidence records its type and byte count, and the terminal count goes up. No record fails the schema. | | E6 | A tool result delayed across a pause and a reconnect | Reconciled without a manual refresh | | E7 | The mediated terminal as observer, then as controller | Renders the same stream as the library client, and submits only as controller | The browser leg of the return flow is CHAT-05. I1 proves it at the library and the terminal. #### 10. Checks and mutation pass - **Suites.** These must pass: - `node --test packages/conversation/tests/`; - the control-board, webui and seat package tests; - `node docs/plans/chat-00/check.mjs`, `chat-01/check.mjs` and `chat-01c/check.mjs`; - every `scripts/test-*.sh` suite at the candidate's base, including `test-queue.sh`, since queue A1 landed in `34a72af9`. - **Nested runners.** A test that starts `node --test` itself clears `NODE_TEST_CONTEXT` for that child, or it can hide failures. The DEFERRED entry on nested runners (Filbert's A1 review `6933b885`, note N13) was closed at `34a72af9`, and new tests keep to the same rule. - **Isolation.** Every fixture runs in temporary directories. No fixture touches live registrations, `.pi/state`, `~/.claude` or a live seat (plan line 272). - **Mutation pass.** Run in a scratch copy, as for the CHAT-02 Console. Each of these mutants must fail at least one test: 1. dispatch skips the generation recheck; 2. `abort` is sent before `clear_queue`; 3. a revision is published by rename or overwrite, so an existing revision name can be replaced; 4. the late-event filter ignores the incarnation; 5. the text policy checks only the first character, missing leading whitespace; 6. an acknowledged SIGTERM promotes a stop to `stopped`; 7. the process-group fallback can reach `stopped`; 8. a confirmation is not consumed on use; 9. disconnect releases control; 10. the composer isn't cleared on transfer; 11. engine output is read with `readline`; 12. a revision is published by exclusive create in place of the temp-file-then-`link()` sequence, so a partial file becomes visible; 13. a second controller reclassifies a claim whose owner is alive; 14. the live-session guard skips the real-path check; 15. the pending slot is released at `agent_start` instead of at reconciliation; 16. a pipe is written again after an unknown write outcome; 17. an ack followed by no run marks the item `working` or `finished` instead of `delivery-unknown` (N5); 18. admission reopens without the post-settle empty clear; 19. `abort` is sent after a `clear_queue` timeout; 20. the incarnation token check is skipped; 21. a missing cgroup path counts as empty; 22. a unit with a different invocation ID is signalled; 23. a restart during force stop records the kill phase as done; 24. an eligibility record can be used twice; 26. a non-empty `clear_queue` reaches `reconciled`; 27. the pending slot's in-flight preflight is abandoned at the fence instead of awaited, so a late ack's run is never aborted (N4); 28. the fake engine queues a Mosaic prompt while streaming instead of throwing (N10's conformance check must catch it); 29. a no-unit observation frees a pair that has a spawn marker (W20); 30. a settled slot, an idle engine and empty clears are taken as interruption evidence, so `turnState: interrupted` is written without a final `stopReason: aborted` (N14 and N15 must catch it); 31. a receipt whose run ended `stop` is relabelled `failed` (reason `interrupted`) because a stop was in progress (N14); 32. `turnState: input-reconciled` is used to reconcile an Interrupt (N3, N14, N15); 33. Interrupt with no slot held and no visible run creates a stop record or writes to the engine (N16); 34. a run is attributed to the slot with the overlap checks skipped, so the receipt settles `finished` or `failed` after an overlap signal, or the binding stays `active` (N8, N19, N20); 35. a settle that doesn't close the slot's own run (no `agent_start` since the ack, or an `agent_start` with no `agent_end`) is taken as the slot's settle (N8, N19); 36. an `agent_start` or `agent_settled` with no slot held is ignored, so admission stays open (N13, N21); 37. the binding launches without the `--no-*` flags, or with an `--extension` argument (N24); 38. an `ack-without-start` run settles `failed` (N11); 39. an empty clear is recorded as proof that no external input was removed, instead of naming the seal as the basis (N22). Mutant 25 (drift) left with the drift check for CHAT-07, and its number isn't reused. I4 adds two more: the pin check is skipped, and the decoy file is listed. #### 11. Choices in this brief, for review | Choice | Tradeoff | |---|---| | No broker queue; a busy engine refuses `busy` | No queued follow-ups until CHAT-04, and no native queueing of mediated input | | No `streamingBehavior`, `steer` or `follow_up`, and one pending slot | Same as above. On pinned Pi a Mosaic prompt then never enters a native queue (lines 860–863). It can still lose a race in preflight to another run (lines 860–949), which the seal and the overlap signals handle (§3). | | Sealed engine: `--no-extensions` and no `--extension` (`unsealed-engine`; lead decision 31) | Without run IDs, only excluding every other source of turns lets a run be attributed to the slot by order. The cost is that no explicit extension loads in CHAT-03, goal included (Limit 11). They come back in CHAT-06. The overlap signals watch for a seal failure. | | Overlap signals close admission and make the binding `uncertain` (`run-overlap`) | One overlap needs a force stop to move on, and never produces a guessed receipt. Some seal failures give no signal (Limits 7 and 11). | | The slot settles only on evidence tied to its own prompt, never from `clear_queue`, idle or a settle alone | An ack with no run, and a run that fails before its user message, stay `delivery-unknown`: no "unsent" state and no safe-resend button | | Receipt settlement and the stop outcome are classified separately (§3 rules 4 and 5) | A turn that finishes during an Interrupt is reported as finished, not interrupted. The cost is that the stop can't reconcile: until C-5 it stays `uncertain` and needs a force stop, even though nothing was lost. | | Interrupt with nothing running refuses `no-turn` | Narrows the completion race to the clear and abort window. The controller can refuse just before it reads a new run's `agent_start`; the run keeps going and the actor retries. | | Poison the pipe on an unknown write | One transport glitch needs a force stop, with no guessing about half-written lines | | Live-session guard by path | Fixture-only is enforced in code. The guard needs a reviewed change to lift at CHAT-07. | | No default claim root | Nothing live can use I1 before the CHAT-07 data-map decision | | Cohort is an `engine` child cgroup in a delegated user scope, held by a shim, with a cgroup namespace against migration | Depends on `systemd-run --user`, delegation, `cgroup.freeze`, `cgroup.kill` and user namespaces. Without them, or if K13 fails, stops end at `uncertain`. The verifier is fixture-grade until B3/B4. | | One actor, `local-operator`, guarded only by socket-directory mode | Same-uid agents are not kept out (Limit 1) | | Dedup lasts one controller incarnation, fenced by a token (deviation V-1, accepted with limits in lead decision 25) | A retry across a restart refuses `stale-incarnation` instead of returning the receipt. The client shows "outcome unknown, check the transcript" and never resends. CHAT-04's durable receipts must restore the contract behavior. | | Three increments, with I3/I4 gated on B1 | CHAT-03 can close Pi-only, with that outcome recorded | | Claude catalogue from launcher receipts only | A Claude session started outside the launcher never appears | ### Carry-forwards | # | Sage's item | Where | Acceptance evidence | |---|---|---|---| | 1 | D1 writer-claim record and R3-1 (lead decisions item 8) | §1, §2, §3 | W1–W9, W11–W17, W20, G1–G3, N1–N25 and H18–H23 pass. Mutants 2, 3, 12–20 and 26–39 fail tests. R3-1 binds only sealed engines, settles each receipt only on evidence tied to its own prompt (§3 rule 4), and classifies the stop separately (rule 5). Any overlap signal closes admission. Only an interrupted turn reconciles. The other stop outcomes leave the stop `uncertain` until C-5 lands in CHAT-04 (lead decision 30). | | 2 | The Claude catalogue after B1 | §8; increments I3 and I4 | The B1 packet (parts 1–5), approved by Filbert and Rocko. L1–L7 pass. | | 3 | The board send that can become a Pi slash command | §4 | S1–S7 pass. Mutants 5 and 10 fail tests. The DEFERRED entry stays open for seats still on tmux. | | 4 | The CHAT-01/01C contracts, by path and hash | "Contracts implemented" | The hash table matches `sha256sum` at approval. C-3 is a separate reviewed item if B1 needs it, and C-5 moves to CHAT-04. C-1, C-2 and C-4 are cut (lead decision 30). Deviation V-1 is accepted with limits (lead decision 25): H21 and H23 pass, the client never resends, and V-1's closure is a required CHAT-04 item (lead decision 25). CHAT-04's brief must list it. | | 5 | Path authority | "Files owned" | The overlap table. Every changed path in each candidate is inside "Files owned". | | 6 | The source author | "Owner and reviewer" | ``, named by Darkwing after this brief is approved | ### Limits 1. **Same-uid control.** Any process running as the same user can connect to the socket and take control. Agent seats run as that user, so an agent could take over another conversation. CHAT-03 is fixture-only, so no live conversation is exposed. Live use waits for a reviewed local owner-channel design (B4). 2. **Fake engines.** They model the pinned docs. The real-binary smoke covers framing only. I1 makes no model call, so real-engine behavior is unproven until CHAT-06. 3. **B2 untouched.** Tool isolation and trust and settings parity (CHAT-00 lines 196–197) are untouched. Any real engine for a real seat waits on B2. 4. **Containment.** The cgroup covers local processes. Network and remote effects stay `uncertain`. A same-uid process outside the cohort can still move members out. The verifier is the fixture registry. No CHAT-03 stop proof is live authority until B3/B4 and B2. 5. **Sole writer.** CHAT-03 has no session drift check; it moves to CHAT-07 (lead decision 30). The claim binds only controllers that use it. Real sessions stay off-limits through the live-session guard. 6. **Durability.** W5 kills processes at each barrier. Real power loss isn't tested; the fsync-then-`link()` order is the stated basis. 7. **What the clear can't see.** Under the seal no extension queues input, so the stop proof rests on the seal. If the seal fails, three things escape it. Input queued after the final empty clear is outside the proof (§3 rule 7), and it shows up only if it starts a run (O1) or queues visibly (O5). Custom messages an extension queues straight into the agent while it streams (`agent-session.js` lines 1112–1118) are dropped by the clear without being returned (lines 1195–1203) and aren't counted in `get_state` (line 1206). `nextTurn` messages (line 1110) survive clear and abort and attach to the next Mosaic prompt (lines 910–913). The last two give no signal at all (N22, N23). 8. **Takeover has nothing to recover.** With no broker queue, a takeover has no queued input to turn into drafts. Q13 is CHAT-04's. 9. **Interrupt racing completion.** A turn can finish, fail, or never start while the clear and abort are in flight. The receipt reports what happened, but until C-5 the stop can't reconcile and needs a force stop. Sessions with short turns will see this more often. 10. **Restart dedup (V-1).** Within one incarnation a retry returns the receipt. Across a restart it refuses `stale-incarnation`, and the actor has to check the transcript. CHAT-04's durable receipts close this. 11. **Sealed engines only.** Pinned Pi ties no run to a prompt, so CHAT-03 attributes runs by order under the seal. A session that loads any explicit extension, goal included, can't bind (`unsealed-engine`, lead decision 31). The five Pi repository seats (Darkwing, Dewey, Filbert, Researcher, Sage) load goal through `scripts/agent-host-dev.sh` line 137, so none of them can bind with goal loaded. Under lead decision 30, a seat bound to the Console runs without goal, and goal continuation is a CHAT-06 item. The seal rests on the launch argv and the pinned Pi build. A failure there that the overlap signals don't catch goes undetected if it queues straight into the agent (Limit 7). Other failures are caught late. Another run's user `message_start` that arrives before any signal puts the item in `working`, and the signal that follows can only mark it outcome unknown (N8). A spurious settle delayed until after the receipt settled signals only once the receipt is final (N21). CHAT-06 owns explicit and turn-starting extensions, which need run-to-prompt evidence from Pi. ### Out of scope - Durable queues, drafts, uploads, queue edit and cancel, takeover-to-draft (CHAT-04). - Remote control (CHAT-04R) and the chat UI (CHAT-05). - The candidate, cutover and any real seat migration (CHAT-06, CHAT-07). - Changes to `tools/tmux`, the board, the seat package, launchers, `roles/`, `contracts/` and the `docs/plans/chat-0*` contracts. - The CHAT-03I fleet-communications charter (CHAT-01 lines 232–238). - Authenticated multi-actor control and the local owner-channel design (B4). - The Claude adapter and history, unless B1 passes. - Filbert's idsDigest note, which stays in `agents/dewey/work/chat-02/FOLLOWUPS.md`. - Explicit extensions, goal continuation among them, and any pinning of them (CHAT-06; Limit 11; lead decisions 30 and 31). - Cut by lead decision 30: Pi native dialogs and the CHAT-03D companion (C-2), the unknown-event type (C-4) and the removed-input `turnProof` (C-1). C-5 moves to CHAT-04, and the idle session drift check to CHAT-07. ### Gate - **Brief.** - Filbert approves the exact `BRIEF.md` hash. - Rocko's pass on §5 and §6 leaves no blocking finding open. - Sage commits and pins the brief. Darkwing then names the author. - **Each increment.** - Filbert approves the exact candidate hashes. Rocko passes the §5 and §6 code in I1 and the B1 packet in I3. - Every §10 suite is green on an index export, and every mutant is killed. - Sage commits. A push needs Jason's word. - **Contract items.** Each is a separate reviewed change to the CHAT-01 files, never folded into an increment. - C-3, if B1 shows a gap, is approved before I4 starts. - C-5 moves to CHAT-04 (lead decision 30). CHAT-03 ships the interim rule: every stop outcome except Interrupted stays `uncertain`. - Deviation V-1: accepted with limits (lead decision 25). I1 is approved only with H21 and H23 passing and the client's "outcome unknown, check the transcript" display. V-1 stays open until CHAT-04's durable receipts restore the contract behavior; CHAT-04's brief must list that as a required item. - **I3.** Jason's go comes before any model call. - **CHAT-03 done.** All of these: - I1 is approved and committed; - I4 is approved and committed, or B1 is recorded as not passed (the triggers are in "What ships"), with Claude still refusing and CHAT-06 blocked. - **No live check.** No real seat migrates, so there is none. The live proof belongs to CHAT-07.