Files
stack/agents/dewey/work/chat-03/BRIEF.md
T

107 KiB
Raw Blame History

CHAT-03 brief: live adapters and mediated terminal (#1507, row 5)

Author: Dewey, 2026-09-26/27. Final text for pinning: R3 (2c5be6b4, frozen as BRIEF-r3-2c5be6b4.md) with the one edit lead decision 32 orders. It is not a review round (lead decision 27). This is the brief only. It includes no source, no contract edits and no seat changes. R1 (5dd447f7) is frozen as BRIEF-r1-5dd447f7.md, and R2 (5c5b45a2) as BRIEF-r2-5c5b45a2.md. §0 lists what changed.

Jason chartered CHAT-03 on 2026-09-26 (docs/plans/2026-09-26_lead-decisions.md item 22). Plan row: docs/plans/2026-09-13_webui-session-chat.md line 215. It depends on CHAT-01 (28d4e98a) and implements part of the CHAT-01C companion (b023841c). The layout follows docs/plans/BRIEF-TEMPLATE.md, and the depth follows the CHAT-02 brief (agents/dewey/work/chat-02/BRIEF.md, sha256 636b0fac…). The six carry-forwards Sage set are mapped to their evidence under "Carry-forwards".

Line numbers and hashes below are at f2b9e622. Between 40a02d2b (the R2 base) and f2b9e622, no contract, source or pinned Pi file changed. Queue A1 landed (34a72af9) with docs/plans/BRIEF-TEMPLATE.md, lead decisions gained items 26 and 27, and the goals review (docs/plans/2026-09-27_goals-review.md) was added.

Lead decision 27 made R3 the last review round, and its answer to anything pinned Pi can't prove is to refuse or report unknown, never new machinery to prove it. Sage rescoped CHAT-03 against Gate E from R3's section map (lead decision 30), ruled on Rocko's R3 finding (lead decision 31) and ordered this edit (lead decision 32).

0. Changes since R3 (lead decision 32)

Filbert approved R3 on the sections Gate E keeps, with no blocking finding (agents/filbert/work/chat-03-brief-review-r3-2026-09-27.md, 48447592). Rocko closed his R2 finding and raised one blocking finding in the seal (agents/rocko/work/chat-03-r3-adversarial-2026-09-27.md, 19e3fcff), which lead decision 31 resolves. This edit does the five things lead decision 32 lists, and nothing else.

Item What changed
1. Lead decision 30 cuts, removed and not reworded Removed: C-1, C-2 and C-4; increment I2 and §7's Pi dialogs (P1, P2, P4–P6); increment I1b (C-5 moves to CHAT-04, and the interim rule stays); the §2 idle drift check, W10, W18, W19, mutant 25, the drift choice and the session-drift name. H5–H8 stay for Claude in I4. References to removed items were adjusted where they stood: What ships, E5, carry-forwards 1 and 4, §11, Limits 5 and 11, Out of scope and Gate. Two sentences were kept by moving them: the claim's reach, into the live-session guard bullet, and P3's disabled Pi dialog, which is now §7's only Pi rule.
2. Lead decision 31: no explicit extensions The seal is --no-extensions, --no-prompt-templates, --no-themes and no --extension. The registry, the input-silent review and the tree hashes are removed, not replaced with pinning. N5 and N15 are fake-only. N24 and mutant 37 refuse any --extension. Limit 11 and §11 say explicit extensions return with goal in CHAT-06.
3. Filbert n1: the Pi pin Pi is pinned by the package-lock.json integrity of @earendil-works/pi-coding-agent 0.85.1, checked against npm's installed record. The listed dist/core and dist/extensions files are now labelled as citation sources, since pi runs dist/bundle. The built-in llama.cpp is pinned with the package.
4. Filbert n2: the clear's own queue_update O5 counts the empty queue_update Pi's clear emits before its response as the controller's own. N10's fake emits it, and N25 proves that an ordinary Interrupt reconciles.
5. Rocko's R3 note A build note under What ships: the no-turn cleanup lifts only its own fence. H10 covers the force-stop, overlap and revocation branches.

Filbert's n3 (an aborted with no stop in progress) isn't in lead decision 32, so it isn't applied.

0a. Changes from R2 to R3 (historical)

R3 answers Rocko's review of R2 (agents/rocko/work/chat-03-r2-adversarial-2026-09-26.md, 07b938fb), Filbert's review of R2 (agents/filbert/work/chat-03-brief-review-r2-2026-09-26.md, d149e8cc), Sage's ruling on deviation V-1 (docs/plans/2026-09-26_lead-decisions.md item 25) and Sage's rule for all three R2 blocking findings: a receipt settles only on positive evidence tied to its own prompt. Rocko closed R1 findings 1–7, and Filbert closed B2–B6 and his R1 notes. The full disposition table is in REVIEW-REQUEST.md.

Finding What changed
Rocko R2 finding 1 (blocking): a settled slot, an idle engine and empty clears don't prove an interrupted turn §3 splits the old slot rule. Rule 4 settles the item's receipt from its own evidence, and routes a late ack through the same table. Rule 5 classifies the stop: Interrupted, Completed first, Failed on its own, No run, Unknown. A normal completion stays finished. Only Interrupted gives turnState: interrupted and reconciles; input-reconciled isn't borrowed. The rest leave the stop uncertain, admission closed and force stop available, until contract item C-5 (new, separate, reviewed). Rule 6 (reopening) requires Interrupted until C-5. Interrupt with nothing running refuses no-turn. N14 and N15 cover Rocko's completion and handled-ack schedules, and N3 now covers his no-run preflight error. N16–N18 add no-turn, Failed on its own and Unknown. N3, N4 and N9 assert the receipt and the stop. A receipt never moves back from working, and rule 8 states the receipt order. Mutants 30–33. Limit 9.
Rocko R2 note 2: N11 wording The ack-without-start note says Pi's ordering is local to Pi, and that the ordered output chain says nothing about which prompt a settle belongs to. With F1, the outcome is now delivery-unknown, never failed (next row). The pi-agent-core run-failure lines (349–364) are cited, and its two files are pinned by hash. N11 adds the unparseable-line case.
Filbert R2 F1 (blocking): an extension run can overlap a Mosaic prompt's preflight, so order proves nothing §3 R3-1 drops the "at once or not at all" premise and states the verified window (lines 860–949 and their awaits), the acked loser, the swallowed throw, the spurious settle and the multi-pair run. A binding needs a sealed engine: --no-extensions and a registry of hash-pinned, input-silent extensions, otherwise unsealed-engine. The seal is named as the basis for attributing by order. Overlap signals O1–O6 close admission and make the binding uncertain (run-overlap). R2's failed row for a run that fails before its user message becomes delivery-unknown / ack-without-start, and working requires that no overlap signal has been seen. A misattribution to working that a later signal exposes is shown as outcome unknown. Goal can't bind in CHAT-03 (Limit 11, a CHAT-06 item). N8, N10, N13 and N19–N21, N24; mutants 34–37.
Filbert R2 F2 (blocking): an empty clear_queue doesn't prove that nothing was removed Rule 3 narrowed to Pi's steer and follow-up queues, with the seal as the basis. C-1's premise doesn't survive, so C-1 is withdrawn and restated for CHAT-06 ("Contracts implemented"), and I1b is C-5 only. Limit 7 names the agent-level and nextTurn gaps. N22 and N23.
Filbert R2 notes n1–n4 n1: the ordering note cites rpc-mode.js 28–29 and output-guard.js 71, and makes no transport claim. n2: recovery reads the file before get_entries (§2). n3: the drift baseline is taken after load; W18 adds the truncated-line, empty-file and migration cases; "before the first assistant message" applies to new files only. n4: H17 cites CHAT-01C 92–102; §1 cites rpc.md 56–65 and 78–104; agent.js, output-guard.js and the goal extension are pinned.
Rocko R2 remark on W20 A marker on either key counts, and completing a pair copies it. W20 runs with the marker on one key only.
Sage, lead decision 25: V-1 accepted with limits V-1 is accepted. The client shows stale-incarnation as "outcome unknown, check the transcript" and never resends. H23 proves that no request pending across a restart reaches the engine. V-1 closes only when CHAT-04's durable receipts restore the contract behavior, and CHAT-04's brief must list it. Gate, carry-forward 4, §11 and Limit 10 updated.
Consistency passes Two passes over the R3 draft found 22 defects of wording and cross-reference. A third pass over the finished draft found 17, including three where a fixture's expected result didn't follow from the rules (N6, N8, N20). All are fixed.
Base Line numbers and hashes move to f2b9e622. No contract, source or pinned Pi file changed since 40a02d2b. Queue A1 is now committed, so the overlap table names its commit.

0b. Changes from R1 to R2 (historical)

The rule numbers in this table are R2's. R2's rule 6 is R3's rule 7.

R2 answered three reviews of R1:

  • Rocko's adversarial review (agents/rocko/work/chat-03-r1-adversarial-2026-09-26.md, 89752c2b) and his addendum narrowing finding 3 (…-addendum-2026-09-26.md, e5b6bd00);
  • Filbert's review (agents/filbert/work/chat-03-brief-review-2026-09-26.md, ec00544e).

Sage's rule for R2: where a guarantee can't be proven on pinned Pi, the brief refuses or reports uncertainty instead of claiming it. The full R2 disposition table is in REVIEW-REQUEST-r2-c75ad86f.md.

Finding What changed
Filbert B1, Rocko 3 and addendum: R3-1 premise §3 R3-1 rewritten from the source. A mediated prompt can't enter Pi's queues: without streamingBehavior it throws while isStreaming (agent-session.js lines 860–863, 617, 773, 348). clear_queue returns only external input. The residual risks are a fence during preflight and an ack with no run. No text-identity machinery. External input queued after the last clear is stated as a limit (rule 6, Limit 7). N1–N13 rebuilt.
Filbert B2, Rocko 6: no entry IDs on the stream §2: the drift check compares disk entries with the engine's own get_entries list (rpc.md lines 717–745), not with stream IDs. It runs at idle. §3 advertises replay as unavailable. W10, W18 and W19.
Filbert B3: schema gaps C-1 is repurposed as a turnProof field for external items that clear_queue removed. Until it lands, a non-empty clear leaves the stop uncertain. C-4 is an optional event type for unknown native events, with an interim rule.
Filbert B4 §2 live-session guard (live-session-refused, symlinks resolved), G1–G3, mutant 14. CHAT-07 lifts it.
Filbert B5: order and gate Each increment starts when its "Needs first" column is met. Gate adds C-1 adoption or carry, a B1-not-passed trigger and owner, and what happens if C-2 is declined. Increments renamed I1–I4 so they don't collide with CHAT-03D.
Filbert B6, Rocko 1: claim gaps §2 rewritten: one claim ID across both keys, atomic link() publication, the more conservative key wins, restart rules for reserved and half-done pairs, a scope named from the claim ID, and a no-unit path to stopped that a spawn marker closes. W4, W5, W12–W17, W20. The restart dedup is recorded as deviation V-1 for Sage.
Rocko 2: pending prompts, partial writes §1 pending dispatch slot and three write outcomes. A partial or failed write poisons the pipe. H3, H18–H20.
Rocko 4: cohort proof §6 rewritten: shim-held scope, engine child cgroup, invocation-ID epoch, freeze, enumerate, cgroup.kill, unavailable distinct from empty, the K13 migration test, and fixture-grade verification (CHAT-01 lines 330–333). K1–K16.
Rocko 5: dedup across restart §1 incarnation token and stale-incarnation. H21, H22. Deviation V-1.
Rocko 7 (non-blocking) §6 single-use eligibility with a new reservation. K8, K17, K18.
Filbert's notes N1–N12 (non-blocking; not the N fixtures) Skills are live (§4, S2). Refusal names generation and controller reused. Citations corrected. B1 floating sources fixed by hash. Smoke isolated. Suite count by glob.

CHAT-03: live adapters and mediated terminal

Problem

CHAT-02 made every repository Pi conversation readable. Nothing can drive a conversation yet except the old paths: the Pi TUI in a tmux pane, and board Reply or agent-send, which paste into that pane. Those paths have no controller, no generation fence and no stop proof. At 2026-09-26T20:10:57Z a leftover / in Pi's composer was concatenated in front of a board reply. Pi read the result as plain text that time, but the same path could run a reply as a slash command, because tools/tmux/send-message.sh (lines 45–54) pastes onto whatever the composer holds and then presses Enter (docs/plans/DEFERRED.md, "Board send can turn into a Pi slash command").

CHAT-01 defines the records and commands for one controller per conversation. It defers the writer-claim record (chat-01/README.md:342), and CHAT-01C defers R3-1 (chat-01c/README.md:217). Lead decisions item 8 moves both here. The Claude catalogue still refuses unsupported-harness (packages/conversation/src/reader.mjs: constant at line 24, refusals at lines 51 and 65). Jason ruled that CHAT-03 owns it after B1. Claude on this host reports 2.1.283, but the plan's evidence named 2.1.269 (plan line 159). The binary changes under any adapter that does not pin it.

Owner and reviewer

  • Source author: <slot>. The plan has Darkwing name the backend author (line 215; lead decisions item 22). Darkwing is on queue A1/A2, so the author is named after this brief is approved. The brief does not pick one.
  • Brief author: Dewey.
  • Filbert reviews this brief, then the exact candidate of each increment.
  • Rocko does an adversarial pass on §5 (control races) and §6 (stop and recovery), in this brief and later in the code of each increment. Rocko also reviews the B1 packet (§8), because it is evidence that a false pass would turn into a wrong adapter.
  • No one reviews their own work. If Darkwing names Filbert or Rocko as author, Sage names a replacement for that review.
  • Tracking: #1507.

Files owned

The source author may create or change only these paths.

Path Contents
packages/conversation/src/** New modules for the controller, claim, supervisor shim, Pi adapter, events, transport and mediated terminal. The Claude parser and adapter come in I4. reader.mjs changes only to lift the Claude refusal for the pinned version.
packages/conversation/tests/** Tests, fake engines, fixture extensions and redacted recordings
packages/conversation/README.md, packages/conversation/package.json Documentation and the new refusal names. No new dependencies.
agents/<author>/work/chat-03/** Review packets, evidence and the B1 packet

Module names inside src/ are the author's call (plan lines 150–151). Any path outside this table needs a brief amendment.

Excluded, and why:

  • tools/tmux/**. Plan line 244 excludes it, and §4 shows the mediated path needs no change there.
  • packages/control-board/**. No board edit is needed (§4, S6). Tests import replyToRow read-only by relative path, the same way the queue-as-data plan imports scan.mjs.
  • packages/seat/**, scripts/agent-host-dev.sh, agents/*/launch.sh. No seat migrates, so no seat or launcher changes. The launcher handoff belongs to CHAT-07.
  • packages/webui/**. The chat UI is CHAT-05.
  • docs/plans/chat-0*/**. Contracts change only through the C items below.
  • roles/**, contracts/**, root files, auth and ~/.mosaic.

No overlap with queue-as-data or the ledger:

Owner Paths Overlap
A1, Darkwing, committed 34a72af9 (agents/darkwing/work/queue-a1/build-manifest.sha256) packages/queue/**, scripts/queue-commit.sh, scripts/test-queue.sh, scripts/git-hooks/pre-commit, docs/plans/BRIEF-TEMPLATE.md none
A2, Darkwing (Filbert's queue-as-data plan 282fabbb… §8.1 and lines 130 and 776–777; lead decisions item 20) scripts/mosaic dispatch, docs/plans/queue.json, docs/plans/QUEUE.md, scripts/test-{darkwing,rocko}-launch.mjs none. The mediated terminal runs as node packages/conversation/src/<cli>.mjs. It is not a scripts/mosaic verb.
Piece B, after A AGENTS.md, agents/*/CONTEXT.md none
Ledger packages/ledger/** none

The author may import the process-identity helpers in packages/discord/src/journal.mjs read-only, as A1 does. That changes neither package.

Contracts implemented

These are unchanged since the commits named. The sha256 values are at f2b9e622, and git status shows no local edits.

Path sha256 Commit
docs/plans/chat-00/README.md 991663e607404c2f022716bd45b40d5b3092d53c3f9f61bbafd455c0659b832f 370823b3
docs/plans/chat-00/sources.json 1a07ae88de45fd3219eca10acae13637ca75416cb1c598aa9080825bd7598af9 370823b3
docs/plans/chat-00/fixtures.json ad9d4fe94131b2c4bb30691ad846698c5b0818f0935d20a38de73fa3f267bd80 370823b3
docs/plans/chat-00/check.mjs 568f004f5a0cfce64d65ff202c420ab491e716919a93989377a7bb9e7b8be622 370823b3
docs/plans/chat-01/README.md 61aba7d60f380ff8a135a04a2e851f11c250795ead11cca2e594566849483163 28d4e98a
docs/plans/chat-01/contracts.schema.json 38382e08c97635f8864e6953b6cf44ec324040397a1aa8fc3f86e279abe36db1 28d4e98a
docs/plans/chat-01/fixtures.json 403c8ae91963a379de17805380bc1425bc4e96c1ef34094385b703d17cf5ce7d 28d4e98a
docs/plans/chat-01/check.mjs 2e164e4bfa61963bb5dd8639e26e67407e64d19278cc56ed8a4f77e430f1dee5 28d4e98a
docs/plans/chat-01c/README.md 63d9f9edffc2aa1aa3bc99364b864e8c6f234eedc8ad2f9da231531970d54250 b023841c
docs/plans/chat-01c/contracts.schema.json da132f0a02f29281344cc5350396af53c893d7895308bc430f9bf3ca3b3785b9 b023841c
docs/plans/chat-01c/fixtures.json 00639a219b00b67c5e0ab904163a6d945481f6b0df45a058b32e00f1f1326611 b023841c
docs/plans/chat-01c/check.mjs 5971ed5f11a32f0dff5720db03d8d4bc4c007a5d2654ba50726bb35ba753eba6 b023841c

Design inputs, which CHAT-03 does not implement in full: docs/plans/foundation-v1-candidate/RUNTIME.md §§3–5 (b1a2b4d0…) and the plan itself (48142829…).

What CHAT-03 takes from each contract:

  • CHAT-01. The records binding, connection, clientRequest, request, receipt, event, stop, cohortProof, effectReport, turnProof, nativeDecision, approval and confirmation. The commands observe, prompt, takeover, acquire-recovery-control, approval, interrupt, force-stop, recover, issue-confirmation and answer-confirmation. The draft, upload and queue-edit commands belong to CHAT-04.
  • CHAT-01C. Only the confirmation restoration rules (lines 92–104). The private-state pages, upload ranges and recover-refused-draft need CHAT-04's durable store.

The CHAT-01 text at line 342 stays as published. Lead decisions item 8 records the move.

Contract changes, each a separate reviewed item. Each one lands and is approved before the code that needs it. Sage assigns the author; Filbert and Dewey review, as they did for CHAT-01C.

  • C-3: Claude record fields. Needed only if B1 shows that CHAT-01 lacks a field, for example the capability-negotiation result on binding. Otherwise C-3 is void.

  • C-5: interrupt outcomes other than an interrupted turn (R3). turnProof.turnState allows interrupted, unknown and input-reconciled, and reconcile-interrupt passes only with interrupted (check.mjs line 315). A turn that completed or failed before the abort took effect, or an Interrupt that found no run, has no honest value. C-5 adds values for those (for example completed-before-interrupt, failed-before-interrupt, no-run-at-interrupt), allowed by reconcile-interrupt without recording turn-interrupted. Until C-5 lands, those stops stay uncertain (§3 rule 5). C-5 moves to CHAT-04 with its own contract review (lead decision 30), so CHAT-03 ships that interim rule only.

Deviation V-1, accepted with limits (lead decision 25). CHAT-01 line 145 says an exact retry after reconnect returns the existing receipt. CHAT-01C line 212 says a reconnect or restart retry returns the original result, for recovery requests. CHAT-03 has no durable receipts, so an exact retry across a controller restart refuses stale-incarnation (§1). That is safe, because nothing is replayed, but it departs from the contract. It is recorded here the way the CHAT-02 newer deviation was. Sage accepted it for CHAT-03 with three limits, which the brief carries:

  • the client shows stale-incarnation as "outcome unknown, check the transcript" and never resends on its own (§1);
  • a fixture proves that no retry after a restart reaches the engine (H21, H23);
  • V-1 closes only when CHAT-04's durable receipts restore the contract behavior. CHAT-04's brief lists that as a required item.

Not contract changes:

  • The writer-claim storage record (§2) is internal and never crosses the wire. It stores the CHAT-01 binding fields it needs. If one is missing, that becomes a C item, not an edit.
  • New refusal and reason names (busy, stale-incarnation, engine-pin-mismatch, live-session-refused, foreign-host, transport-unknown, handled-without-run, ack-without-start, interrupted, no-turn, unsealed-engine, run-overlap) are bounded draft IDs, like the existing names CHAT-01 line 363 describes. Where CHAT-01 already has a name, it is reused: generation (check.mjs:205) and controller (check.mjs:230). The package README lists them.

What ships

Three increments: I1, I3 and I4. Each is its own candidate and review round, as with the queue-as-data A1/A2 split. An increment starts when everything in its "Needs first" column is met. Lead decision 30 cut I2 (Pi dialogs) and moved I1b (C-5 adoption) to CHAT-04.

Increment Content Needs first
I1 Controller, transport, writer claim, Pi adapter on a fake Pi engine, incremental events, mediated terminal, interrupt, force stop and recovery. Fixtures in §2–§6 and §9, except the approval races H5–H8, which run in I4; the §10 checks. Brief approved; author named
I3 The B1 evidence packet for Claude (§8) I1 approved. Jason's go for any run that calls a model or reads Claude auth.
I4 Claude adapter, catalogue and history (§7, §8), and H5–H8 on Claude permission requests I3 approved, so B1 passed; C-3 if it exists

Build note (Rocko, R3). The no-turn refusal lifts only the fence its own Interrupt set. It never reopens admission that a concurrent force stop, overlap signal or revocation closed. H10 covers those branches.

If B1 does not pass. Sage records B1 as not passed when any of these happens:

  • Jason declines the go for the I3 recordings;
  • Filbert or Rocko refuses the B1 packet, and no path to a pass exists on the pinned version;
  • Jason says to close CHAT-03 without Claude.

CHAT-03 then closes without I4. Claude keeps refusing unsupported-harness, and CHAT-06's both-harness gate stays blocked with CHAT-03 as its owner. The outcome is recorded, not a silent narrowing.

1. Controller and transport

  • One controller process per execution owns the engine's stdin. Pi runs in --mode rpc with no TUI, so a mediated engine has no native composer or terminal left to fence. That is how CHAT-03 meets plan lines 133–134: a losing native TUI can't stay writable because none exists.

  • Engine output is split on LF only (rpc.md lines 30–38). Node readline is not used, because it also splits on U+2028 and U+2029. The ledger refused a live line for that class of bug (lead decisions item 18). Fixture E2 covers it.

  • Clients connect over a Unix socket in a 0700 directory. Every connection starts as an observer. The browser does not reach the socket in CHAT-03; the WebUI side is CHAT-05.

  • Actor. There is one actor, local-operator, as in CHAT-02. The socket directory's mode is the only boundary, and every process with the same uid passes it. That includes agent seats, which run as the same user. So CHAT-03 cannot stop a same-uid agent from taking control of another conversation. See Limit 1. This blocks any live use until a local owner-channel design passes review (plan lines 181–185; B4).

  • No broker queue. CHAT-04 owns durable queues (plan line 216). In CHAT-03, a prompt sent while the engine is busy is refused with busy, and the text stays in the client. The controller never sends streamingBehavior, steer or follow_up, so mediated input never enters Pi's native queues (rpc.md lines 56–65 for streamingBehavior, 78–104 for steer and follow_up). Tradeoff: there are no queued follow-ups (Q6) until CHAT-04, but I1 has no native-queue ambiguity of its own making.

  • Admission and dispatch both recheck, under one dispatch lock: binding state, open admission, the controlling connection and its generation, the execution incarnation, the controller incarnation token and the text policy. CHAT-01 line 133 requires serialized dispatch, lines 153–160 the dispatch recheck of controller, fence and policy, and lines 181–183 the text policy. The single lock is this brief's way of meeting them.

  • Pending dispatch slot. There is one slot. Under the dispatch lock, the controller reserves it before writing a prompt. The engine counts as busy while the slot is held, and not only from agent_start. The slot is released only by reconciliation: an error response; or an ack, then the agent_settled of its run with no overlap signal (§3); or an ack, then a get_state showing no run with no agent_start or agent_settled in between (handled without a run, §3); or a refusal at the dispatch recheck before the write. On an overlap signal the slot stays held, the binding goes to uncertain, and only force stop moves it on. A second prompt sent while the slot is held refuses busy. This keeps "at most one unstarted item" true, which R3-1 depends on.

  • Write outcomes. The controller records three outcomes:

    • written: the whole line was accepted by the pipe;
    • acknowledged: Pi's prompt response arrived;
    • unknown: the write was partial, returned an error (EPIPE), the controller died before recording the result, or the line was written but no response came within the bound.

    A write accepted by the pipe is not native consumption. An unknown outcome always comes before working, since no ack has arrived. After an unknown outcome, the pipe counts as poisoned, because a partial line would join the next write. The controller never writes to it again. The receipt becomes delivery-unknown with reason transport-unknown, admission closes, and the binding goes to uncertain. Only force stop moves it on. Nothing is retried automatically and nothing is called unsent.

  • Dedup key and incarnation token. The key is the actor, conversation and client request ID (CHAT-01 line 144). The index lives only as long as one controller incarnation. Each controller start mints a random incarnation token, and observe returns it. Every command carries it. A command with an old token refuses stale-incarnation, whatever its request ID or text. So after a restart, an old retry is refused, even though the empty index can't recognize its ID. The client library marks requests pending under the old token as outcome-unknown, and shows a stale-incarnation refusal as "outcome unknown, check the transcript". It never resubmits them under the new token. Durable receipts are CHAT-04, which must restore the contract behavior (deviation V-1, lead decision 25).

  • Text only. Images and files are CHAT-04.

2. Writer-claim record (D1)

CHAT-01 lines 336–343 allow at most one non-stopped binding per conversation or native session identity, and they defer the record. RUNTIME.md §3 item 3 adds one claim per (agent, project, workspace). The record:

  • Keys and claim ID. Each claim has two keys, and both must be held: the seat tuple (seat, project, workspace) and the native session identity (the Pi header ID or the Claude session UUID). A claim ID, random and minted at acquisition, ties them together. Every revision on either key names it. Keys are always taken seat first, then session, and released in the reverse order. A contender reads both keys before publishing. If the session key turns out held after it has published on the seat key, it follows that revision with stopped and a no-unit proof reference (it never spawned). A key counts as held when its highest revision is anything other than stopped with a proof reference. The pair is free only when both keys are free.

  • Pair state. Neither key's chain is authoritative alone. The pair's state is the more conservative of the two, in this order: uncertain, stopping, active, reserved, stopped. Two keys naming different claim IDs are both held, and acquisition refuses unsafe-replacement until each is resolved under its own claim ID.

  • Atomic publication. Each key is a directory under the claim root, and each revision is a numbered file in it. To publish, the controller writes the complete revision to a temporary file in the same directory, fsyncs it, then calls link() to the next revision name and fsyncs the directory. link() fails if the name exists, which makes publication exclusive, and a revision name only ever points at complete contents. A losing link() means another writer published first: the controller re-reads and re-decides. Revisions are never rewritten.

  • Unreadable revision. If the highest revision exists but won't parse, for example after disk damage, the key stays held as uncertain, and acquisition refuses unsafe-replacement. The controller never skips a damaged revision to reuse an older stopped one.

  • Fields. A revision stores:

    • the key, the claim ID and the binding ID;
    • the harness, conversation, and branch and leaf at launch;
    • the config pins: engine version, engine pin (the Pi package integrity of §3, or the Claude binary sha256 of §8) and launch argv digest;
    • the host identity (/etc/machine-id) and boot ID;
    • the owning controller: pid, process start time and incarnation token;
    • the intended scope unit name, derived from the claim ID, recorded before anything is spawned;
    • once known, the scope's systemd invocation ID and the engine pid and start time;
    • the controller generation;
    • the state: reserved, active, stopping, uncertain or stopped;
    • the proof reference that allowed stopped. There are three kinds: a cohortProof (§6), a boot proof (§6), and a no-unit observation, which is valid only when no spawn marker was published.
  • Lifecycle. Each step publishes the new state on the seat key, then on the session key:

    1. reserve both keys, recording the intended unit name;
    2. publish a spawn marker on both keys (still reserved), then start the engine inside that scope (§6), then record the invocation ID and engine identity;
    3. publish active on both keys.

    Release publishes stopping, then stopped with the proof, on each key.

  • Owner check. A controller acts on a claim only if it owns the claim, or if the recorded owner is proven gone: a different boot on the same host, or no process with the recorded pid and start time. A paused or slow owner that still exists is never repaired, reclassified or orphaned by a second controller. The second controller refuses already-active.

  • Restart and half-done pairs. A controller that finds a claim whose owner is proven gone does not launch. It completes or classifies the pair under the same claim ID:

    • Host differs from the recorded machine ID. The pair is treated as uncertain and held, as the queue lock treats a foreign host (queue-as-data plan 8.4), and acquisition refuses foreign-host. Nothing is written, and a copied claim root never becomes stopped.
    • Same host, different boot ID. Every process of the recorded boot is gone. The supervisor issues a boot proof (§6), and both keys move to stopped. Tool calls with no recorded end become uncertain effects.
    • Same boot, reserved, keys possibly half-done. The intended unit name is looked up.
      • No spawn marker and no unit: nothing was started under that claim. The reservation moves to stopped with a no-unit observation.
      • Spawn marker and no unit: the engine may have run and exited, and a collected scope is an absent observation (§6), not proof that every process ended. The pair stays uncertain. Only a boot proof moves it on in CHAT-03. Any other rule is a reviewed change. A marker on either key counts, and completing a half-done pair copies the marker to the other key; it never drops it.
      • A unit exists: the pair is uncertain, and only force stop (§6) moves it on.
    • Same boot, engine or unit alive. This is an orphan: the stdin pipe is gone, so no controller can attach. The state is uncertain, and only a confirmed force stop of that recorded scope moves it on.
    • Same boot, keys disagree (for example one stopped, one stopping). The pair stays held. The restart continues the unfinished transition for that claim ID and never starts a new one.
  • Location and live-session guard. The claim root and every session path are constructor arguments. CHAT-03 has no production default. At construction and again at bind time, the controller resolves real paths. It refuses live-session-refused for any session file or claim root that is not inside the explicit fixture root it was given. It also refuses any path under the repository's .pi/state/, ~/.pi, ~/.claude, the configured data root, or a path named in any seat registration. This is what makes CHAT-03 fixture-only in code and not just by convention. The guard comes out only at cutover (CHAT-07), when the live location is reviewed as a data-map change. The claim excludes only controllers that use it: a pi --session run outside Mosaic isn't prevented, and the guard is what keeps CHAT-03 off real sessions. The idle session drift check moves to CHAT-07 (lead decision 30).

  • What never releases a claim. Disconnect never changes a claim. agent_settled, EOF, SIGTERM, idle and an abort acknowledgement never release one. CHAT-00 line 65 says settled must not release a writer claim. CHAT-01 line 325 says SIGTERM, EOF, an abort acknowledgement and idle don't prove death, so none of them can support a stopped revision either.

  • No writes to sessions. The controller never writes a session file and never uses SessionManager.open (CHAT-00 line 67).

# Fixture Expected
W1 Two processes acquire the same pair at once Exactly one claim. The other refuses already-active.
W2 Acquire while a claim is active or reserved already-active
W3 Acquire while stopping, uncertain, or stopped without proof unsafe-replacement
W4 Same session with a different seat tuple, and the reverse Both refuse. A loser that already published on the seat key follows it with stopped (no-unit). It never spawns, and the winner's revisions are untouched.
W5 SIGKILL between every publication barrier of acquire, transition and release: after the temp write, after its fsync, after link(), after the directory fsync, and between the two keys After restart: never two holders and never a lost claim. Every visible revision is complete. A half-done pair is completed or classified under its claim ID, as in "Restart".
W6 Controller killed mid-turn and restarted while the engine is alive uncertain. No launch, and prompts refuse.
W7 Recorded boot ID differs, same machine ID stopped with a boot proof. Open tool calls become uncertain.
W8 Resume after a proven stop with the same pins New claim ID, generation +1, same conversation, branch and leaf
W9 Resume with a changed binary, argv digest, branch or leaf Refused. The claim is unchanged.
W11 Controller writes to session files None. The CHAT-02 F17 check runs over the fixture session directory. The fake engine's own appends are recorded separately and excluded.
W12 A live owner paused with SIGSTOP; a second controller starts The second refuses already-active. The paused owner's revisions are unchanged after it resumes.
W13 Crash after the engine spawns but before active is published Restart finds the reservation and the live unit: uncertain, force stop only. No second spawn.
W14 Crash after reservation, before the spawn marker No unit and no marker: stopped with a no-unit observation. The pair is free.
W20 Crash after the spawn marker; the engine exits and the scope is collected before restart. Run twice: marker on both keys, and marker on the seat key only. No unit, but a marker: uncertain in both runs. The session key gains the marker when the pair is completed. No launch until a boot proof.
W15 Crash between the two keys during release The pair stays held. Restart finishes the release under the same claim ID.
W16 A highest revision that won't parse Held as uncertain, and acquisition refuses. The older stopped revision is not reused.
W17 A claim root copied from a fixture "other host" (different machine ID) foreign-host. Nothing is promoted.
G1 Session path or claim root under .pi/state/, ~/.claude, the data root, or named in a registration live-session-refused at construction
G2 A symlink inside the fixture root pointing at a live session file live-session-refused at bind (real-path check)
G3 A fixture path that is swapped for a live path after construction Refused at bind

3. Pi adapter, incremental events and R3-1

The Pi pin (lead decision 32). Pi 0.85.1 is pinned by the package-lock.json integrity of @earendil-works/pi-coding-agent, sha512-FGRN+OHbWaefBPGaTggAdLjrIHW+s2PzLyglz/5dfLzb9of7uuXMXYC0fJIeZTw+shS32o2cuQ9jF7YSDuL/oQ==. pi runs dist/bundle/cli.js (the package's bin) and its chunks, not the dist/core and dist/extensions files listed below. At launch the controller checks that package-lock.json and npm's installed record (node_modules/.package-lock.json) both name 0.85.1 with that integrity, and otherwise refuses engine-pin-mismatch. That ties the install to the package through npm's record. It isn't a hash of the files on disk.

The files below are the sources this brief cites for line numbers, hashed at f2b9e622, under node_modules/@earendil-works/pi-coding-agent/. Filbert's R3 review found that the bundle matches them on every point the brief relies on.

  • docs/rpc.md (15fcd26bee72777b373fd5f2edd77091a01cadd4de95e48b08422ced0552a28d)
  • dist/modes/rpc/rpc-mode.js (e7e4724aa55c5aac73cf36793653b26736200e5c59d58373990fc31028f86477)
  • dist/modes/rpc/rpc-types.d.ts (e968e5be01dc7ad9615f938ae867ef136fa495f13dcf169942e9f781a299d9eb)
  • dist/core/agent-session.js (fb8a3981c20c8c0bbd42231b1c99a10335fb3858b659056b341954de9cfa467f)
  • dist/core/session-manager.js (ccace64949db25379a43971ecea750c1b7ec6344e1bc31b9d5fe596ac2f1c9f3)
  • dist/core/output-guard.js (e860db94650c57e07582c300983671737bf9e796682193b498f75e3dd72e9024)
  • dist/core/resource-loader.js (8e8a1bc1c5bc9e955f6a2314dd1db02071be56d48b1fea7b8b2cacb4fc9a0628)
  • dist/cli/args.js (bfb311d2c5d919fa4015d6aaa3c5a71a90b90011320e5e40f56b12e448c44dfc)
  • dist/main.js (f0b7e5a8419af8d149ffe367af2992c76ce70b73484c15492bd50787d4f4962a)
  • dist/extensions/index.js (f980647d447657237cb12b189cec903dc093c55f7a5942a57930c620b01420cc) and dist/extensions/llama/index.js (446b17f49d6197de5aaa6548f78da5934e4e5dfc83acdeed7119a2971ce5e8c1)
  • docs/usage.md (588896ba21944ff002d637444edc22698fd24959c59fe25f95010b9707b47d92)
  • node_modules/@earendil-works/pi-agent-core/dist/agent.js (d84351e451b9fef40fe2532c446aca90d26a4be9038b2d77d3d45dd6eab21d41) and dist/agent-loop.js (6732a1c65c09577d2ffcb716b48e4f4673e57e3e333f10ebfce5132d82e4d7a2), version 0.85.1
  • The goal extension, extensions/goal/index.ts (5ccf78ce7e285ce290add9798b34f0e4e4b30fc8a154c006e34d2494e51a06ae). scripts/sync-dev-extensions.sh copies it into .pi/extensions/goal/, which is untracked (.pi/.gitignore line 1), and the five Pi repository seats load that copy through scripts/agent-host-dev.sh line 137. The copy has the same hash. CHAT-03 cites goal as the example of a turn-starting extension and never loads it.

The plan requires reading the Pi docs completely before Pi implementation (lines 163–164). The author does that before any I1 code.

  • The fake engine. A fixture process that speaks the pinned protocol:

    • prompt acceptance and refusal;
    • message and tool events;
    • agent_end arriving before agent_settled, and retries;
    • clear_queue and abort;
    • dialogs;
    • scripted pause points, so a test can land a race at an exact step;
    • the overlap model of N10: preflight awaits between the busy check and the run start, an acked loser whose throw is swallowed and which settles with no agent_start, several agent_start … agent_end pairs in one run, agent-level messages that clear_queue doesn't return, and nextTurn messages that survive a clear. Tests use these to simulate an unsealed extension without loading one.

    Its behavior is checked against the pinned files, not guessed.

  • Real-binary smoke. The pinned pi starts in --mode rpc in a scratch directory, with a scratch HOME and a scratch Pi agent directory, so the default ~/.pi auth can't be found, and no model call. The setup checks that no auth file is reachable before starting the binary. It answers get_state, get_commands, and clear_queue and abort while idle. The recorded exchange shows that the fake's framing matches the binary. If the binary won't start without auth, the author records that, and the packet says plainly that I1 rests on the fake alone.

  • Events. Native events map onto CHAT-01 event records: message-start, text and thinking deltas, tool-start, tool-update and tool-end, message-end with updateMode: replace, and run-settled.

    • Every event carries the execution incarnation and a sequence number for that execution.
    • An unknown native event gets no client event. It is recorded in controller evidence and counted, and the terminal shows the count. It is never dropped silently and never passed through raw.
    • agent_settled maps to run-settled, never to cohort termination.
  • Joining history and the stream. CHAT-01 lines 105–112 require an atomic cut between a history page and the stream; otherwise streaming is advertised as unavailable. Pinned Pi emits message_end before it persists the entry (agent-session.js lines 386–398), so a page read just after the event can miss that message. Unless the author shows a cut from the pinned source, I1 advertises replay as unavailable. A client gets the page, then live events from the moment it subscribes, with a reconcile marker at the seam. Stream events carry no entry ID, so the seam can't be deduplicated by ID; the marker tells the client to reconcile by re-reading the page once the run settles. Fixture E4 covers both cases.

R3-1: dispatched input when a fence lands. R1 assumed a Mosaic prompt could wait in Pi's queue. It can't. R2 replaced that with a second premise, that Pi runs a Mosaic prompt at once or not at all. Filbert's R2 review (agents/filbert/work/chat-03-brief-review-r2-2026-09-26.md, d149e8cc, F1) showed that premise is false too. R3 builds only on what the pinned source shows (dist/core/agent-session.js, hash above):

  • isStreaming is _isAgentRunActive (lines 616–617). _runAgentPrompt sets it at line 773. Its finally clears it and then emits agent_settled, running extension handlers before the RPC event (lines 780–784 and 347–351).
  • prompt() without streamingBehavior throws while isStreaming is true (lines 860–863). The controller never sends streamingBehavior, so a Mosaic prompt never enters Pi's steer or follow-up queues.
  • The check at line 860 isn't atomic with the run start at line 949. Between them, prompt() always awaits emitBeforeAgentStart (line 915). It also awaits _checkCompaction (line 895) once a turn exists, and emitInput (line 843) when an input handler exists. In that window another run can start: an extension's sendCustomMessage with triggerTurn (lines 1120–1121), or an extension prompt through sendUserMessage (lines 1161–1187) that passed its own line-860 check first.
  • The prompt that loses is still acked (preflightResult(true), line 948). Its agent.prompt throws (pi-agent-core dist/agent.js line 228). The RPC handler swallows the throw because preflight succeeded (rpc-mode.js lines 314–317). Its finally then emits agent_settled and clears isStreaming while the other run goes on. Filbert confirmed this on the installed method with a stub receiver. No real engine was run.
  • One prompt's run can hold several agent_start … agent_end pairs before its single agent_settled, because retries and compaction continue it (_handlePostAgentRun, lines 787–810). In this brief, "a run" means one prompt's run, from its first agent_start to its agent_settled.
  • Events carry no prompt or run ID. The only native evidence tied to a Mosaic prompt is its own response: the ack or the error.

So agent_settled, an idle get_state, and the first user message_start after an ack aren't tied to a Mosaic prompt whenever other code can start a turn. The goal extension that repository seats load does that from every agent_settled. Its handler calls sendUserMessage (goal lines 468–483 and 189), and Pi dispatches the call without awaiting it (lines 2020–2021).

clear_queue doesn't see all external input either (Filbert F2). clearQueue() returns only _steeringMessages and _followUpMessages (lines 1195–1203). An extension's sendMessage while streaming goes straight to agent.steer or agent.followUp (lines 1112–1118). clearAllQueues drops it without returning it, and the pending count in get_state (line 1206) misses it too. nextTurn messages (line 1110) survive clear and abort, and attach to the next prompt (lines 910–913).

The rule R3 follows (Sage). A receipt settles only on positive evidence tied to its own prompt. Idle, settled or empty settle nothing on their own. Where pinned Pi can't tell the cases apart, the controller refuses or reports unknown.

Sealed engine, a binding precondition. Pinned Pi gives no evidence that ties a run to a prompt. So CHAT-03 removes every other source of turns and input, and doesn't guess. A binding needs a sealed engine:

  • The controller builds the launch argv. Pi starts with --no-extensions, --no-prompt-templates and --no-themes, and with no --extension argument (lead decision 31; dist/cli/args.js lines 135–140, 167 and 170; docs/usage.md lines 224 and 233–236). With --no-extensions, Pi loads only the paths given on the command line and ignores settings and packages (dist/core/resource-loader.js lines 316–318), so no explicit extension loads. Skills stay allowed: they expand text and start no turn (§4).
  • Pi also always loads its built-in extensions, whatever the flags (dist/main.js line 439). Pi 0.85.1 has one, llama.cpp (dist/extensions/index.js). It registers a provider and a /llama command, and no event handler or input call (llama/index.js lines 37 and 163). It is part of the pinned package (the Pi pin above), so it is pinned with it.
  • Any --extension argument, or an argv without the three --no-* flags, refuses the binding with unsealed-engine. Explicit extensions aren't pinned or reviewed for binding: Rocko's R3 review showed that an extension can import code outside any hashed tree. They come back with goal in CHAT-06 (lead decision 31). The goal extension starts turns, so a session that loads it can't bind in CHAT-03 (Limit 11). Under lead decision 30, a seat bound to the Console runs without goal, and goal continuation is a CHAT-06 item.

Under the seal, the Mosaic prompt in the slot is the only thing that can start a run, so a run that follows its ack is its run. That exclusion is the evidence basis for attributing by order, and the brief names it as the basis. The seal rests on the launch argv and the pinned Pi build, so the controller also watches for signs that it failed.

Overlap signals. Each of these means the seal failed, or Pi behaved outside the pinned model:

  • O1: an agent_start while no slot is held, before the slot's ack, or after the slot's run has settled;
  • O2: an agent_settled while no slot is held, while the slot is held but before its ack, or a second agent_settled for one ack;
  • O3: an agent_settled while an agent_start has no matching agent_end;
  • O4: an agent_settled after an ack, with no agent_start and no failure message between them;
  • O5: a non-empty clear_queue, a non-zero pending count, or a queue_update the controller didn't cause. Mosaic never queues, so under the seal these are always empty. Pi's own clear emits a queue_update before the clear's response (agent-session.js line 1201, rpc-mode.js line 334). A queue_update read after the controller writes clear_queue and before that response, with empty steering and followUp, is the clear's own, and O5 counts it as caused (N25);
  • O6: a second user message_start in one run.

On any overlap signal, admission closes, the binding goes to uncertain with reason run-overlap, and force stop is the way on. The slot's item becomes delivery-unknown with reason run-overlap if it hasn't reached working. If it has, it stays working, shown as "outcome unknown". A stop in progress gets the Unknown outcome. Because a receipt never moves back from working, a seal failure can misattribute a run to the slot: another run's user message_start arrives before any signal, the item goes to working, and the signal that follows can only mark it "outcome unknown" (N8). Some violations give no signal at all: custom messages queued straight into the agent, and nextTurn messages (Limit 7). A spurious settle delayed until after the item's receipt has settled signals only once the receipt is final (N21, Limit 11).

The rule:

  1. Order. Interrupt closes admission, sends clear_queue, then abort. abort would run anything still queued (rpc.md line 158).

  2. Clear timeout or failure. abort is not sent, because it would run whatever is queued. The turnProof records nativeQueue: unknown, the stop outcome is Unknown, the stop is uncertain, and admission stays closed. Force stop stays available. It needs no cooperation from the engine (§6).

  3. What the clear returned. An empty clear gives nativeQueue: cleared, which here means that Pi's steer and follow-up queues returned nothing. It doesn't prove that no external input was removed, because agent-level custom messages are dropped without being returned (F2). The seal is what keeps those out, and the evidence names the seal as that basis, not the clear. A non-empty clear is overlap signal O5. The stop stays uncertain, reconciled isn't published, admission stays closed (CHAT-01 lines 311–312), and force stop is the way on. The removed items go to controller evidence as digests and byte counts. They are never attributed to a request and never resent.

  4. Receipt settlement. The Mosaic item's receipt settles from evidence tied to its own prompt, never from clear_queue and never from the stop. The prompt's own response is tied to it by ID. A run is tied to it only through the seal, and only while no overlap signal has been seen. A late ack that arrives after the fence goes through the same table as an earlier one. An ack alone never means a run.

    Evidence for the item Receipt
    Error response to the prompt, before or after the fence failed, native error kept
    Refused at the dispatch recheck before any write (fence set) dispatch-refused; no engine bytes
    Ack; then get_state says not streaming, with no agent_start or agent_settled between the ack and that reply delivery-unknown, reason handled-without-run
    Ack; then agent_start and a user message_start, with no overlap signal working. At the run's agent_settled, if the observation from the ack to the settle is complete and ordered and holds no overlap signal, the stopReason of the last assistant message_end decides: stop, length or toolUse gives finished; error gives failed; aborted gives failed with reason interrupted, and evidence links the stop ID. Otherwise (no final assistant message_end, another stopReason such as pending or deferred, or a gap), the receipt stays working, shown as "outcome unknown", and the binding goes to uncertain.
    Ack; then a run that ends in a failure message before any user message_start; then agent_settled delivery-unknown, reason ack-without-start
    Any overlap signal before working delivery-unknown, reason run-overlap
    Write outcome unknown, a missing or unparseable line, or no ack, agent_start or get_state reply within the bound Before working: delivery-unknown, reason transport-unknown. Once working: it stays working, shown as "outcome unknown", because a receipt never moves back from working. Either way the pipe is poisoned (§1) and the binding goes to uncertain.
    Ack; then agent_settled after any sequence no row above matches delivery-unknown, reason ack-without-start

    The transport row and the overlap row take precedence over the rest. Otherwise the rows are exclusive. A line lost without a trace can't be detected, and the table doesn't claim to detect it. An unparseable line is detected.

    A preflight still in flight when the fence lands is awaited, bounded, and then classified by this table. If the classification shows a run that is still active, the controller repeats clear, then abort (rule 1), at most three times. If the run is still going after the third, nativeQueue: unknown, the stop is uncertain, and force stop is the way on.

    Why ack-without-start isn't failed. A run that fails before its user message emits a failure assistant message from Pi's run-failure handler (pi-agent-core dist/agent.js lines 349–364). Pi persists a run's user message only in the message_end handler (agent-session.js lines 386–398), so the session gains no user entry. That ordering is inside Pi. Output reaches the controller through one ordered promise chain (rpc-mode.js lines 28–29, output-guard.js line 71), so a received agent_settled means that every earlier event from the process was written first. It doesn't mean that the settle belongs to this prompt (F1). It also doesn't mean that input or before-agent-start extensions had no effect. The controller has no positive evidence that the item had no effect, so the receipt is delivery-unknown, and a failed that invites a resend is never shown. This answers Rocko's R2 note 2 and Filbert's F1 row 4.

    A delivery-unknown item is never relabelled unsent and never resent. The client shows it as "outcome unknown". The actor may copy the text into a new request, and the client doesn't present that as a safe resend.

  5. The stop outcome. The Interrupt stop is classified separately from the receipt. A run is active at a moment if its first agent_start was read before that moment and its agent_settled wasn't. A run is in scope if it was active when the fence was set, or if its first agent_start was read before the last abort was written. The Interrupted, Completed first and Failed on its own rows also need all of this: exactly one run in scope, its agent_settled read, and an observation from the fence to that settle that is complete and ordered and holds no overlap signal. Unknown's conditions are checked first. Otherwise the first row that matches wins.

    Stop outcome Evidence CHAT-01 reconciliation
    Interrupted The run was active when an abort was written, and its last assistant message_end has stopReason: aborted. What caused the abort isn't claimed: an extension's abort in the same window still interrupted the turn. turnState: interrupted. reconciled may follow when the other rule 6 conditions hold (check.mjs line 315).
    Completed first The run's last stopReason is stop, length or toolUse: it finished before the abort took effect None honest. The receipt stays finished, and its effects stay the run's own. The completion is never relabelled as an interruption.
    Failed on its own The run's last stopReason is error None honest
    No run No run is in scope. The slot held a written item that settled failed or handled-without-run, and the observation from the fence to the last abort has no gap and no overlap signal. None honest
    Unknown Anything else: no abort written (rule 2), any overlap signal, more than one run in scope, no final assistant message, any other stopReason, a transport gap, a missing settle, or a run whose first agent_start was read only after the last abort None

    CHAT-01 lets reconcile-interrupt pass only with turnState: interrupted (check.mjs line 315). input-reconciled belongs to revocation, and CHAT-03 doesn't borrow it. So until C-5 gives the other outcomes an honest value, every outcome except Interrupted leaves the stop uncertain: no reconciled, prompts refuse (CHAT-01 lines 311–312), and force stop is the way on. The stop advances to uncertain (check.mjs line 332), and no turnProof is used for reconciliation. Controller evidence records the outcome, so the client can say "the turn had already finished" instead of "interrupted". The author confirms from pi-agent-core that every run ended by an abort emits a final assistant message_end with stopReason: aborted (dist/agent.js line 357, dist/agent-loop.js line 124). Any path that doesn't is classified Unknown.

    Nothing to interrupt. Interrupt sets the fence, takes the dispatch lock, and then checks the slot. The dispatch recheck reads the fence under the same lock immediately before the write (§1), and a write already under way counts as written. A slot that was reserved but not written therefore refuses at the recheck: the item becomes dispatch-refused, the slot is released, and no bytes reach the engine (H9). If, after that, no written item holds the slot and no run is active, Interrupt refuses no-turn. The fence is lifted, admission reopens, and no stop record is created. The refusal has one recorded effect, the dispatch-refused item if there was one, and names it. This narrowing happens on the controller side, like busy. CHAT-01's checker creates a stop for every interrupt it admits (check.mjs line 244). The narrowing leaves only the race with a turn that finishes during the clear and abort exchange. If the controller refuses just before it reads a new run's agent_start, the run keeps going, and the actor can interrupt again or use force stop.

  6. Reopening. reconciled is published and admission reopens only when all of these hold:

    • the stop outcome is Interrupted (rule 5). Once C-5 is adopted, Completed first, Failed on its own and No run qualify too;
    • if the slot held a Mosaic item, its receipt is settled by rule 4 as finished, failed or delivery-unknown (reason handled-without-run or ack-without-start). A transport-unknown or run-overlap receipt, or a poisoned pipe, keeps the stop uncertain;
    • the in-scope run's agent_settled was read (No run has none), and a get_state sent after the last abort shows not streaming;
    • a clear_queue sent after that returns empty;
    • every clear in the sequence was empty (rule 3), and no overlap signal was seen.

    Under the seal, a non-empty post-settle clear is overlap signal O5: the stop stays uncertain and admission stays closed.

  7. What the proof covers. The proof covers Pi's steer and follow-up queues as of the last clear, and the turnProof says so with its observedAt. Under the seal no extension queues input. If the seal fails, three things stay invisible: input queued after the last clear, custom messages queued straight into the agent, and nextTurn messages that attach to the next Mosaic prompt. The first can show up later as an overlap signal; the other two can't (Limit 7).

  8. Late native events. A user message_start that arrives after the fence is recorded as an event. CHAT-01 says only that a receipt is monotonically revised (README line 36) and has no transition table, so the brief fixes the order: admitted < dispatched < acknowledged < working < finished or failed. dispatch-refused is reachable only before dispatched, and delivery-unknown only before working. A receipt never moves to a state claiming the item wasn't sent.

Outside a fence, the slot's item becomes working at the first user message_start after its ack and agent_start, if no overlap signal has been seen. The attribution rests on the seal, not on order alone. The adapter doesn't compare text, because skills and input handlers can change it. The README says that working is attributed through the seal.

The fake engine models the overlap (N10). Fixtures that break the seal don't load a real extension, since the binding would refuse it. The fake simulates the extension's effect instead.

# Fixture Expected
N1 During an active Mosaic run, Interrupt; after the fence is set and before the clear, the fake simulates an unsealed extension that queues a follow-up clear_queue goes before abort, and the follow-up doesn't run. The queue_update and the non-empty clear are O5: run-overlap, stop outcome Unknown, stop uncertain, prompts refuse, and evidence holds the item's digest. The Mosaic receipt stays working, shown as outcome unknown. Queued before the fence, the queue_update is O5 at once: the binding goes uncertain, and CHAT-01 refuses the Interrupt as fenced (check.mjs line 243).
N2 N1 with abort sent first (mutant) The fake runs the external item, so the test fails. This is the ordering guard.
N3 Fence while the Mosaic prompt is in preflight; preflight then errors, and no run exists Receipt failed with the native error. Stop outcome No run: no turnState: interrupted, stop uncertain, prompts refuse until C-5, force stop available.
N4 Fence while the Mosaic prompt is in preflight; the ack arrives after the first abort, and a run starts Classified by rule 4 as a run. Clear, then abort again. The run ends aborted: receipt failed, reason interrupted; stop outcome Interrupted; reconciled after a post-settle empty clear.
N5 Mosaic prompt acked but handled by an input handler the fake simulates, with no run get_state shows no run, and no agent_start or agent_settled arrived. Receipt delivery-unknown, reason handled-without-run. Nothing is resent.
N6 During an active Mosaic run, Interrupt; the fake simulates an unsealed extension that queues between clear_queue and abort The queue_update is O5. abort continues the queued item inside the same run (agent-session.js lines 787–810), so its user message_start is O6. run-overlap, stop outcome Unknown, stop uncertain, admission closed, force stop is the way on.
N7 clear_queue times out during an active run No abort sent. Stop outcome Unknown. nativeQueue: unknown, stop uncertain, admission closed. A confirmed force stop still ends the cohort (K1).
N8 Filbert's order one: during Mosaic preflight (paused at the line-915 await), the fake simulates an extension prompt that starts a run first. The wire shows the Mosaic ack, the other run's agent_start and user message_start, then the losing Mosaic prompt's agent_settled The other run's user message_start arrives before any signal, so the receipt goes to working. The losing settle is O3: the receipt stays working, shown as outcome unknown, never finished or failed. Binding uncertain, admission closed, nothing resent. The fixture pins the misattribution window (Limit 11).
N9 The item's run started before the fence and ends aborted Receipt failed, reason interrupted, not delivery-unknown. Stop outcome Interrupted (turnState: interrupted).
N10 Fake conformance The fake throws on a prompt while streaming and acks before running. It emits agent_settled from a finally, and it can hold several agent_start … agent_end pairs in one run. It pauses at the preflight awaits (lines 843, 895, 915) between the line-860 check and the run start. A colliding prompt is acked, its throw is swallowed, it settles with no agent_start, and it leaves isStreaming false while the other run goes on. clearQueue emits an empty queue_update before its response, doesn't return agent-level custom messages, and leaves nextTurn messages in place. A fake that queues a Mosaic prompt, or can't produce the overlap, fails the test.
N11 Ack; then the run ends in a failure message before any user message_start; agent_settled arrives; the stream is complete and ordered Receipt delivery-unknown, reason ack-without-start, never failed. The session file gains no user entry. The client shows outcome unknown and no resend offer. The same schedule with the failure line unparseable gives delivery-unknown / transport-unknown.
N12 During an active run, the fake simulates an extension that queues after the final empty clear The stop records the clear's observedAt. The queued item's run starts with no slot held: O1, run-overlap, binding uncertain, admission closed. It is not part of the stopped turn's proof.
N13 The fake emits an agent_start with no slot held (a simulated turn-starting extension) O1: binding uncertain, run-overlap, admission closed. A prompt sent afterwards refuses with zero engine bytes.
N14 The run completes normally (stopReason: stop) while clear_queue is in flight; abort then reaches an idle engine; all clears empty Receipt finished, effects kept. Stop outcome Completed first. No turnState: interrupted, no reconciled, stop uncertain, prompts refuse until C-5, force stop available.
N15 Fence while the Mosaic prompt is in preflight; after the first clear and abort, an input handler the fake simulates handles it and Pi acks with no run Receipt delivery-unknown, reason handled-without-run. No second clear or abort is sent. Stop outcome No run: uncertain until C-5.
N16 Interrupt with no slot held and no visible run Refused no-turn. No stop record, no engine bytes. Admission is open afterwards, and the next prompt is admitted.
N17 The run fails on its own (stopReason: error) during the clear and abort exchange Receipt failed. Stop outcome Failed on its own: uncertain until C-5.
N18 The run settles with no final assistant message_end, or a line is lost during the exchange Receipt working, shown as outcome unknown. A lost line after working never moves it back to delivery-unknown. A lost line before working gives delivery-unknown / transport-unknown. Stop outcome Unknown: uncertain.
N19 Filbert's order two: the Mosaic prompt wins, and a simulated extension prompt that lost emits its agent_settled before the Mosaic run's user message_start O3 (or O4 if it lands before the Mosaic agent_start). Receipt delivery-unknown, reason run-overlap, never failed ack-without-start. Binding uncertain.
N20 A simulated extension timer calls sendCustomMessage with triggerTurn during Mosaic preflight As N8. triggerTurn goes to _runAgentPrompt with no preflight (agent-session.js lines 1120–1121), so the extension's run always starts first. Never finished or failed. The same call while the Mosaic run streams is queued straight into the agent (lines 1112–1118) and gives no signal (Limit 7, N22).
N21 The losing prompt's agent_settled is delayed until after the Mosaic receipt settled finished O2 on arrival: binding uncertain, run-overlap, admission closed. The receipt stays finished, since receipts are monotonic, and evidence records the overlap against it (Limit 11).
N22 During an active run, the fake simulates an extension sendMessage queued straight into the agent; Interrupt The clear returns empty and the item is gone. Evidence records agent-level queues as unobservable, with the seal as the basis, and never as "nothing removed". This fixture pins the Limit 7 gap: no signal fires.
N23 The fake simulates a nextTurn message queued before an Interrupt It survives clear and abort and attaches to the next Mosaic prompt. No signal fires. The fixture pins the Limit 7 gap.
N24 Launch with any --extension argument (a local path, an npm or git source), or an argv missing --no-extensions, --no-prompt-templates or --no-themes unsealed-engine at bind, for each case. No engine is started.
N25 An ordinary Interrupt of an active Mosaic run: the fake emits Pi's empty queue_update before each clear_queue response, and the run ends aborted No overlap signal. Stop outcome Interrupted, reconciled after the post-settle empty clear, admission reopens. A non-empty queue_update in the same window is O5.

4. The slash path (carry-forward 3)

There are two separate hazards: text left in a composer gets concatenated with a new message, and Pi interprets certain prefixes.

  • Concatenation. send-message.sh pastes onto whatever the composer holds. The mediated path has no engine-side composer. Each prompt is one JSONL record, and its message field holds the whole text. Nothing left over can prefix it. The mediated terminal's own composer is a local buffer. It is empty at start, cleared after each submit and on control transfer, and it submits only while its connection is the controller.
  • Interpretation. RPC still acts on a leading /. Extension commands run immediately, even while streaming, and skill and template commands expand (rpc.md lines 67–69).
    • Repository seats load one extension command, /goal (extensions/goal/index.ts:230, loaded through the .pi/extensions path at scripts/agent-host-dev.sh:137). Prompt templates are off (line 136). Skills are live: --no-skills stops discovery only, and the ten explicit --skill paths still load (agent-host-dev.sh lines 77–83 and 138; Pi docs/skills.md line 42). So /skill:<name> expands on repository seats.
    • A CHAT-03 binding can't load goal (§3, sealed engine), so /goal has no command behind it there. The only built-in command, /llama, acts only in TUI mode. Admission doesn't rely on either fact.
    • Admission refuses text-policy when the first non-whitespace character is / (CHAT-01 lines 181–183). Dispatch checks again. The rule holds whatever extensions are loaded, and S1 tests it.
    • CHAT-01 lines 368–370 leave !, @ and slashes on later lines open under B3. I1 settles them from the pinned source. The author lists every prefix that Pi 0.85.1 RPC prompt interprets, and the list becomes a fixture. Admission refuses each listed prefix, and any prefix the author can't classify.
  • Board and agent-send for a mediated seat. A mediated engine has no tmux pane. The board already refuses a reply to a registration without a tmux session before the tool runs (packages/control-board/src/serve.mjs:84; test serve.test.mjs:810). CHAT-03 relies on that and changes neither the board nor tools/tmux. No real seat registers as mediated in CHAT-03, so the refusal first applies to a real seat at cutover.
  • Seats still on tmux keep the hazard. The DEFERRED entry stays open. It closes when a seat moves to the mediated path (CHAT-07), or when a CHAT-03I charter covers tools/tmux (CHAT-01 lines 232–238). Neither happens in CHAT-03.
# Fixture Expected
S1 /goal x, and the same with leading spaces or a tab text-policy at admission. Zero bytes reach the fake engine.
S2 Each prefix on the author's list, including /skill:ms-unslop Refused. The test reads the same list the code uses.
S3 /goal on the second line The fixture pins the result from the pinned source: admitted if Pi doesn't interpret later lines, refused otherwise
S4 The terminal composer holds a / from an abandoned edit, then control transfers and returns Composer cleared. The next submit sends only the new text, and the fake engine records the exact bytes.
S5 An observer terminal gets a paste and then Enter, as send-message.sh does Not admitted: controller. Zero engine bytes.
S6 A mediated-shaped registration (no tmux) passed to replyToRow with a recording exec 409, "no tmux session". exec is never called.
S7 A prompt containing ESC, bracketed-paste markers or U+2028 Sent as one JSON string. The fake engine receives the exact text, and no record splits.

5. Control races (Rocko)

Each race uses the fake engine's pause points to land at an exact step. Every race asserts the engine bytes, the receipts and the events. H5–H8 run in I4 on Claude permission requests. Pi dialogs are cut (lead decision 30).

# Race Expected
H1 Two takeovers with the same expected generation One wins, and the generation goes up by 1. The other refuses generation.
H2 The old controller's prompt arrives after a takeover commits Refused. Zero engine bytes.
H3 A takeover commits while a prompt holds the dispatch lock The prompt either completed its write before the commit, and the receipt keeps the old actor, or it is refused. If the write outcome is unknown, the §1 rule applies: delivery-unknown, pipe poisoned, uncertain. Never dispatched under both controllers.
H4 Self-takeover Refused (CHAT-01 line 269)
H5 The old controller answers an approval after a takeover Refused. The decision is answered once, through the new projection.
H6 Duplicate approval answers with the same content One native response
H7 Conflicting approval answers with the same request ID The second is refused
H8 An approval answer after a native timeout or cancel Refused. The state comes from native evidence.
H9 Interrupt racing a prompt's dispatch Fence set before the write: the item is dispatch-refused with no engine bytes, and with no run active the Interrupt then refuses no-turn (no stop record). Fence set after the write began: §3 rules 4–6.
H10 Interrupt and force stop at the same time; again with an overlap signal or a revocation closing admission before the Interrupt finds no slot and refuses no-turn One stop chain. Force stop supersedes (CHAT-01 lines 319–322). The no-turn cleanup lifts only its own fence: admission stays closed under the surviving reason.
H11 The controller disconnects mid-turn Work continues and the claim is unchanged. Control stays with the disconnected connection until an observer takes over explicitly. Nothing happens automatically.
H12 Exact retry of a prompt after reconnecting to the same controller incarnation The same receipt. No second dispatch.
H13 Retry with the same request ID and different text Refused
H14 Late stdout from the old engine after a replacement Dropped by execution incarnation and counted in the evidence. Never rendered.
H15 A revoked connection sends a command Refused, with the revocation fence of CHAT-01 lines 271–280
H16 A second controller process for the same session Refused as in W1 and W2. The first controller is untouched.
H17 A confirmation reused, answered from another connection, or answered after the stop context changed Refused. Confirmations are single-use (CHAT-01 lines 264–268; CHAT-01C lines 92–102).
H18 Two prompts sent before any native output from the first The second refuses busy. One engine write.
H19 A large prompt line under backpressure; the pipe closes (EPIPE) mid-line, or the controller dies mid-write delivery-unknown, reason transport-unknown. Pipe poisoned, no later write, uncertain. No retry.
H20 The line is written completely, but the native acknowledgement is lost when the controller dies After restart: the claim is an orphan (W6). The request's outcome is unknown, and nothing is resent.
H21 Crash after native dispatch, before the client gets its receipt; restart; the client reconnects and retries the exact request with the old token stale-incarnation. No second engine write.
H22 After H21 and a valid recovery, the client sends a new request with the new token Admitted normally
H23 Requests pending when the controller restarts; the client library reconnects and gets the new token The library resends none of them. The fake engine records zero bytes for them. Each shows "outcome unknown, check the transcript".

6. Stop, cohort proof and recovery (Rocko)

CHAT-01 lines 330–333: a self-posted hash is not trust, and a real producer/verifier and complete cohort containment remain B3/B4. CHAT-03 builds the producer and its observation procedure. The verifier stays the fixture's trusted digest registry, as in CHAT-01. So no CHAT-03 proof is live authority. Where the procedure below can't be carried out, the stop ends at uncertain, not stopped.

  • Producer. A supervisor shim in packages/conversation/src/ starts with systemd-run --user --scope, delegated, under the intended unit name recorded in the claim (§2). Inside the scope it moves itself into a supervisor child cgroup and execs the engine in an engine child cgroup. So the engine is contained from its first instruction, with no window before it could fork. The cohort is the engine cgroup. The shim is not a member. It holds the scope open, so systemd can't garbage-collect the cgroup before emptiness is read.

  • Identity and epoch. The cohort reference is the host (/etc/machine-id), the boot ID, the unit name and the scope's systemd invocation ID. The invocation ID is the membership epoch. A unit with the right name but a different invocation ID is a different cohort. The supervisor sends it no signals and reports evidence unavailable.

  • Containment against migration. A same-uid process can move a pid between cgroups in the user's delegated tree, so a member could leave the scope and survive a "complete" kill. The candidate defence: run the engine in a cgroup namespace rooted at engine. This host mounts cgroup2 with nsdelegate and allows unprivileged user namespaces, and with nsdelegate a namespaced process can't migrate pids outside its namespace root. The author must show this with K13. If K13 can't be made to refuse the escape, real cohorts never reach stopped in CHAT-03: K1 then expects uncertain, and the packet says so. A same-uid process outside the cohort that moves members out is Limit 4 and B2, not something CHAT-03 claims to stop.

  • Unavailable is not empty. Emptiness is a readable engine/cgroup.events showing populated 0, for a scope whose invocation ID matches, read by the live shim. A missing or unreadable path, a gone shim, or an invocation-ID mismatch is evidence unavailable, and the stop stays uncertain. The shim stays a member of the scope in its own leaf, so systemd doesn't collect the scope while the shim reads engine. A collected scope or an absent engine cgroup is an absent observation, never an empty one.

  • Proof. stopped needs a cohortProof with every CHAT-01 field (schema cohortProof): stop, authority, conversation, execution, cohort reference and membership epoch, membershipComplete, members with boot, pid and start time and a death time each, observation time, and verification digest. Membership is complete as of the freeze in the kill phase below, when no member can fork. A process that exited before the freeze isn't listed; its effects belong to the effect report, not the death proof. membershipComplete: true is written only when the freeze, enumeration and populated 0 all succeeded. SIGTERM, EOF, an abort acknowledgement, agent_settled and idle never promote a stop to stopped (line 325).

  • Effects. A tool-start with no tool-end at stop time is an uncertain effect. Killing never counts as rollback (line 334). Network and remote effects stay uncertain.

  • Interrupt (Q20). As in §3 rules 1–8: close admission, clear_queue, abort, settle the item's receipt (rule 4), classify the stop outcome (rule 5), a post-settle empty clear, the turnProof, reconciled, and reopen admission under current control (lines 307–313). Only the Interrupted outcome reconciles until C-5. If the turn completed or failed first, or no run was active, or any overlap signal was seen (a non-empty clear included), or the outcome is Unknown, or the pipe is poisoned, the stop stays uncertain, reconciled isn't published, and prompts refuse, as CHAT-01 lines 311–312 require before that transition. Force stop is the way on. With nothing running, Interrupt refuses no-turn. If it hangs or the clear times out, force stop stays available.

  • Force stop (Q8). Needs an answered confirmation. The stop ID, the answered confirmation and the current phase go into the stopping revision before any signal. Then:

    1. check the scope's invocation ID against the claim;
    2. TERM phase: SIGTERM to every member of engine, then a bounded grace period;
    3. kill phase: freeze engine (cgroup.freeze, waiting for frozen 1), enumerate every member with pid and start time, write cgroup.kill, and wait for populated 0;
    4. build the proof.

    It acts only on the selected execution's engine cgroup, and other executions survive. If a kernel feature the author relies on (cgroup.freeze, cgroup.kill, delegation) isn't available, the stop ends at uncertain.

  • Controller death during force stop. A restart (owner proven gone, §2) never assumes TERM or kill happened. It checks the invocation ID. If that matches, it re-runs the escalation from the TERM phase for the same stop record and the same scope. That continues the recorded stop; it doesn't reuse the confirmation for a new one. If identity can't be checked, it sends no signals, and the stop is uncertain.

  • Boot proof. For a claim whose recorded boot differs on the same host, the supervisor issues a cohortProof whose evidence is the boot change. It goes through the same verifier. A different machine ID is foreign-host, never a boot proof.

  • Recover (Q15). Recover reports eligibility only. It needs:

    • a stopped proof;
    • effects reconciled or explicitly uncertain;
    • an exact confirmation;
    • the same pins (lines 336–343).

    Recover publishes a new reserved claim, under a new claim ID, on both keys. The eligibility record names that claim, the stop, the pins, the leaf and the controller incarnation token, and it is single-use. So no one can take the pair between eligibility and launch. Launching is a separate library call made by a launcher. It revalidates the record against the claim and the current leaf, consumes it once, and spawns. No client command launches, and a browser or controller crash can't start an engine.

  • Pi resume. Resume reuses the same session file and leaf. The author confirms from the pinned docs how Pi selects the leaf. Before admission opens, the adapter checks get_state and get_tree and refuses if the engine loaded any other leaf (RUNTIME.md §3 item 4). A wrongly loaded engine stays under the claim and its scope until a force stop proves it stopped.

# Fixture Expected
K1 Force stop an engine whose tool child calls setsid The child is killed. stopped with a proof accepted by the fixture verifier if K13 refuses the escape; otherwise uncertain, recorded.
K2 K1 on the process-group fallback (no scope) uncertain, never stopped
K3 SIGTERM acknowledged while a member is still alive Stays stopping until the kill phase. Never stopped from TERM alone.
K4 Two fake engines; force stop one The other survives, shown by independent observation: its own cgroup is populated and it still answers get_state
K5 Stop during a tool call The effect is uncertain and shown
K6 Recover without proof, without confirmation, or with changed pins Refused
K7 Recover after proof, then launch New claim ID and incarnation, generation +1, same leaf. The cancelled prompt is not replayed.
K8 The engine loads a different leaf on resume Refused before admission opens. The engine stays claimed and contained until force stop proves it stopped.
K9 An interrupt that never settles Stays uncertain. Force stop stays available, and ordinary takeover is refused while fenced.
K10 Controller killed between the TERM and kill phases Restart checks the invocation ID and re-runs from TERM for the same stop. Nothing is recorded as killed that wasn't observed.
K11 Controller killed after the confirmation is recorded, before TERM Same as K10
K12 A member that forks in a loop during enumeration and kill The freeze stops forking. The enumeration is complete, and populated 0 follows cgroup.kill.
K13 A member writes its own pid to another cgroup's cgroup.procs Refused by the namespace, and the kill is complete. If not refused, stopped is unavailable for real cohorts, and K1 expects uncertain.
K14 A unit with the recorded name but a different invocation ID Evidence unavailable. No signals, uncertain.
K15 engine cgroup path missing or unreadable, or the shim gone Evidence unavailable, not empty. uncertain.
K16 Boot proof requested for a claim from a different machine ID foreign-host. No proof.
K17 Two launcher calls with one eligibility record One launch. The other refuses, and no second engine starts.
K18 The leaf changes after eligibility, before launch Launch refused. The reservation stays until released with proof.

7. Claude permission requests (I4)

  • Pi native dialogs are cut from CHAT-03 (lead decision 30). CHAT-01 lines 239–242 and 292 block non-permission forms until CHAT-03D exists, so a Pi extension_ui_request is shown disabled, with a reason, and is never answered.
  • Claude, in I4 after B1. can_use_tool maps to allow-once and deny. Choices that widen permissions are disabled, and there is no invented cancel (lines 294–295). Pending permission requests are re-read from initialize on reconnect, but only if B1 shows that the pinned version supports it.
# Fixture Expected
P3 A Pi confirm, select, input or editor dialog Disabled, with a reason. No response is sent.
P7 A Claude permission request (I4, recorded fixture) Allow-once and deny only. Widening choices disabled.

8. Claude: B1 and the catalogue (carry-forward 2)

What B1 is. docs/plans/chat-00/README.md lines 193–195: prove the exact Claude CLI and protocol combination, capability negotiation, transcript branch selection and the image and file input shapes. SDK and source examples narrow the uncertainty, but they don't certify the installed binary. CHAT-00 line 66 names the gap for history: Claude's persisted branch format and leaf selection.

What proves it (I3). A B1 packet under agents/<author>/work/chat-03/b1/, approved by Filbert and Rocko, with five parts:

  1. Pin. The exact version string and the sha256 of the binary the adapter will run. Today claude --version reports 2.1.283; the plan saw 2.1.269. The adapter checks both at launch and otherwise refuses with engine-pin-mismatch. It runs the engine with auto-update off, and the author cites the setting from the pinned docs.

  2. Sources. The official protocol sources for that exact version, by hash. These are the C-* entries in chat-00/sources.json, re-pinned to that version. C-HEADLESS and C-REF are floating documents (sources.json lines 106–119) and can't be tied to a version. The packet fixes them as fetched copies, each with its sha256 and fetch date, and says plainly that they aren't version-bound.

  3. Recordings. Made from that binary, in a scratch config directory with no repository credentials:

    • initialize and its capability answer, including whether interrupt_cancel_queued_v1 and pending_permission_requests exist;
    • a text prompt and an image prompt;
    • a can_use_tool allow and a deny;
    • an interrupt.

    They are redacted, hash-pinned and used as the I4 fixtures. The plan rules out model calls for discovery (line 166). Any recording that calls a model or reads Claude auth therefore needs Jason's go first, because it spends money and touches credentials. Without that go, I3 stops after parts 1 and 2, and B1 stays open.

  4. Branch selection. Two persisted transcripts from that binary, one of them with a fork (an edit or a rewind). The packet states the leaf rule and shows that --resume continues the leaf the parser picks.

  5. Negative. The parser refuses a transcript from another version with unsupported-harness.

The catalogue after B1 (I4).

  • Rocko's launcher runs Claude with the default config directory (agents/rocko/launch.sh:86, no CLAUDE_CONFIG_DIR). So its transcripts sit in the same ~/.claude/projects/<cwd slug>/ directory as every other Claude session started in this checkout, including T3 threads and Jason's own sessions.
  • The catalogue therefore never lists that directory. It opens only session IDs the repository launcher recorded for that seat (.pi/state/rocko/session-id and launches/*/session-id). Each ID must resolve to exactly one file under the approved root, opened with the CHAT-02 safe-open rules.
  • A launcher receipt is a hint, not membership. CHAT-01 lines 62–64 require an approved project and workspace mapping, and say OS readability, cwd and seat labels don't establish it. The access mapping stays under B4 review.
  • The CHAT-02 guarantees still apply: read-only, no writes (F17), cursors, byte caps, continuation parts and inert rendering.
# Fixture Expected
L1 A Claude transcript at the pinned version Default leaf shown and other branches readable, as in F12
L2 A transcript from another version unsupported-harness
L3 A decoy session file in the same directory, named in no receipt Never listed, never opened
L4 A receipt naming a file outside the root, a symlink or another project Refused, as in F7–F9
L5 Sidechain or subagent entries Kept separate, not merged into the main branch
L6 No-write check F17 over the Claude directory
L7 The binary's sha256 or version differs at launch engine-pin-mismatch. No launch.

The I4 adapter maps Claude stream-json onto CHAT-01 events. The §5 and §9 fixtures then run again against the recordings.

9. Return flow and events

# Fixture Expected
E1 The plan §6 return flow on the fake engine: send, acknowledge, user, toolCall, toolResult, new final answer Shown once in the same conversation, with no refresh, no duplication, and the draft and reading position kept
E2 U+2028 and U+2029 inside JSON strings, and CRLF Each parsed as one record
E3 Multipart final, two blocks, null request correlation, duplicate delivery The CHAT-01 cases at lines 110–112
E4 A page read just after a message_end but before its entry is persisted, then a subscription; again with a gap or a new epoch Replay is advertised unavailable. The client shows a reconcile marker at the seam and re-reads the page after agent_settled. After that, each message appears exactly once. A gap or new epoch also reconciles. If the author shows an atomic cut from the source, overlap is deduplicated instead, and the fixture pins that case.
E5 An unknown native event No client event. Controller evidence records its type and byte count, and the terminal count goes up. No record fails the schema.
E6 A tool result delayed across a pause and a reconnect Reconciled without a manual refresh
E7 The mediated terminal as observer, then as controller Renders the same stream as the library client, and submits only as controller

The browser leg of the return flow is CHAT-05. I1 proves it at the library and the terminal.

10. Checks and mutation pass

  • Suites. These must pass:

    • node --test packages/conversation/tests/;
    • the control-board, webui and seat package tests;
    • node docs/plans/chat-00/check.mjs, chat-01/check.mjs and chat-01c/check.mjs;
    • every scripts/test-*.sh suite at the candidate's base, including test-queue.sh, since queue A1 landed in 34a72af9.
  • Nested runners. A test that starts node --test itself clears NODE_TEST_CONTEXT for that child, or it can hide failures. The DEFERRED entry on nested runners (Filbert's A1 review 6933b885, note N13) was closed at 34a72af9, and new tests keep to the same rule.

  • Isolation. Every fixture runs in temporary directories. No fixture touches live registrations, .pi/state, ~/.claude or a live seat (plan line 272).

  • Mutation pass. Run in a scratch copy, as for the CHAT-02 Console. Each of these mutants must fail at least one test:

    1. dispatch skips the generation recheck;
    2. abort is sent before clear_queue;
    3. a revision is published by rename or overwrite, so an existing revision name can be replaced;
    4. the late-event filter ignores the incarnation;
    5. the text policy checks only the first character, missing leading whitespace;
    6. an acknowledged SIGTERM promotes a stop to stopped;
    7. the process-group fallback can reach stopped;
    8. a confirmation is not consumed on use;
    9. disconnect releases control;
    10. the composer isn't cleared on transfer;
    11. engine output is read with readline;
    12. a revision is published by exclusive create in place of the temp-file-then-link() sequence, so a partial file becomes visible;
    13. a second controller reclassifies a claim whose owner is alive;
    14. the live-session guard skips the real-path check;
    15. the pending slot is released at agent_start instead of at reconciliation;
    16. a pipe is written again after an unknown write outcome;
    17. an ack followed by no run marks the item working or finished instead of delivery-unknown (N5);
    18. admission reopens without the post-settle empty clear;
    19. abort is sent after a clear_queue timeout;
    20. the incarnation token check is skipped;
    21. a missing cgroup path counts as empty;
    22. a unit with a different invocation ID is signalled;
    23. a restart during force stop records the kill phase as done;
    24. an eligibility record can be used twice;
    25. a non-empty clear_queue reaches reconciled;
    26. the pending slot's in-flight preflight is abandoned at the fence instead of awaited, so a late ack's run is never aborted (N4);
    27. the fake engine queues a Mosaic prompt while streaming instead of throwing (N10's conformance check must catch it);
    28. a no-unit observation frees a pair that has a spawn marker (W20);
    29. a settled slot, an idle engine and empty clears are taken as interruption evidence, so turnState: interrupted is written without a final stopReason: aborted (N14 and N15 must catch it);
    30. a receipt whose run ended stop is relabelled failed (reason interrupted) because a stop was in progress (N14);
    31. turnState: input-reconciled is used to reconcile an Interrupt (N3, N14, N15);
    32. Interrupt with no slot held and no visible run creates a stop record or writes to the engine (N16);
    33. a run is attributed to the slot with the overlap checks skipped, so the receipt settles finished or failed after an overlap signal, or the binding stays active (N8, N19, N20);
    34. a settle that doesn't close the slot's own run (no agent_start since the ack, or an agent_start with no agent_end) is taken as the slot's settle (N8, N19);
    35. an agent_start or agent_settled with no slot held is ignored, so admission stays open (N13, N21);
    36. the binding launches without the --no-* flags, or with an --extension argument (N24);
    37. an ack-without-start run settles failed (N11);
    38. an empty clear is recorded as proof that no external input was removed, instead of naming the seal as the basis (N22).

    Mutant 25 (drift) left with the drift check for CHAT-07, and its number isn't reused. I4 adds two more: the pin check is skipped, and the decoy file is listed.

11. Choices in this brief, for review

Choice Tradeoff
No broker queue; a busy engine refuses busy No queued follow-ups until CHAT-04, and no native queueing of mediated input
No streamingBehavior, steer or follow_up, and one pending slot Same as above. On pinned Pi a Mosaic prompt then never enters a native queue (lines 860–863). It can still lose a race in preflight to another run (lines 860–949), which the seal and the overlap signals handle (§3).
Sealed engine: --no-extensions and no --extension (unsealed-engine; lead decision 31) Without run IDs, only excluding every other source of turns lets a run be attributed to the slot by order. The cost is that no explicit extension loads in CHAT-03, goal included (Limit 11). They come back in CHAT-06. The overlap signals watch for a seal failure.
Overlap signals close admission and make the binding uncertain (run-overlap) One overlap needs a force stop to move on, and never produces a guessed receipt. Some seal failures give no signal (Limits 7 and 11).
The slot settles only on evidence tied to its own prompt, never from clear_queue, idle or a settle alone An ack with no run, and a run that fails before its user message, stay delivery-unknown: no "unsent" state and no safe-resend button
Receipt settlement and the stop outcome are classified separately (§3 rules 4 and 5) A turn that finishes during an Interrupt is reported as finished, not interrupted. The cost is that the stop can't reconcile: until C-5 it stays uncertain and needs a force stop, even though nothing was lost.
Interrupt with nothing running refuses no-turn Narrows the completion race to the clear and abort window. The controller can refuse just before it reads a new run's agent_start; the run keeps going and the actor retries.
Poison the pipe on an unknown write One transport glitch needs a force stop, with no guessing about half-written lines
Live-session guard by path Fixture-only is enforced in code. The guard needs a reviewed change to lift at CHAT-07.
No default claim root Nothing live can use I1 before the CHAT-07 data-map decision
Cohort is an engine child cgroup in a delegated user scope, held by a shim, with a cgroup namespace against migration Depends on systemd-run --user, delegation, cgroup.freeze, cgroup.kill and user namespaces. Without them, or if K13 fails, stops end at uncertain. The verifier is fixture-grade until B3/B4.
One actor, local-operator, guarded only by socket-directory mode Same-uid agents are not kept out (Limit 1)
Dedup lasts one controller incarnation, fenced by a token (deviation V-1, accepted with limits in lead decision 25) A retry across a restart refuses stale-incarnation instead of returning the receipt. The client shows "outcome unknown, check the transcript" and never resends. CHAT-04's durable receipts must restore the contract behavior.
Three increments, with I3/I4 gated on B1 CHAT-03 can close Pi-only, with that outcome recorded
Claude catalogue from launcher receipts only A Claude session started outside the launcher never appears

Carry-forwards

# Sage's item Where Acceptance evidence
1 D1 writer-claim record and R3-1 (lead decisions item 8) §1, §2, §3 W1–W9, W11–W17, W20, G1–G3, N1–N25 and H18–H23 pass. Mutants 2, 3, 12–20 and 26–39 fail tests. R3-1 binds only sealed engines, settles each receipt only on evidence tied to its own prompt (§3 rule 4), and classifies the stop separately (rule 5). Any overlap signal closes admission. Only an interrupted turn reconciles. The other stop outcomes leave the stop uncertain until C-5 lands in CHAT-04 (lead decision 30).
2 The Claude catalogue after B1 §8; increments I3 and I4 The B1 packet (parts 1–5), approved by Filbert and Rocko. L1–L7 pass.
3 The board send that can become a Pi slash command §4 S1–S7 pass. Mutants 5 and 10 fail tests. The DEFERRED entry stays open for seats still on tmux.
4 The CHAT-01/01C contracts, by path and hash "Contracts implemented" The hash table matches sha256sum at approval. C-3 is a separate reviewed item if B1 needs it, and C-5 moves to CHAT-04. C-1, C-2 and C-4 are cut (lead decision 30). Deviation V-1 is accepted with limits (lead decision 25): H21 and H23 pass, the client never resends, and V-1's closure is a required CHAT-04 item (lead decision 25). CHAT-04's brief must list it.
5 Path authority "Files owned" The overlap table. Every changed path in each candidate is inside "Files owned".
6 The source author "Owner and reviewer" <slot>, named by Darkwing after this brief is approved

Limits

  1. Same-uid control. Any process running as the same user can connect to the socket and take control. Agent seats run as that user, so an agent could take over another conversation. CHAT-03 is fixture-only, so no live conversation is exposed. Live use waits for a reviewed local owner-channel design (B4).
  2. Fake engines. They model the pinned docs. The real-binary smoke covers framing only. I1 makes no model call, so real-engine behavior is unproven until CHAT-06.
  3. B2 untouched. Tool isolation and trust and settings parity (CHAT-00 lines 196–197) are untouched. Any real engine for a real seat waits on B2.
  4. Containment. The cgroup covers local processes. Network and remote effects stay uncertain. A same-uid process outside the cohort can still move members out. The verifier is the fixture registry. No CHAT-03 stop proof is live authority until B3/B4 and B2.
  5. Sole writer. CHAT-03 has no session drift check; it moves to CHAT-07 (lead decision 30). The claim binds only controllers that use it. Real sessions stay off-limits through the live-session guard.
  6. Durability. W5 kills processes at each barrier. Real power loss isn't tested; the fsync-then-link() order is the stated basis.
  7. What the clear can't see. Under the seal no extension queues input, so the stop proof rests on the seal. If the seal fails, three things escape it. Input queued after the final empty clear is outside the proof (§3 rule 7), and it shows up only if it starts a run (O1) or queues visibly (O5). Custom messages an extension queues straight into the agent while it streams (agent-session.js lines 1112–1118) are dropped by the clear without being returned (lines 1195–1203) and aren't counted in get_state (line 1206). nextTurn messages (line 1110) survive clear and abort and attach to the next Mosaic prompt (lines 910–913). The last two give no signal at all (N22, N23).
  8. Takeover has nothing to recover. With no broker queue, a takeover has no queued input to turn into drafts. Q13 is CHAT-04's.
  9. Interrupt racing completion. A turn can finish, fail, or never start while the clear and abort are in flight. The receipt reports what happened, but until C-5 the stop can't reconcile and needs a force stop. Sessions with short turns will see this more often.
  10. Restart dedup (V-1). Within one incarnation a retry returns the receipt. Across a restart it refuses stale-incarnation, and the actor has to check the transcript. CHAT-04's durable receipts close this.
  11. Sealed engines only. Pinned Pi ties no run to a prompt, so CHAT-03 attributes runs by order under the seal. A session that loads any explicit extension, goal included, can't bind (unsealed-engine, lead decision 31). The five Pi repository seats (Darkwing, Dewey, Filbert, Researcher, Sage) load goal through scripts/agent-host-dev.sh line 137, so none of them can bind with goal loaded. Under lead decision 30, a seat bound to the Console runs without goal, and goal continuation is a CHAT-06 item. The seal rests on the launch argv and the pinned Pi build. A failure there that the overlap signals don't catch goes undetected if it queues straight into the agent (Limit 7). Other failures are caught late. Another run's user message_start that arrives before any signal puts the item in working, and the signal that follows can only mark it outcome unknown (N8). A spurious settle delayed until after the receipt settled signals only once the receipt is final (N21). CHAT-06 owns explicit and turn-starting extensions, which need run-to-prompt evidence from Pi.

Out of scope

  • Durable queues, drafts, uploads, queue edit and cancel, takeover-to-draft (CHAT-04).
  • Remote control (CHAT-04R) and the chat UI (CHAT-05).
  • The candidate, cutover and any real seat migration (CHAT-06, CHAT-07).
  • Changes to tools/tmux, the board, the seat package, launchers, roles/, contracts/ and the docs/plans/chat-0* contracts.
  • The CHAT-03I fleet-communications charter (CHAT-01 lines 232–238).
  • Authenticated multi-actor control and the local owner-channel design (B4).
  • The Claude adapter and history, unless B1 passes.
  • Filbert's idsDigest note, which stays in agents/dewey/work/chat-02/FOLLOWUPS.md.
  • Explicit extensions, goal continuation among them, and any pinning of them (CHAT-06; Limit 11; lead decisions 30 and 31).
  • Cut by lead decision 30: Pi native dialogs and the CHAT-03D companion (C-2), the unknown-event type (C-4) and the removed-input turnProof (C-1). C-5 moves to CHAT-04, and the idle session drift check to CHAT-07.

Gate

  • Brief.
    • Filbert approves the exact BRIEF.md hash.
    • Rocko's pass on §5 and §6 leaves no blocking finding open.
    • Sage commits and pins the brief. Darkwing then names the author.
  • Each increment.
    • Filbert approves the exact candidate hashes. Rocko passes the §5 and §6 code in I1 and the B1 packet in I3.
    • Every §10 suite is green on an index export, and every mutant is killed.
    • Sage commits. A push needs Jason's word.
  • Contract items. Each is a separate reviewed change to the CHAT-01 files, never folded into an increment.
    • C-3, if B1 shows a gap, is approved before I4 starts.
    • C-5 moves to CHAT-04 (lead decision 30). CHAT-03 ships the interim rule: every stop outcome except Interrupted stays uncertain.
    • Deviation V-1: accepted with limits (lead decision 25). I1 is approved only with H21 and H23 passing and the client's "outcome unknown, check the transcript" display. V-1 stays open until CHAT-04's durable receipts restore the contract behavior; CHAT-04's brief must list that as a required item.
  • I3. Jason's go comes before any model call.
  • CHAT-03 done. All of these:
    • I1 is approved and committed;
    • I4 is approved and committed, or B1 is recorded as not passed (the triggers are in "What ships"), with Claude still refusing and CHAT-06 blocked.
  • No live check. No real seat migrates, so there is none. The live proof belongs to CHAT-07.