107 KiB
CHAT-03 brief: live adapters and mediated terminal (#1507, row 5)
Author: Dewey, 2026-09-26/27. Final text for pinning: R3 (2c5be6b4,
frozen as BRIEF-r3-2c5be6b4.md) with the one edit lead decision 32
orders. It is not a review round (lead decision 27). This is the brief
only. It includes no source, no contract edits and no seat changes. R1
(5dd447f7) is frozen as BRIEF-r1-5dd447f7.md, and R2 (5c5b45a2) as
BRIEF-r2-5c5b45a2.md. §0 lists what changed.
Jason chartered CHAT-03 on 2026-09-26 (docs/plans/2026-09-26_lead-decisions.md
item 22). Plan row: docs/plans/2026-09-13_webui-session-chat.md line 215.
It depends on CHAT-01 (28d4e98a) and implements part of the CHAT-01C
companion (b023841c). The layout follows docs/plans/BRIEF-TEMPLATE.md,
and the depth follows the CHAT-02 brief (agents/dewey/work/chat-02/BRIEF.md,
sha256 636b0fac…). The six carry-forwards
Sage set are mapped to their evidence under "Carry-forwards".
Line numbers and hashes below are at f2b9e622. Between 40a02d2b
(the R2 base) and f2b9e622, no contract, source or pinned Pi file
changed. Queue A1 landed (34a72af9) with docs/plans/BRIEF-TEMPLATE.md,
lead decisions gained items 26 and 27, and the goals review
(docs/plans/2026-09-27_goals-review.md) was added.
Lead decision 27 made R3 the last review round, and its answer to anything pinned Pi can't prove is to refuse or report unknown, never new machinery to prove it. Sage rescoped CHAT-03 against Gate E from R3's section map (lead decision 30), ruled on Rocko's R3 finding (lead decision 31) and ordered this edit (lead decision 32).
0. Changes since R3 (lead decision 32)
Filbert approved R3 on the sections Gate E keeps, with no blocking
finding (agents/filbert/work/chat-03-brief-review-r3-2026-09-27.md,
48447592). Rocko closed his R2 finding and raised one blocking finding
in the seal (agents/rocko/work/chat-03-r3-adversarial-2026-09-27.md,
19e3fcff), which lead decision 31 resolves. This edit does the five
things lead decision 32 lists, and nothing else.
| Item | What changed |
|---|---|
| 1. Lead decision 30 cuts, removed and not reworded | Removed: C-1, C-2 and C-4; increment I2 and §7's Pi dialogs (P1, P2, P4–P6); increment I1b (C-5 moves to CHAT-04, and the interim rule stays); the §2 idle drift check, W10, W18, W19, mutant 25, the drift choice and the session-drift name. H5–H8 stay for Claude in I4. References to removed items were adjusted where they stood: What ships, E5, carry-forwards 1 and 4, §11, Limits 5 and 11, Out of scope and Gate. Two sentences were kept by moving them: the claim's reach, into the live-session guard bullet, and P3's disabled Pi dialog, which is now §7's only Pi rule. |
| 2. Lead decision 31: no explicit extensions | The seal is --no-extensions, --no-prompt-templates, --no-themes and no --extension. The registry, the input-silent review and the tree hashes are removed, not replaced with pinning. N5 and N15 are fake-only. N24 and mutant 37 refuse any --extension. Limit 11 and §11 say explicit extensions return with goal in CHAT-06. |
| 3. Filbert n1: the Pi pin | Pi is pinned by the package-lock.json integrity of @earendil-works/pi-coding-agent 0.85.1, checked against npm's installed record. The listed dist/core and dist/extensions files are now labelled as citation sources, since pi runs dist/bundle. The built-in llama.cpp is pinned with the package. |
4. Filbert n2: the clear's own queue_update |
O5 counts the empty queue_update Pi's clear emits before its response as the controller's own. N10's fake emits it, and N25 proves that an ordinary Interrupt reconciles. |
| 5. Rocko's R3 note | A build note under What ships: the no-turn cleanup lifts only its own fence. H10 covers the force-stop, overlap and revocation branches. |
Filbert's n3 (an aborted with no stop in progress) isn't in lead
decision 32, so it isn't applied.
0a. Changes from R2 to R3 (historical)
R3 answers Rocko's review of R2
(agents/rocko/work/chat-03-r2-adversarial-2026-09-26.md, 07b938fb),
Filbert's review of R2
(agents/filbert/work/chat-03-brief-review-r2-2026-09-26.md, d149e8cc),
Sage's ruling on deviation V-1 (docs/plans/2026-09-26_lead-decisions.md
item 25) and Sage's rule for all three R2 blocking findings: a receipt
settles only on positive evidence tied to its own prompt. Rocko closed R1
findings 1–7, and Filbert closed B2–B6 and his R1 notes. The full
disposition table is in REVIEW-REQUEST.md.
| Finding | What changed |
|---|---|
| Rocko R2 finding 1 (blocking): a settled slot, an idle engine and empty clears don't prove an interrupted turn | §3 splits the old slot rule. Rule 4 settles the item's receipt from its own evidence, and routes a late ack through the same table. Rule 5 classifies the stop: Interrupted, Completed first, Failed on its own, No run, Unknown. A normal completion stays finished. Only Interrupted gives turnState: interrupted and reconciles; input-reconciled isn't borrowed. The rest leave the stop uncertain, admission closed and force stop available, until contract item C-5 (new, separate, reviewed). Rule 6 (reopening) requires Interrupted until C-5. Interrupt with nothing running refuses no-turn. N14 and N15 cover Rocko's completion and handled-ack schedules, and N3 now covers his no-run preflight error. N16–N18 add no-turn, Failed on its own and Unknown. N3, N4 and N9 assert the receipt and the stop. A receipt never moves back from working, and rule 8 states the receipt order. Mutants 30–33. Limit 9. |
| Rocko R2 note 2: N11 wording | The ack-without-start note says Pi's ordering is local to Pi, and that the ordered output chain says nothing about which prompt a settle belongs to. With F1, the outcome is now delivery-unknown, never failed (next row). The pi-agent-core run-failure lines (349–364) are cited, and its two files are pinned by hash. N11 adds the unparseable-line case. |
| Filbert R2 F1 (blocking): an extension run can overlap a Mosaic prompt's preflight, so order proves nothing | §3 R3-1 drops the "at once or not at all" premise and states the verified window (lines 860–949 and their awaits), the acked loser, the swallowed throw, the spurious settle and the multi-pair run. A binding needs a sealed engine: --no-extensions and a registry of hash-pinned, input-silent extensions, otherwise unsealed-engine. The seal is named as the basis for attributing by order. Overlap signals O1–O6 close admission and make the binding uncertain (run-overlap). R2's failed row for a run that fails before its user message becomes delivery-unknown / ack-without-start, and working requires that no overlap signal has been seen. A misattribution to working that a later signal exposes is shown as outcome unknown. Goal can't bind in CHAT-03 (Limit 11, a CHAT-06 item). N8, N10, N13 and N19–N21, N24; mutants 34–37. |
Filbert R2 F2 (blocking): an empty clear_queue doesn't prove that nothing was removed |
Rule 3 narrowed to Pi's steer and follow-up queues, with the seal as the basis. C-1's premise doesn't survive, so C-1 is withdrawn and restated for CHAT-06 ("Contracts implemented"), and I1b is C-5 only. Limit 7 names the agent-level and nextTurn gaps. N22 and N23. |
| Filbert R2 notes n1–n4 | n1: the ordering note cites rpc-mode.js 28–29 and output-guard.js 71, and makes no transport claim. n2: recovery reads the file before get_entries (§2). n3: the drift baseline is taken after load; W18 adds the truncated-line, empty-file and migration cases; "before the first assistant message" applies to new files only. n4: H17 cites CHAT-01C 92–102; §1 cites rpc.md 56–65 and 78–104; agent.js, output-guard.js and the goal extension are pinned. |
| Rocko R2 remark on W20 | A marker on either key counts, and completing a pair copies it. W20 runs with the marker on one key only. |
| Sage, lead decision 25: V-1 accepted with limits | V-1 is accepted. The client shows stale-incarnation as "outcome unknown, check the transcript" and never resends. H23 proves that no request pending across a restart reaches the engine. V-1 closes only when CHAT-04's durable receipts restore the contract behavior, and CHAT-04's brief must list it. Gate, carry-forward 4, §11 and Limit 10 updated. |
| Consistency passes | Two passes over the R3 draft found 22 defects of wording and cross-reference. A third pass over the finished draft found 17, including three where a fixture's expected result didn't follow from the rules (N6, N8, N20). All are fixed. |
| Base | Line numbers and hashes move to f2b9e622. No contract, source or pinned Pi file changed since 40a02d2b. Queue A1 is now committed, so the overlap table names its commit. |
0b. Changes from R1 to R2 (historical)
The rule numbers in this table are R2's. R2's rule 6 is R3's rule 7.
R2 answered three reviews of R1:
- Rocko's adversarial review (
agents/rocko/work/chat-03-r1-adversarial-2026-09-26.md,89752c2b) and his addendum narrowing finding 3 (…-addendum-2026-09-26.md,e5b6bd00); - Filbert's review (
agents/filbert/work/chat-03-brief-review-2026-09-26.md,ec00544e).
Sage's rule for R2: where a guarantee can't be proven on pinned Pi, the
brief refuses or reports uncertainty instead of claiming it. The full
R2 disposition table is in REVIEW-REQUEST-r2-c75ad86f.md.
| Finding | What changed |
|---|---|
| Filbert B1, Rocko 3 and addendum: R3-1 premise | §3 R3-1 rewritten from the source. A mediated prompt can't enter Pi's queues: without streamingBehavior it throws while isStreaming (agent-session.js lines 860–863, 617, 773, 348). clear_queue returns only external input. The residual risks are a fence during preflight and an ack with no run. No text-identity machinery. External input queued after the last clear is stated as a limit (rule 6, Limit 7). N1–N13 rebuilt. |
| Filbert B2, Rocko 6: no entry IDs on the stream | §2: the drift check compares disk entries with the engine's own get_entries list (rpc.md lines 717–745), not with stream IDs. It runs at idle. §3 advertises replay as unavailable. W10, W18 and W19. |
| Filbert B3: schema gaps | C-1 is repurposed as a turnProof field for external items that clear_queue removed. Until it lands, a non-empty clear leaves the stop uncertain. C-4 is an optional event type for unknown native events, with an interim rule. |
| Filbert B4 | §2 live-session guard (live-session-refused, symlinks resolved), G1–G3, mutant 14. CHAT-07 lifts it. |
| Filbert B5: order and gate | Each increment starts when its "Needs first" column is met. Gate adds C-1 adoption or carry, a B1-not-passed trigger and owner, and what happens if C-2 is declined. Increments renamed I1–I4 so they don't collide with CHAT-03D. |
| Filbert B6, Rocko 1: claim gaps | §2 rewritten: one claim ID across both keys, atomic link() publication, the more conservative key wins, restart rules for reserved and half-done pairs, a scope named from the claim ID, and a no-unit path to stopped that a spawn marker closes. W4, W5, W12–W17, W20. The restart dedup is recorded as deviation V-1 for Sage. |
| Rocko 2: pending prompts, partial writes | §1 pending dispatch slot and three write outcomes. A partial or failed write poisons the pipe. H3, H18–H20. |
| Rocko 4: cohort proof | §6 rewritten: shim-held scope, engine child cgroup, invocation-ID epoch, freeze, enumerate, cgroup.kill, unavailable distinct from empty, the K13 migration test, and fixture-grade verification (CHAT-01 lines 330–333). K1–K16. |
| Rocko 5: dedup across restart | §1 incarnation token and stale-incarnation. H21, H22. Deviation V-1. |
| Rocko 7 (non-blocking) | §6 single-use eligibility with a new reservation. K8, K17, K18. |
| Filbert's notes N1–N12 (non-blocking; not the N fixtures) | Skills are live (§4, S2). Refusal names generation and controller reused. Citations corrected. B1 floating sources fixed by hash. Smoke isolated. Suite count by glob. |
CHAT-03: live adapters and mediated terminal
Problem
CHAT-02 made every repository Pi conversation readable. Nothing can drive a
conversation yet except the old paths: the Pi TUI in a tmux pane, and board
Reply or agent-send, which paste into that pane. Those paths have no
controller, no generation fence and no stop proof. At 2026-09-26T20:10:57Z a
leftover / in Pi's composer was concatenated in front of a board reply.
Pi read the result as plain text that time, but the same path could run a
reply as a slash command, because tools/tmux/send-message.sh (lines 45–54) pastes onto
whatever the composer holds and then presses Enter (docs/plans/DEFERRED.md,
"Board send can turn into a Pi slash command").
CHAT-01 defines the records and commands for one controller per
conversation. It defers the writer-claim record (chat-01/README.md:342),
and CHAT-01C defers R3-1 (chat-01c/README.md:217). Lead decisions item 8
moves both here. The Claude catalogue still refuses unsupported-harness
(packages/conversation/src/reader.mjs: constant at line 24, refusals at
lines 51 and 65). Jason ruled that CHAT-03 owns it
after B1. Claude on this host reports 2.1.283, but the plan's evidence named
2.1.269 (plan line 159). The binary changes under any adapter that does not
pin it.
Owner and reviewer
- Source author:
<slot>. The plan has Darkwing name the backend author (line 215; lead decisions item 22). Darkwing is on queue A1/A2, so the author is named after this brief is approved. The brief does not pick one. - Brief author: Dewey.
- Filbert reviews this brief, then the exact candidate of each increment.
- Rocko does an adversarial pass on §5 (control races) and §6 (stop and recovery), in this brief and later in the code of each increment. Rocko also reviews the B1 packet (§8), because it is evidence that a false pass would turn into a wrong adapter.
- No one reviews their own work. If Darkwing names Filbert or Rocko as author, Sage names a replacement for that review.
- Tracking: #1507.
Files owned
The source author may create or change only these paths.
| Path | Contents |
|---|---|
packages/conversation/src/** |
New modules for the controller, claim, supervisor shim, Pi adapter, events, transport and mediated terminal. The Claude parser and adapter come in I4. reader.mjs changes only to lift the Claude refusal for the pinned version. |
packages/conversation/tests/** |
Tests, fake engines, fixture extensions and redacted recordings |
packages/conversation/README.md, packages/conversation/package.json |
Documentation and the new refusal names. No new dependencies. |
agents/<author>/work/chat-03/** |
Review packets, evidence and the B1 packet |
Module names inside src/ are the author's call (plan lines 150–151). Any
path outside this table needs a brief amendment.
Excluded, and why:
tools/tmux/**. Plan line 244 excludes it, and §4 shows the mediated path needs no change there.packages/control-board/**. No board edit is needed (§4, S6). Tests importreplyToRowread-only by relative path, the same way the queue-as-data plan importsscan.mjs.packages/seat/**,scripts/agent-host-dev.sh,agents/*/launch.sh. No seat migrates, so no seat or launcher changes. The launcher handoff belongs to CHAT-07.packages/webui/**. The chat UI is CHAT-05.docs/plans/chat-0*/**. Contracts change only through the C items below.roles/**,contracts/**, root files, auth and~/.mosaic.
No overlap with queue-as-data or the ledger:
| Owner | Paths | Overlap |
|---|---|---|
A1, Darkwing, committed 34a72af9 (agents/darkwing/work/queue-a1/build-manifest.sha256) |
packages/queue/**, scripts/queue-commit.sh, scripts/test-queue.sh, scripts/git-hooks/pre-commit, docs/plans/BRIEF-TEMPLATE.md |
none |
A2, Darkwing (Filbert's queue-as-data plan 282fabbb… §8.1 and lines 130 and 776–777; lead decisions item 20) |
scripts/mosaic dispatch, docs/plans/queue.json, docs/plans/QUEUE.md, scripts/test-{darkwing,rocko}-launch.mjs |
none. The mediated terminal runs as node packages/conversation/src/<cli>.mjs. It is not a scripts/mosaic verb. |
| Piece B, after A | AGENTS.md, agents/*/CONTEXT.md |
none |
| Ledger | packages/ledger/** |
none |
The author may import the process-identity helpers in
packages/discord/src/journal.mjs read-only, as A1 does. That changes
neither package.
Contracts implemented
These are unchanged since the commits named. The sha256 values are at
f2b9e622, and git status shows no local edits.
| Path | sha256 | Commit |
|---|---|---|
docs/plans/chat-00/README.md |
991663e607404c2f022716bd45b40d5b3092d53c3f9f61bbafd455c0659b832f |
370823b3 |
docs/plans/chat-00/sources.json |
1a07ae88de45fd3219eca10acae13637ca75416cb1c598aa9080825bd7598af9 |
370823b3 |
docs/plans/chat-00/fixtures.json |
ad9d4fe94131b2c4bb30691ad846698c5b0818f0935d20a38de73fa3f267bd80 |
370823b3 |
docs/plans/chat-00/check.mjs |
568f004f5a0cfce64d65ff202c420ab491e716919a93989377a7bb9e7b8be622 |
370823b3 |
docs/plans/chat-01/README.md |
61aba7d60f380ff8a135a04a2e851f11c250795ead11cca2e594566849483163 |
28d4e98a |
docs/plans/chat-01/contracts.schema.json |
38382e08c97635f8864e6953b6cf44ec324040397a1aa8fc3f86e279abe36db1 |
28d4e98a |
docs/plans/chat-01/fixtures.json |
403c8ae91963a379de17805380bc1425bc4e96c1ef34094385b703d17cf5ce7d |
28d4e98a |
docs/plans/chat-01/check.mjs |
2e164e4bfa61963bb5dd8639e26e67407e64d19278cc56ed8a4f77e430f1dee5 |
28d4e98a |
docs/plans/chat-01c/README.md |
63d9f9edffc2aa1aa3bc99364b864e8c6f234eedc8ad2f9da231531970d54250 |
b023841c |
docs/plans/chat-01c/contracts.schema.json |
da132f0a02f29281344cc5350396af53c893d7895308bc430f9bf3ca3b3785b9 |
b023841c |
docs/plans/chat-01c/fixtures.json |
00639a219b00b67c5e0ab904163a6d945481f6b0df45a058b32e00f1f1326611 |
b023841c |
docs/plans/chat-01c/check.mjs |
5971ed5f11a32f0dff5720db03d8d4bc4c007a5d2654ba50726bb35ba753eba6 |
b023841c |
Design inputs, which CHAT-03 does not implement in full:
docs/plans/foundation-v1-candidate/RUNTIME.md §§3–5 (b1a2b4d0…) and the
plan itself (48142829…).
What CHAT-03 takes from each contract:
- CHAT-01. The records
binding,connection,clientRequest,request,receipt,event,stop,cohortProof,effectReport,turnProof,nativeDecision,approvalandconfirmation. The commandsobserve,prompt,takeover,acquire-recovery-control,approval,interrupt,force-stop,recover,issue-confirmationandanswer-confirmation. The draft, upload and queue-edit commands belong to CHAT-04. - CHAT-01C. Only the confirmation restoration rules (lines 92–104). The
private-state pages, upload ranges and
recover-refused-draftneed CHAT-04's durable store.
The CHAT-01 text at line 342 stays as published. Lead decisions item 8 records the move.
Contract changes, each a separate reviewed item. Each one lands and is approved before the code that needs it. Sage assigns the author; Filbert and Dewey review, as they did for CHAT-01C.
-
C-3: Claude record fields. Needed only if B1 shows that CHAT-01 lacks a field, for example the capability-negotiation result on
binding. Otherwise C-3 is void. -
C-5: interrupt outcomes other than an interrupted turn (R3).
turnProof.turnStateallowsinterrupted,unknownandinput-reconciled, andreconcile-interruptpasses only withinterrupted(check.mjsline 315). A turn that completed or failed before the abort took effect, or an Interrupt that found no run, has no honest value. C-5 adds values for those (for examplecompleted-before-interrupt,failed-before-interrupt,no-run-at-interrupt), allowed byreconcile-interruptwithout recordingturn-interrupted. Until C-5 lands, those stops stayuncertain(§3 rule 5). C-5 moves to CHAT-04 with its own contract review (lead decision 30), so CHAT-03 ships that interim rule only.
Deviation V-1, accepted with limits (lead decision 25). CHAT-01 line 145 says an exact
retry after reconnect returns the existing receipt. CHAT-01C line 212 says
a reconnect or restart retry returns the original result, for recovery
requests. CHAT-03 has no durable receipts, so an exact retry across a
controller restart refuses stale-incarnation (§1). That is safe, because
nothing is replayed, but it departs from the contract. It is recorded here
the way the CHAT-02 newer deviation was. Sage accepted it for CHAT-03
with three limits, which the brief carries:
- the client shows
stale-incarnationas "outcome unknown, check the transcript" and never resends on its own (§1); - a fixture proves that no retry after a restart reaches the engine (H21, H23);
- V-1 closes only when CHAT-04's durable receipts restore the contract behavior. CHAT-04's brief lists that as a required item.
Not contract changes:
- The writer-claim storage record (§2) is internal and never crosses the
wire. It stores the CHAT-01
bindingfields it needs. If one is missing, that becomes a C item, not an edit. - New refusal and reason names (
busy,stale-incarnation,engine-pin-mismatch,live-session-refused,foreign-host,transport-unknown,handled-without-run,ack-without-start,interrupted,no-turn,unsealed-engine,run-overlap) are bounded draft IDs, like the existing names CHAT-01 line 363 describes. Where CHAT-01 already has a name, it is reused:generation(check.mjs:205) andcontroller(check.mjs:230). The package README lists them.
What ships
Three increments: I1, I3 and I4. Each is its own candidate and review round, as with the queue-as-data A1/A2 split. An increment starts when everything in its "Needs first" column is met. Lead decision 30 cut I2 (Pi dialogs) and moved I1b (C-5 adoption) to CHAT-04.
| Increment | Content | Needs first |
|---|---|---|
| I1 | Controller, transport, writer claim, Pi adapter on a fake Pi engine, incremental events, mediated terminal, interrupt, force stop and recovery. Fixtures in §2–§6 and §9, except the approval races H5–H8, which run in I4; the §10 checks. | Brief approved; author named |
| I3 | The B1 evidence packet for Claude (§8) | I1 approved. Jason's go for any run that calls a model or reads Claude auth. |
| I4 | Claude adapter, catalogue and history (§7, §8), and H5–H8 on Claude permission requests | I3 approved, so B1 passed; C-3 if it exists |
Build note (Rocko, R3). The no-turn refusal lifts only the fence
its own Interrupt set. It never reopens admission that a concurrent force
stop, overlap signal or revocation closed. H10 covers those branches.
If B1 does not pass. Sage records B1 as not passed when any of these happens:
- Jason declines the go for the I3 recordings;
- Filbert or Rocko refuses the B1 packet, and no path to a pass exists on the pinned version;
- Jason says to close CHAT-03 without Claude.
CHAT-03 then closes without I4. Claude keeps refusing
unsupported-harness, and CHAT-06's both-harness gate stays blocked with
CHAT-03 as its owner. The outcome is recorded, not a silent narrowing.
1. Controller and transport
-
One controller process per execution owns the engine's stdin. Pi runs in
--mode rpcwith no TUI, so a mediated engine has no native composer or terminal left to fence. That is how CHAT-03 meets plan lines 133–134: a losing native TUI can't stay writable because none exists. -
Engine output is split on LF only (
rpc.mdlines 30–38). Nodereadlineis not used, because it also splits on U+2028 and U+2029. The ledger refused a live line for that class of bug (lead decisions item 18). Fixture E2 covers it. -
Clients connect over a Unix socket in a 0700 directory. Every connection starts as an observer. The browser does not reach the socket in CHAT-03; the WebUI side is CHAT-05.
-
Actor. There is one actor,
local-operator, as in CHAT-02. The socket directory's mode is the only boundary, and every process with the same uid passes it. That includes agent seats, which run as the same user. So CHAT-03 cannot stop a same-uid agent from taking control of another conversation. See Limit 1. This blocks any live use until a local owner-channel design passes review (plan lines 181–185; B4). -
No broker queue. CHAT-04 owns durable queues (plan line 216). In CHAT-03, a prompt sent while the engine is busy is refused with
busy, and the text stays in the client. The controller never sendsstreamingBehavior,steerorfollow_up, so mediated input never enters Pi's native queues (rpc.mdlines 56–65 forstreamingBehavior, 78–104 forsteerandfollow_up). Tradeoff: there are no queued follow-ups (Q6) until CHAT-04, but I1 has no native-queue ambiguity of its own making. -
Admission and dispatch both recheck, under one dispatch lock: binding state, open admission, the controlling connection and its generation, the execution incarnation, the controller incarnation token and the text policy. CHAT-01 line 133 requires serialized dispatch, lines 153–160 the dispatch recheck of controller, fence and policy, and lines 181–183 the text policy. The single lock is this brief's way of meeting them.
-
Pending dispatch slot. There is one slot. Under the dispatch lock, the controller reserves it before writing a prompt. The engine counts as busy while the slot is held, and not only from
agent_start. The slot is released only by reconciliation: an error response; or an ack, then theagent_settledof its run with no overlap signal (§3); or an ack, then aget_stateshowing no run with noagent_startoragent_settledin between (handled without a run, §3); or a refusal at the dispatch recheck before the write. On an overlap signal the slot stays held, the binding goes touncertain, and only force stop moves it on. A second prompt sent while the slot is held refusesbusy. This keeps "at most one unstarted item" true, which R3-1 depends on. -
Write outcomes. The controller records three outcomes:
- written: the whole line was accepted by the pipe;
- acknowledged: Pi's
promptresponse arrived; - unknown: the write was partial, returned an error (EPIPE), the controller died before recording the result, or the line was written but no response came within the bound.
A write accepted by the pipe is not native consumption. An unknown outcome always comes before
working, since no ack has arrived. After an unknown outcome, the pipe counts as poisoned, because a partial line would join the next write. The controller never writes to it again. The receipt becomesdelivery-unknownwith reasontransport-unknown, admission closes, and the binding goes touncertain. Only force stop moves it on. Nothing is retried automatically and nothing is called unsent. -
Dedup key and incarnation token. The key is the actor, conversation and client request ID (CHAT-01 line 144). The index lives only as long as one controller incarnation. Each controller start mints a random incarnation token, and
observereturns it. Every command carries it. A command with an old token refusesstale-incarnation, whatever its request ID or text. So after a restart, an old retry is refused, even though the empty index can't recognize its ID. The client library marks requests pending under the old token as outcome-unknown, and shows astale-incarnationrefusal as "outcome unknown, check the transcript". It never resubmits them under the new token. Durable receipts are CHAT-04, which must restore the contract behavior (deviation V-1, lead decision 25). -
Text only. Images and files are CHAT-04.
2. Writer-claim record (D1)
CHAT-01 lines 336–343 allow at most one non-stopped binding per conversation or native session identity, and they defer the record. RUNTIME.md §3 item 3 adds one claim per (agent, project, workspace). The record:
-
Keys and claim ID. Each claim has two keys, and both must be held: the seat tuple (seat, project, workspace) and the native session identity (the Pi header ID or the Claude session UUID). A claim ID, random and minted at acquisition, ties them together. Every revision on either key names it. Keys are always taken seat first, then session, and released in the reverse order. A contender reads both keys before publishing. If the session key turns out held after it has published on the seat key, it follows that revision with
stoppedand a no-unit proof reference (it never spawned). A key counts as held when its highest revision is anything other thanstoppedwith a proof reference. The pair is free only when both keys are free. -
Pair state. Neither key's chain is authoritative alone. The pair's state is the more conservative of the two, in this order:
uncertain,stopping,active,reserved,stopped. Two keys naming different claim IDs are both held, and acquisition refusesunsafe-replacementuntil each is resolved under its own claim ID. -
Atomic publication. Each key is a directory under the claim root, and each revision is a numbered file in it. To publish, the controller writes the complete revision to a temporary file in the same directory, fsyncs it, then calls
link()to the next revision name and fsyncs the directory.link()fails if the name exists, which makes publication exclusive, and a revision name only ever points at complete contents. A losinglink()means another writer published first: the controller re-reads and re-decides. Revisions are never rewritten. -
Unreadable revision. If the highest revision exists but won't parse, for example after disk damage, the key stays held as
uncertain, and acquisition refusesunsafe-replacement. The controller never skips a damaged revision to reuse an olderstoppedone. -
Fields. A revision stores:
- the key, the claim ID and the binding ID;
- the harness, conversation, and branch and leaf at launch;
- the config pins: engine version, engine pin (the Pi package integrity of §3, or the Claude binary sha256 of §8) and launch argv digest;
- the host identity (
/etc/machine-id) and boot ID; - the owning controller: pid, process start time and incarnation token;
- the intended scope unit name, derived from the claim ID, recorded before anything is spawned;
- once known, the scope's systemd invocation ID and the engine pid and start time;
- the controller generation;
- the state:
reserved,active,stopping,uncertainorstopped; - the proof reference that allowed
stopped. There are three kinds: a cohortProof (§6), a boot proof (§6), and a no-unit observation, which is valid only when no spawn marker was published.
-
Lifecycle. Each step publishes the new state on the seat key, then on the session key:
- reserve both keys, recording the intended unit name;
- publish a spawn marker on both keys (still
reserved), then start the engine inside that scope (§6), then record the invocation ID and engine identity; - publish
activeon both keys.
Release publishes
stopping, thenstoppedwith the proof, on each key. -
Owner check. A controller acts on a claim only if it owns the claim, or if the recorded owner is proven gone: a different boot on the same host, or no process with the recorded pid and start time. A paused or slow owner that still exists is never repaired, reclassified or orphaned by a second controller. The second controller refuses
already-active. -
Restart and half-done pairs. A controller that finds a claim whose owner is proven gone does not launch. It completes or classifies the pair under the same claim ID:
- Host differs from the recorded machine ID. The pair is treated as
uncertainand held, as the queue lock treats a foreign host (queue-as-data plan 8.4), and acquisition refusesforeign-host. Nothing is written, and a copied claim root never becomesstopped. - Same host, different boot ID. Every process of the recorded boot is
gone. The supervisor issues a boot proof (§6), and both keys move to
stopped. Tool calls with no recorded end becomeuncertaineffects. - Same boot,
reserved, keys possibly half-done. The intended unit name is looked up.- No spawn marker and no unit: nothing was started under that claim.
The reservation moves to
stoppedwith a no-unit observation. - Spawn marker and no unit: the engine may have run and exited, and a
collected scope is an absent observation (§6), not proof that every
process ended. The pair stays
uncertain. Only a boot proof moves it on in CHAT-03. Any other rule is a reviewed change. A marker on either key counts, and completing a half-done pair copies the marker to the other key; it never drops it. - A unit exists: the pair is
uncertain, and only force stop (§6) moves it on.
- No spawn marker and no unit: nothing was started under that claim.
The reservation moves to
- Same boot, engine or unit alive. This is an orphan: the stdin pipe
is gone, so no controller can attach. The state is
uncertain, and only a confirmed force stop of that recorded scope moves it on. - Same boot, keys disagree (for example one
stopped, onestopping). The pair stays held. The restart continues the unfinished transition for that claim ID and never starts a new one.
- Host differs from the recorded machine ID. The pair is treated as
-
Location and live-session guard. The claim root and every session path are constructor arguments. CHAT-03 has no production default. At construction and again at bind time, the controller resolves real paths. It refuses
live-session-refusedfor any session file or claim root that is not inside the explicit fixture root it was given. It also refuses any path under the repository's.pi/state/,~/.pi,~/.claude, the configured data root, or a path named in any seat registration. This is what makes CHAT-03 fixture-only in code and not just by convention. The guard comes out only at cutover (CHAT-07), when the live location is reviewed as a data-map change. The claim excludes only controllers that use it: api --sessionrun outside Mosaic isn't prevented, and the guard is what keeps CHAT-03 off real sessions. The idle session drift check moves to CHAT-07 (lead decision 30). -
What never releases a claim. Disconnect never changes a claim.
agent_settled, EOF, SIGTERM, idle and an abort acknowledgement never release one. CHAT-00 line 65 says settled must not release a writer claim. CHAT-01 line 325 says SIGTERM, EOF, an abort acknowledgement and idle don't prove death, so none of them can support astoppedrevision either. -
No writes to sessions. The controller never writes a session file and never uses
SessionManager.open(CHAT-00 line 67).
| # | Fixture | Expected |
|---|---|---|
| W1 | Two processes acquire the same pair at once | Exactly one claim. The other refuses already-active. |
| W2 | Acquire while a claim is active or reserved |
already-active |
| W3 | Acquire while stopping, uncertain, or stopped without proof |
unsafe-replacement |
| W4 | Same session with a different seat tuple, and the reverse | Both refuse. A loser that already published on the seat key follows it with stopped (no-unit). It never spawns, and the winner's revisions are untouched. |
| W5 | SIGKILL between every publication barrier of acquire, transition and release: after the temp write, after its fsync, after link(), after the directory fsync, and between the two keys |
After restart: never two holders and never a lost claim. Every visible revision is complete. A half-done pair is completed or classified under its claim ID, as in "Restart". |
| W6 | Controller killed mid-turn and restarted while the engine is alive | uncertain. No launch, and prompts refuse. |
| W7 | Recorded boot ID differs, same machine ID | stopped with a boot proof. Open tool calls become uncertain. |
| W8 | Resume after a proven stop with the same pins | New claim ID, generation +1, same conversation, branch and leaf |
| W9 | Resume with a changed binary, argv digest, branch or leaf | Refused. The claim is unchanged. |
| W11 | Controller writes to session files | None. The CHAT-02 F17 check runs over the fixture session directory. The fake engine's own appends are recorded separately and excluded. |
| W12 | A live owner paused with SIGSTOP; a second controller starts | The second refuses already-active. The paused owner's revisions are unchanged after it resumes. |
| W13 | Crash after the engine spawns but before active is published |
Restart finds the reservation and the live unit: uncertain, force stop only. No second spawn. |
| W14 | Crash after reservation, before the spawn marker | No unit and no marker: stopped with a no-unit observation. The pair is free. |
| W20 | Crash after the spawn marker; the engine exits and the scope is collected before restart. Run twice: marker on both keys, and marker on the seat key only. | No unit, but a marker: uncertain in both runs. The session key gains the marker when the pair is completed. No launch until a boot proof. |
| W15 | Crash between the two keys during release | The pair stays held. Restart finishes the release under the same claim ID. |
| W16 | A highest revision that won't parse | Held as uncertain, and acquisition refuses. The older stopped revision is not reused. |
| W17 | A claim root copied from a fixture "other host" (different machine ID) | foreign-host. Nothing is promoted. |
| G1 | Session path or claim root under .pi/state/, ~/.claude, the data root, or named in a registration |
live-session-refused at construction |
| G2 | A symlink inside the fixture root pointing at a live session file | live-session-refused at bind (real-path check) |
| G3 | A fixture path that is swapped for a live path after construction | Refused at bind |
3. Pi adapter, incremental events and R3-1
The Pi pin (lead decision 32). Pi 0.85.1 is pinned by the
package-lock.json integrity of @earendil-works/pi-coding-agent,
sha512-FGRN+OHbWaefBPGaTggAdLjrIHW+s2PzLyglz/5dfLzb9of7uuXMXYC0fJIeZTw+shS32o2cuQ9jF7YSDuL/oQ==.
pi runs dist/bundle/cli.js (the package's bin) and its chunks, not
the dist/core and dist/extensions files listed below. At launch the
controller checks that package-lock.json and npm's installed record
(node_modules/.package-lock.json) both name 0.85.1 with that integrity,
and otherwise refuses engine-pin-mismatch. That ties the install to the
package through npm's record. It isn't a hash of the files on disk.
The files below are the sources this brief cites for line numbers, hashed
at f2b9e622, under node_modules/@earendil-works/pi-coding-agent/.
Filbert's R3 review found that the bundle matches them on every point the
brief relies on.
docs/rpc.md(15fcd26bee72777b373fd5f2edd77091a01cadd4de95e48b08422ced0552a28d)dist/modes/rpc/rpc-mode.js(e7e4724aa55c5aac73cf36793653b26736200e5c59d58373990fc31028f86477)dist/modes/rpc/rpc-types.d.ts(e968e5be01dc7ad9615f938ae867ef136fa495f13dcf169942e9f781a299d9eb)dist/core/agent-session.js(fb8a3981c20c8c0bbd42231b1c99a10335fb3858b659056b341954de9cfa467f)dist/core/session-manager.js(ccace64949db25379a43971ecea750c1b7ec6344e1bc31b9d5fe596ac2f1c9f3)dist/core/output-guard.js(e860db94650c57e07582c300983671737bf9e796682193b498f75e3dd72e9024)dist/core/resource-loader.js(8e8a1bc1c5bc9e955f6a2314dd1db02071be56d48b1fea7b8b2cacb4fc9a0628)dist/cli/args.js(bfb311d2c5d919fa4015d6aaa3c5a71a90b90011320e5e40f56b12e448c44dfc)dist/main.js(f0b7e5a8419af8d149ffe367af2992c76ce70b73484c15492bd50787d4f4962a)dist/extensions/index.js(f980647d447657237cb12b189cec903dc093c55f7a5942a57930c620b01420cc) anddist/extensions/llama/index.js(446b17f49d6197de5aaa6548f78da5934e4e5dfc83acdeed7119a2971ce5e8c1)docs/usage.md(588896ba21944ff002d637444edc22698fd24959c59fe25f95010b9707b47d92)node_modules/@earendil-works/pi-agent-core/dist/agent.js(d84351e451b9fef40fe2532c446aca90d26a4be9038b2d77d3d45dd6eab21d41) anddist/agent-loop.js(6732a1c65c09577d2ffcb716b48e4f4673e57e3e333f10ebfce5132d82e4d7a2), version 0.85.1- The goal extension,
extensions/goal/index.ts(5ccf78ce7e285ce290add9798b34f0e4e4b30fc8a154c006e34d2494e51a06ae).scripts/sync-dev-extensions.shcopies it into.pi/extensions/goal/, which is untracked (.pi/.gitignoreline 1), and the five Pi repository seats load that copy throughscripts/agent-host-dev.shline 137. The copy has the same hash. CHAT-03 cites goal as the example of a turn-starting extension and never loads it.
The plan requires reading the Pi docs completely before Pi implementation (lines 163–164). The author does that before any I1 code.
-
The fake engine. A fixture process that speaks the pinned protocol:
- prompt acceptance and refusal;
- message and tool events;
agent_endarriving beforeagent_settled, and retries;clear_queueandabort;- dialogs;
- scripted pause points, so a test can land a race at an exact step;
- the overlap model of N10: preflight awaits between the busy check and
the run start, an acked loser whose throw is swallowed and which
settles with no
agent_start, severalagent_start…agent_endpairs in one run, agent-level messages thatclear_queuedoesn't return, andnextTurnmessages that survive a clear. Tests use these to simulate an unsealed extension without loading one.
Its behavior is checked against the pinned files, not guessed.
-
Real-binary smoke. The pinned
pistarts in--mode rpcin a scratch directory, with a scratchHOMEand a scratch Pi agent directory, so the default~/.piauth can't be found, and no model call. The setup checks that no auth file is reachable before starting the binary. It answersget_state,get_commands, andclear_queueandabortwhile idle. The recorded exchange shows that the fake's framing matches the binary. If the binary won't start without auth, the author records that, and the packet says plainly that I1 rests on the fake alone. -
Events. Native events map onto CHAT-01
eventrecords: message-start, text and thinking deltas, tool-start, tool-update and tool-end, message-end withupdateMode: replace, and run-settled.- Every event carries the execution incarnation and a sequence number for that execution.
- An unknown native event gets no client event. It is recorded in controller evidence and counted, and the terminal shows the count. It is never dropped silently and never passed through raw.
agent_settledmaps to run-settled, never to cohort termination.
-
Joining history and the stream. CHAT-01 lines 105–112 require an atomic cut between a history page and the stream; otherwise streaming is advertised as unavailable. Pinned Pi emits
message_endbefore it persists the entry (agent-session.jslines 386–398), so a page read just after the event can miss that message. Unless the author shows a cut from the pinned source, I1 advertises replay as unavailable. A client gets the page, then live events from the moment it subscribes, with a reconcile marker at the seam. Stream events carry no entry ID, so the seam can't be deduplicated by ID; the marker tells the client to reconcile by re-reading the page once the run settles. Fixture E4 covers both cases.
R3-1: dispatched input when a fence lands. R1 assumed a Mosaic prompt
could wait in Pi's queue. It can't. R2 replaced that with a second premise,
that Pi runs a Mosaic prompt at once or not at all. Filbert's R2 review
(agents/filbert/work/chat-03-brief-review-r2-2026-09-26.md, d149e8cc,
F1) showed that premise is false too. R3 builds only on what the pinned
source shows (dist/core/agent-session.js, hash above):
isStreamingis_isAgentRunActive(lines 616–617)._runAgentPromptsets it at line 773. Itsfinallyclears it and then emitsagent_settled, running extension handlers before the RPC event (lines 780–784 and 347–351).prompt()withoutstreamingBehaviorthrows whileisStreamingis true (lines 860–863). The controller never sendsstreamingBehavior, so a Mosaic prompt never enters Pi's steer or follow-up queues.- The check at line 860 isn't atomic with the run start at line 949.
Between them,
prompt()always awaitsemitBeforeAgentStart(line 915). It also awaits_checkCompaction(line 895) once a turn exists, andemitInput(line 843) when an input handler exists. In that window another run can start: an extension'ssendCustomMessagewithtriggerTurn(lines 1120–1121), or an extension prompt throughsendUserMessage(lines 1161–1187) that passed its own line-860 check first. - The prompt that loses is still acked (
preflightResult(true), line 948). Itsagent.promptthrows (pi-agent-coredist/agent.jsline 228). The RPC handler swallows the throw because preflight succeeded (rpc-mode.jslines 314–317). Itsfinallythen emitsagent_settledand clearsisStreamingwhile the other run goes on. Filbert confirmed this on the installed method with a stub receiver. No real engine was run. - One prompt's run can hold several
agent_start…agent_endpairs before its singleagent_settled, because retries and compaction continue it (_handlePostAgentRun, lines 787–810). In this brief, "a run" means one prompt's run, from its firstagent_startto itsagent_settled. - Events carry no prompt or run ID. The only native evidence tied to a Mosaic prompt is its own response: the ack or the error.
So agent_settled, an idle get_state, and the first user message_start
after an ack aren't tied to a Mosaic prompt whenever other code can start a
turn. The goal extension that repository seats load does that from every
agent_settled. Its handler calls sendUserMessage (goal lines 468–483
and 189), and Pi dispatches the call without awaiting it (lines
2020–2021).
clear_queue doesn't see all external input either (Filbert F2).
clearQueue() returns only _steeringMessages and _followUpMessages
(lines 1195–1203). An extension's sendMessage while streaming goes
straight to agent.steer or agent.followUp (lines 1112–1118).
clearAllQueues drops it without returning it, and the pending count in
get_state (line 1206) misses it too. nextTurn messages (line 1110)
survive clear and abort, and attach to the next prompt (lines 910–913).
The rule R3 follows (Sage). A receipt settles only on positive evidence tied to its own prompt. Idle, settled or empty settle nothing on their own. Where pinned Pi can't tell the cases apart, the controller refuses or reports unknown.
Sealed engine, a binding precondition. Pinned Pi gives no evidence that ties a run to a prompt. So CHAT-03 removes every other source of turns and input, and doesn't guess. A binding needs a sealed engine:
- The controller builds the launch argv. Pi starts with
--no-extensions,--no-prompt-templatesand--no-themes, and with no--extensionargument (lead decision 31;dist/cli/args.jslines 135–140, 167 and 170;docs/usage.mdlines 224 and 233–236). With--no-extensions, Pi loads only the paths given on the command line and ignores settings and packages (dist/core/resource-loader.jslines 316–318), so no explicit extension loads. Skills stay allowed: they expand text and start no turn (§4). - Pi also always loads its built-in extensions, whatever the flags
(
dist/main.jsline 439). Pi 0.85.1 has one,llama.cpp(dist/extensions/index.js). It registers a provider and a/llamacommand, and no event handler or input call (llama/index.jslines 37 and 163). It is part of the pinned package (the Pi pin above), so it is pinned with it. - Any
--extensionargument, or an argv without the three--no-*flags, refuses the binding withunsealed-engine. Explicit extensions aren't pinned or reviewed for binding: Rocko's R3 review showed that an extension can import code outside any hashed tree. They come back with goal in CHAT-06 (lead decision 31). The goal extension starts turns, so a session that loads it can't bind in CHAT-03 (Limit 11). Under lead decision 30, a seat bound to the Console runs without goal, and goal continuation is a CHAT-06 item.
Under the seal, the Mosaic prompt in the slot is the only thing that can start a run, so a run that follows its ack is its run. That exclusion is the evidence basis for attributing by order, and the brief names it as the basis. The seal rests on the launch argv and the pinned Pi build, so the controller also watches for signs that it failed.
Overlap signals. Each of these means the seal failed, or Pi behaved outside the pinned model:
- O1: an
agent_startwhile no slot is held, before the slot's ack, or after the slot's run has settled; - O2: an
agent_settledwhile no slot is held, while the slot is held but before its ack, or a secondagent_settledfor one ack; - O3: an
agent_settledwhile anagent_starthas no matchingagent_end; - O4: an
agent_settledafter an ack, with noagent_startand no failure message between them; - O5: a non-empty
clear_queue, a non-zero pending count, or aqueue_updatethe controller didn't cause. Mosaic never queues, so under the seal these are always empty. Pi's own clear emits aqueue_updatebefore the clear's response (agent-session.jsline 1201,rpc-mode.jsline 334). Aqueue_updateread after the controller writesclear_queueand before that response, with emptysteeringandfollowUp, is the clear's own, and O5 counts it as caused (N25); - O6: a second user
message_startin one run.
On any overlap signal, admission closes, the binding goes to uncertain
with reason run-overlap, and force stop is the way on. The slot's item
becomes delivery-unknown with reason run-overlap if it hasn't reached
working. If it has, it stays working, shown as "outcome unknown". A
stop in progress gets the Unknown outcome. Because a receipt never moves
back from working, a seal failure can misattribute a run to the slot:
another run's user message_start arrives before any signal, the item
goes to working, and the signal that follows can only mark it "outcome
unknown" (N8). Some violations give no signal at all: custom messages
queued straight into the agent, and nextTurn messages (Limit 7). A
spurious settle delayed until after the item's receipt has settled
signals only once the receipt is final (N21, Limit 11).
The rule:
-
Order. Interrupt closes admission, sends
clear_queue, thenabort.abortwould run anything still queued (rpc.mdline 158). -
Clear timeout or failure.
abortis not sent, because it would run whatever is queued. The turnProof recordsnativeQueue: unknown, the stop outcome is Unknown, the stop isuncertain, and admission stays closed. Force stop stays available. It needs no cooperation from the engine (§6). -
What the clear returned. An empty clear gives
nativeQueue: cleared, which here means that Pi's steer and follow-up queues returned nothing. It doesn't prove that no external input was removed, because agent-level custom messages are dropped without being returned (F2). The seal is what keeps those out, and the evidence names the seal as that basis, not the clear. A non-empty clear is overlap signal O5. The stop staysuncertain,reconciledisn't published, admission stays closed (CHAT-01 lines 311–312), and force stop is the way on. The removed items go to controller evidence as digests and byte counts. They are never attributed to a request and never resent. -
Receipt settlement. The Mosaic item's receipt settles from evidence tied to its own prompt, never from
clear_queueand never from the stop. The prompt's own response is tied to it by ID. A run is tied to it only through the seal, and only while no overlap signal has been seen. A late ack that arrives after the fence goes through the same table as an earlier one. An ack alone never means a run.Evidence for the item Receipt Error response to the prompt, before or after the fence failed, native error keptRefused at the dispatch recheck before any write (fence set) dispatch-refused; no engine bytesAck; then get_statesays not streaming, with noagent_startoragent_settledbetween the ack and that replydelivery-unknown, reasonhandled-without-runAck; then agent_startand a usermessage_start, with no overlap signalworking. At the run'sagent_settled, if the observation from the ack to the settle is complete and ordered and holds no overlap signal, thestopReasonof the last assistantmessage_enddecides:stop,lengthortoolUsegivesfinished;errorgivesfailed;abortedgivesfailedwith reasoninterrupted, and evidence links the stop ID. Otherwise (no final assistantmessage_end, anotherstopReasonsuch aspendingordeferred, or a gap), the receipt staysworking, shown as "outcome unknown", and the binding goes touncertain.Ack; then a run that ends in a failure message before any user message_start; thenagent_settleddelivery-unknown, reasonack-without-startAny overlap signal before workingdelivery-unknown, reasonrun-overlapWrite outcome unknown, a missing or unparseable line, or no ack, agent_startorget_statereply within the boundBefore working:delivery-unknown, reasontransport-unknown. Onceworking: it staysworking, shown as "outcome unknown", because a receipt never moves back fromworking. Either way the pipe is poisoned (§1) and the binding goes touncertain.Ack; then agent_settledafter any sequence no row above matchesdelivery-unknown, reasonack-without-startThe transport row and the overlap row take precedence over the rest. Otherwise the rows are exclusive. A line lost without a trace can't be detected, and the table doesn't claim to detect it. An unparseable line is detected.
A preflight still in flight when the fence lands is awaited, bounded, and then classified by this table. If the classification shows a run that is still active, the controller repeats clear, then abort (rule 1), at most three times. If the run is still going after the third,
nativeQueue: unknown, the stop isuncertain, and force stop is the way on.Why
ack-without-startisn'tfailed. A run that fails before its user message emits a failure assistant message from Pi's run-failure handler (pi-agent-coredist/agent.jslines 349–364). Pi persists a run's user message only in themessage_endhandler (agent-session.jslines 386–398), so the session gains no user entry. That ordering is inside Pi. Output reaches the controller through one ordered promise chain (rpc-mode.jslines 28–29,output-guard.jsline 71), so a receivedagent_settledmeans that every earlier event from the process was written first. It doesn't mean that the settle belongs to this prompt (F1). It also doesn't mean that input or before-agent-start extensions had no effect. The controller has no positive evidence that the item had no effect, so the receipt isdelivery-unknown, and afailedthat invites a resend is never shown. This answers Rocko's R2 note 2 and Filbert's F1 row 4.A
delivery-unknownitem is never relabelled unsent and never resent. The client shows it as "outcome unknown". The actor may copy the text into a new request, and the client doesn't present that as a safe resend. -
The stop outcome. The Interrupt stop is classified separately from the receipt. A run is active at a moment if its first
agent_startwas read before that moment and itsagent_settledwasn't. A run is in scope if it was active when the fence was set, or if its firstagent_startwas read before the lastabortwas written. The Interrupted, Completed first and Failed on its own rows also need all of this: exactly one run in scope, itsagent_settledread, and an observation from the fence to that settle that is complete and ordered and holds no overlap signal. Unknown's conditions are checked first. Otherwise the first row that matches wins.Stop outcome Evidence CHAT-01 reconciliation Interrupted The run was active when an abortwas written, and its last assistantmessage_endhasstopReason: aborted. What caused the abort isn't claimed: an extension's abort in the same window still interrupted the turn.turnState: interrupted.reconciledmay follow when the other rule 6 conditions hold (check.mjsline 315).Completed first The run's last stopReasonisstop,lengthortoolUse: it finished before the abort took effectNone honest. The receipt stays finished, and its effects stay the run's own. The completion is never relabelled as an interruption.Failed on its own The run's last stopReasoniserrorNone honest No run No run is in scope. The slot held a written item that settled failedorhandled-without-run, and the observation from the fence to the lastaborthas no gap and no overlap signal.None honest Unknown Anything else: no abortwritten (rule 2), any overlap signal, more than one run in scope, no final assistant message, any otherstopReason, a transport gap, a missing settle, or a run whose firstagent_startwas read only after the last abortNone CHAT-01 lets
reconcile-interruptpass only withturnState: interrupted(check.mjsline 315).input-reconciledbelongs to revocation, and CHAT-03 doesn't borrow it. So until C-5 gives the other outcomes an honest value, every outcome except Interrupted leaves the stopuncertain: noreconciled, prompts refuse (CHAT-01 lines 311–312), and force stop is the way on. The stop advances touncertain(check.mjsline 332), and no turnProof is used for reconciliation. Controller evidence records the outcome, so the client can say "the turn had already finished" instead of "interrupted". The author confirms frompi-agent-corethat every run ended by an abort emits a final assistantmessage_endwithstopReason: aborted(dist/agent.jsline 357,dist/agent-loop.jsline 124). Any path that doesn't is classified Unknown.Nothing to interrupt. Interrupt sets the fence, takes the dispatch lock, and then checks the slot. The dispatch recheck reads the fence under the same lock immediately before the write (§1), and a write already under way counts as written. A slot that was reserved but not written therefore refuses at the recheck: the item becomes
dispatch-refused, the slot is released, and no bytes reach the engine (H9). If, after that, no written item holds the slot and no run is active, Interrupt refusesno-turn. The fence is lifted, admission reopens, and no stop record is created. The refusal has one recorded effect, thedispatch-refuseditem if there was one, and names it. This narrowing happens on the controller side, likebusy. CHAT-01's checker creates a stop for every interrupt it admits (check.mjsline 244). The narrowing leaves only the race with a turn that finishes during the clear and abort exchange. If the controller refuses just before it reads a new run'sagent_start, the run keeps going, and the actor can interrupt again or use force stop. -
Reopening.
reconciledis published and admission reopens only when all of these hold:- the stop outcome is Interrupted (rule 5). Once C-5 is adopted, Completed first, Failed on its own and No run qualify too;
- if the slot held a Mosaic item, its receipt is settled by rule 4 as
finished,failedordelivery-unknown(reasonhandled-without-runorack-without-start). Atransport-unknownorrun-overlapreceipt, or a poisoned pipe, keeps the stopuncertain; - the in-scope run's
agent_settledwas read (No run has none), and aget_statesent after the lastabortshows not streaming; - a
clear_queuesent after that returns empty; - every clear in the sequence was empty (rule 3), and no overlap signal was seen.
Under the seal, a non-empty post-settle clear is overlap signal O5: the stop stays
uncertainand admission stays closed. -
What the proof covers. The proof covers Pi's steer and follow-up queues as of the last clear, and the turnProof says so with its
observedAt. Under the seal no extension queues input. If the seal fails, three things stay invisible: input queued after the last clear, custom messages queued straight into the agent, andnextTurnmessages that attach to the next Mosaic prompt. The first can show up later as an overlap signal; the other two can't (Limit 7). -
Late native events. A user
message_startthat arrives after the fence is recorded as an event. CHAT-01 says only that a receipt is monotonically revised (README line 36) and has no transition table, so the brief fixes the order:admitted<dispatched<acknowledged<working<finishedorfailed.dispatch-refusedis reachable only beforedispatched, anddelivery-unknownonly beforeworking. A receipt never moves to a state claiming the item wasn't sent.
Outside a fence, the slot's item becomes working at the first user
message_start after its ack and agent_start, if no overlap signal has
been seen. The attribution rests on the seal, not on order alone. The
adapter doesn't compare text, because skills and input handlers can change
it. The README says that working is attributed through the seal.
The fake engine models the overlap (N10). Fixtures that break the seal don't load a real extension, since the binding would refuse it. The fake simulates the extension's effect instead.
| # | Fixture | Expected |
|---|---|---|
| N1 | During an active Mosaic run, Interrupt; after the fence is set and before the clear, the fake simulates an unsealed extension that queues a follow-up | clear_queue goes before abort, and the follow-up doesn't run. The queue_update and the non-empty clear are O5: run-overlap, stop outcome Unknown, stop uncertain, prompts refuse, and evidence holds the item's digest. The Mosaic receipt stays working, shown as outcome unknown. Queued before the fence, the queue_update is O5 at once: the binding goes uncertain, and CHAT-01 refuses the Interrupt as fenced (check.mjs line 243). |
| N2 | N1 with abort sent first (mutant) |
The fake runs the external item, so the test fails. This is the ordering guard. |
| N3 | Fence while the Mosaic prompt is in preflight; preflight then errors, and no run exists | Receipt failed with the native error. Stop outcome No run: no turnState: interrupted, stop uncertain, prompts refuse until C-5, force stop available. |
| N4 | Fence while the Mosaic prompt is in preflight; the ack arrives after the first abort, and a run starts |
Classified by rule 4 as a run. Clear, then abort again. The run ends aborted: receipt failed, reason interrupted; stop outcome Interrupted; reconciled after a post-settle empty clear. |
| N5 | Mosaic prompt acked but handled by an input handler the fake simulates, with no run | get_state shows no run, and no agent_start or agent_settled arrived. Receipt delivery-unknown, reason handled-without-run. Nothing is resent. |
| N6 | During an active Mosaic run, Interrupt; the fake simulates an unsealed extension that queues between clear_queue and abort |
The queue_update is O5. abort continues the queued item inside the same run (agent-session.js lines 787–810), so its user message_start is O6. run-overlap, stop outcome Unknown, stop uncertain, admission closed, force stop is the way on. |
| N7 | clear_queue times out during an active run |
No abort sent. Stop outcome Unknown. nativeQueue: unknown, stop uncertain, admission closed. A confirmed force stop still ends the cohort (K1). |
| N8 | Filbert's order one: during Mosaic preflight (paused at the line-915 await), the fake simulates an extension prompt that starts a run first. The wire shows the Mosaic ack, the other run's agent_start and user message_start, then the losing Mosaic prompt's agent_settled |
The other run's user message_start arrives before any signal, so the receipt goes to working. The losing settle is O3: the receipt stays working, shown as outcome unknown, never finished or failed. Binding uncertain, admission closed, nothing resent. The fixture pins the misattribution window (Limit 11). |
| N9 | The item's run started before the fence and ends aborted |
Receipt failed, reason interrupted, not delivery-unknown. Stop outcome Interrupted (turnState: interrupted). |
| N10 | Fake conformance | The fake throws on a prompt while streaming and acks before running. It emits agent_settled from a finally, and it can hold several agent_start … agent_end pairs in one run. It pauses at the preflight awaits (lines 843, 895, 915) between the line-860 check and the run start. A colliding prompt is acked, its throw is swallowed, it settles with no agent_start, and it leaves isStreaming false while the other run goes on. clearQueue emits an empty queue_update before its response, doesn't return agent-level custom messages, and leaves nextTurn messages in place. A fake that queues a Mosaic prompt, or can't produce the overlap, fails the test. |
| N11 | Ack; then the run ends in a failure message before any user message_start; agent_settled arrives; the stream is complete and ordered |
Receipt delivery-unknown, reason ack-without-start, never failed. The session file gains no user entry. The client shows outcome unknown and no resend offer. The same schedule with the failure line unparseable gives delivery-unknown / transport-unknown. |
| N12 | During an active run, the fake simulates an extension that queues after the final empty clear | The stop records the clear's observedAt. The queued item's run starts with no slot held: O1, run-overlap, binding uncertain, admission closed. It is not part of the stopped turn's proof. |
| N13 | The fake emits an agent_start with no slot held (a simulated turn-starting extension) |
O1: binding uncertain, run-overlap, admission closed. A prompt sent afterwards refuses with zero engine bytes. |
| N14 | The run completes normally (stopReason: stop) while clear_queue is in flight; abort then reaches an idle engine; all clears empty |
Receipt finished, effects kept. Stop outcome Completed first. No turnState: interrupted, no reconciled, stop uncertain, prompts refuse until C-5, force stop available. |
| N15 | Fence while the Mosaic prompt is in preflight; after the first clear and abort, an input handler the fake simulates handles it and Pi acks with no run | Receipt delivery-unknown, reason handled-without-run. No second clear or abort is sent. Stop outcome No run: uncertain until C-5. |
| N16 | Interrupt with no slot held and no visible run | Refused no-turn. No stop record, no engine bytes. Admission is open afterwards, and the next prompt is admitted. |
| N17 | The run fails on its own (stopReason: error) during the clear and abort exchange |
Receipt failed. Stop outcome Failed on its own: uncertain until C-5. |
| N18 | The run settles with no final assistant message_end, or a line is lost during the exchange |
Receipt working, shown as outcome unknown. A lost line after working never moves it back to delivery-unknown. A lost line before working gives delivery-unknown / transport-unknown. Stop outcome Unknown: uncertain. |
| N19 | Filbert's order two: the Mosaic prompt wins, and a simulated extension prompt that lost emits its agent_settled before the Mosaic run's user message_start |
O3 (or O4 if it lands before the Mosaic agent_start). Receipt delivery-unknown, reason run-overlap, never failed ack-without-start. Binding uncertain. |
| N20 | A simulated extension timer calls sendCustomMessage with triggerTurn during Mosaic preflight |
As N8. triggerTurn goes to _runAgentPrompt with no preflight (agent-session.js lines 1120–1121), so the extension's run always starts first. Never finished or failed. The same call while the Mosaic run streams is queued straight into the agent (lines 1112–1118) and gives no signal (Limit 7, N22). |
| N21 | The losing prompt's agent_settled is delayed until after the Mosaic receipt settled finished |
O2 on arrival: binding uncertain, run-overlap, admission closed. The receipt stays finished, since receipts are monotonic, and evidence records the overlap against it (Limit 11). |
| N22 | During an active run, the fake simulates an extension sendMessage queued straight into the agent; Interrupt |
The clear returns empty and the item is gone. Evidence records agent-level queues as unobservable, with the seal as the basis, and never as "nothing removed". This fixture pins the Limit 7 gap: no signal fires. |
| N23 | The fake simulates a nextTurn message queued before an Interrupt |
It survives clear and abort and attaches to the next Mosaic prompt. No signal fires. The fixture pins the Limit 7 gap. |
| N24 | Launch with any --extension argument (a local path, an npm or git source), or an argv missing --no-extensions, --no-prompt-templates or --no-themes |
unsealed-engine at bind, for each case. No engine is started. |
| N25 | An ordinary Interrupt of an active Mosaic run: the fake emits Pi's empty queue_update before each clear_queue response, and the run ends aborted |
No overlap signal. Stop outcome Interrupted, reconciled after the post-settle empty clear, admission reopens. A non-empty queue_update in the same window is O5. |
4. The slash path (carry-forward 3)
There are two separate hazards: text left in a composer gets concatenated with a new message, and Pi interprets certain prefixes.
- Concatenation.
send-message.shpastes onto whatever the composer holds. The mediated path has no engine-side composer. Each prompt is one JSONL record, and itsmessagefield holds the whole text. Nothing left over can prefix it. The mediated terminal's own composer is a local buffer. It is empty at start, cleared after each submit and on control transfer, and it submits only while its connection is the controller. - Interpretation. RPC still acts on a leading
/. Extension commands run immediately, even while streaming, and skill and template commands expand (rpc.mdlines 67–69).- Repository seats load one extension command,
/goal(extensions/goal/index.ts:230, loaded through the.pi/extensionspath atscripts/agent-host-dev.sh:137). Prompt templates are off (line 136). Skills are live:--no-skillsstops discovery only, and the ten explicit--skillpaths still load (agent-host-dev.shlines 77–83 and 138; Pidocs/skills.mdline 42). So/skill:<name>expands on repository seats. - A CHAT-03 binding can't load goal (§3, sealed engine), so
/goalhas no command behind it there. The only built-in command,/llama, acts only in TUI mode. Admission doesn't rely on either fact. - Admission refuses
text-policywhen the first non-whitespace character is/(CHAT-01 lines 181–183). Dispatch checks again. The rule holds whatever extensions are loaded, and S1 tests it. - CHAT-01 lines 368–370 leave
!,@and slashes on later lines open under B3. I1 settles them from the pinned source. The author lists every prefix that Pi 0.85.1 RPCpromptinterprets, and the list becomes a fixture. Admission refuses each listed prefix, and any prefix the author can't classify.
- Repository seats load one extension command,
- Board and agent-send for a mediated seat. A mediated engine has no
tmux pane. The board already refuses a reply to a registration without a
tmux session before the tool runs (
packages/control-board/src/serve.mjs:84; testserve.test.mjs:810). CHAT-03 relies on that and changes neither the board nortools/tmux. No real seat registers as mediated in CHAT-03, so the refusal first applies to a real seat at cutover. - Seats still on tmux keep the hazard. The DEFERRED entry stays open. It
closes when a seat moves to the mediated path (CHAT-07), or when a CHAT-03I
charter covers
tools/tmux(CHAT-01 lines 232–238). Neither happens in CHAT-03.
| # | Fixture | Expected |
|---|---|---|
| S1 | /goal x, and the same with leading spaces or a tab |
text-policy at admission. Zero bytes reach the fake engine. |
| S2 | Each prefix on the author's list, including /skill:ms-unslop |
Refused. The test reads the same list the code uses. |
| S3 | /goal on the second line |
The fixture pins the result from the pinned source: admitted if Pi doesn't interpret later lines, refused otherwise |
| S4 | The terminal composer holds a / from an abandoned edit, then control transfers and returns |
Composer cleared. The next submit sends only the new text, and the fake engine records the exact bytes. |
| S5 | An observer terminal gets a paste and then Enter, as send-message.sh does |
Not admitted: controller. Zero engine bytes. |
| S6 | A mediated-shaped registration (no tmux) passed to replyToRow with a recording exec |
409, "no tmux session". exec is never called. |
| S7 | A prompt containing ESC, bracketed-paste markers or U+2028 | Sent as one JSON string. The fake engine receives the exact text, and no record splits. |
5. Control races (Rocko)
Each race uses the fake engine's pause points to land at an exact step. Every race asserts the engine bytes, the receipts and the events. H5–H8 run in I4 on Claude permission requests. Pi dialogs are cut (lead decision 30).
| # | Race | Expected |
|---|---|---|
| H1 | Two takeovers with the same expected generation | One wins, and the generation goes up by 1. The other refuses generation. |
| H2 | The old controller's prompt arrives after a takeover commits | Refused. Zero engine bytes. |
| H3 | A takeover commits while a prompt holds the dispatch lock | The prompt either completed its write before the commit, and the receipt keeps the old actor, or it is refused. If the write outcome is unknown, the §1 rule applies: delivery-unknown, pipe poisoned, uncertain. Never dispatched under both controllers. |
| H4 | Self-takeover | Refused (CHAT-01 line 269) |
| H5 | The old controller answers an approval after a takeover | Refused. The decision is answered once, through the new projection. |
| H6 | Duplicate approval answers with the same content | One native response |
| H7 | Conflicting approval answers with the same request ID | The second is refused |
| H8 | An approval answer after a native timeout or cancel | Refused. The state comes from native evidence. |
| H9 | Interrupt racing a prompt's dispatch | Fence set before the write: the item is dispatch-refused with no engine bytes, and with no run active the Interrupt then refuses no-turn (no stop record). Fence set after the write began: §3 rules 4–6. |
| H10 | Interrupt and force stop at the same time; again with an overlap signal or a revocation closing admission before the Interrupt finds no slot and refuses no-turn |
One stop chain. Force stop supersedes (CHAT-01 lines 319–322). The no-turn cleanup lifts only its own fence: admission stays closed under the surviving reason. |
| H11 | The controller disconnects mid-turn | Work continues and the claim is unchanged. Control stays with the disconnected connection until an observer takes over explicitly. Nothing happens automatically. |
| H12 | Exact retry of a prompt after reconnecting to the same controller incarnation | The same receipt. No second dispatch. |
| H13 | Retry with the same request ID and different text | Refused |
| H14 | Late stdout from the old engine after a replacement | Dropped by execution incarnation and counted in the evidence. Never rendered. |
| H15 | A revoked connection sends a command | Refused, with the revocation fence of CHAT-01 lines 271–280 |
| H16 | A second controller process for the same session | Refused as in W1 and W2. The first controller is untouched. |
| H17 | A confirmation reused, answered from another connection, or answered after the stop context changed | Refused. Confirmations are single-use (CHAT-01 lines 264–268; CHAT-01C lines 92–102). |
| H18 | Two prompts sent before any native output from the first | The second refuses busy. One engine write. |
| H19 | A large prompt line under backpressure; the pipe closes (EPIPE) mid-line, or the controller dies mid-write | delivery-unknown, reason transport-unknown. Pipe poisoned, no later write, uncertain. No retry. |
| H20 | The line is written completely, but the native acknowledgement is lost when the controller dies | After restart: the claim is an orphan (W6). The request's outcome is unknown, and nothing is resent. |
| H21 | Crash after native dispatch, before the client gets its receipt; restart; the client reconnects and retries the exact request with the old token | stale-incarnation. No second engine write. |
| H22 | After H21 and a valid recovery, the client sends a new request with the new token | Admitted normally |
| H23 | Requests pending when the controller restarts; the client library reconnects and gets the new token | The library resends none of them. The fake engine records zero bytes for them. Each shows "outcome unknown, check the transcript". |
6. Stop, cohort proof and recovery (Rocko)
CHAT-01 lines 330–333: a self-posted hash is not trust, and a real
producer/verifier and complete cohort containment remain B3/B4. CHAT-03
builds the producer and its observation procedure. The verifier stays the
fixture's trusted digest registry, as in CHAT-01. So no CHAT-03 proof is
live authority. Where the procedure below can't be carried out, the stop
ends at uncertain, not stopped.
-
Producer. A supervisor shim in
packages/conversation/src/starts withsystemd-run --user --scope, delegated, under the intended unit name recorded in the claim (§2). Inside the scope it moves itself into asupervisorchild cgroup and execs the engine in anenginechild cgroup. So the engine is contained from its first instruction, with no window before it could fork. The cohort is theenginecgroup. The shim is not a member. It holds the scope open, so systemd can't garbage-collect the cgroup before emptiness is read. -
Identity and epoch. The cohort reference is the host (
/etc/machine-id), the boot ID, the unit name and the scope's systemd invocation ID. The invocation ID is the membership epoch. A unit with the right name but a different invocation ID is a different cohort. The supervisor sends it no signals and reports evidence unavailable. -
Containment against migration. A same-uid process can move a pid between cgroups in the user's delegated tree, so a member could leave the scope and survive a "complete" kill. The candidate defence: run the engine in a cgroup namespace rooted at
engine. This host mounts cgroup2 withnsdelegateand allows unprivileged user namespaces, and withnsdelegatea namespaced process can't migrate pids outside its namespace root. The author must show this with K13. If K13 can't be made to refuse the escape, real cohorts never reachstoppedin CHAT-03: K1 then expectsuncertain, and the packet says so. A same-uid process outside the cohort that moves members out is Limit 4 and B2, not something CHAT-03 claims to stop. -
Unavailable is not empty. Emptiness is a readable
engine/cgroup.eventsshowingpopulated 0, for a scope whose invocation ID matches, read by the live shim. A missing or unreadable path, a gone shim, or an invocation-ID mismatch is evidence unavailable, and the stop staysuncertain. The shim stays a member of the scope in its own leaf, so systemd doesn't collect the scope while the shim readsengine. A collected scope or an absentenginecgroup is an absent observation, never an empty one. -
Proof.
stoppedneeds a cohortProof with every CHAT-01 field (schemacohortProof): stop, authority, conversation, execution, cohort reference and membership epoch,membershipComplete, members with boot, pid and start time and a death time each, observation time, and verification digest. Membership is complete as of the freeze in the kill phase below, when no member can fork. A process that exited before the freeze isn't listed; its effects belong to the effect report, not the death proof.membershipComplete: trueis written only when the freeze, enumeration andpopulated 0all succeeded. SIGTERM, EOF, an abort acknowledgement,agent_settledand idle never promote a stop tostopped(line 325). -
Effects. A tool-start with no tool-end at stop time is an
uncertaineffect. Killing never counts as rollback (line 334). Network and remote effects stayuncertain. -
Interrupt (Q20). As in §3 rules 1–8: close admission,
clear_queue,abort, settle the item's receipt (rule 4), classify the stop outcome (rule 5), a post-settle empty clear, the turnProof,reconciled, and reopen admission under current control (lines 307–313). Only the Interrupted outcome reconciles until C-5. If the turn completed or failed first, or no run was active, or any overlap signal was seen (a non-empty clear included), or the outcome is Unknown, or the pipe is poisoned, the stop staysuncertain,reconciledisn't published, and prompts refuse, as CHAT-01 lines 311–312 require before that transition. Force stop is the way on. With nothing running, Interrupt refusesno-turn. If it hangs or the clear times out, force stop stays available. -
Force stop (Q8). Needs an answered confirmation. The stop ID, the answered confirmation and the current phase go into the
stoppingrevision before any signal. Then:- check the scope's invocation ID against the claim;
- TERM phase: SIGTERM to every member of
engine, then a bounded grace period; - kill phase: freeze
engine(cgroup.freeze, waiting forfrozen 1), enumerate every member with pid and start time, writecgroup.kill, and wait forpopulated 0; - build the proof.
It acts only on the selected execution's
enginecgroup, and other executions survive. If a kernel feature the author relies on (cgroup.freeze,cgroup.kill, delegation) isn't available, the stop ends atuncertain. -
Controller death during force stop. A restart (owner proven gone, §2) never assumes TERM or kill happened. It checks the invocation ID. If that matches, it re-runs the escalation from the TERM phase for the same stop record and the same scope. That continues the recorded stop; it doesn't reuse the confirmation for a new one. If identity can't be checked, it sends no signals, and the stop is
uncertain. -
Boot proof. For a claim whose recorded boot differs on the same host, the supervisor issues a cohortProof whose evidence is the boot change. It goes through the same verifier. A different machine ID is
foreign-host, never a boot proof. -
Recover (Q15). Recover reports eligibility only. It needs:
- a stopped proof;
- effects reconciled or explicitly uncertain;
- an exact confirmation;
- the same pins (lines 336–343).
Recover publishes a new
reservedclaim, under a new claim ID, on both keys. The eligibility record names that claim, the stop, the pins, the leaf and the controller incarnation token, and it is single-use. So no one can take the pair between eligibility and launch. Launching is a separate library call made by a launcher. It revalidates the record against the claim and the current leaf, consumes it once, and spawns. No client command launches, and a browser or controller crash can't start an engine. -
Pi resume. Resume reuses the same session file and leaf. The author confirms from the pinned docs how Pi selects the leaf. Before admission opens, the adapter checks
get_stateandget_treeand refuses if the engine loaded any other leaf (RUNTIME.md §3 item 4). A wrongly loaded engine stays under the claim and its scope until a force stop proves it stopped.
| # | Fixture | Expected |
|---|---|---|
| K1 | Force stop an engine whose tool child calls setsid |
The child is killed. stopped with a proof accepted by the fixture verifier if K13 refuses the escape; otherwise uncertain, recorded. |
| K2 | K1 on the process-group fallback (no scope) | uncertain, never stopped |
| K3 | SIGTERM acknowledged while a member is still alive | Stays stopping until the kill phase. Never stopped from TERM alone. |
| K4 | Two fake engines; force stop one | The other survives, shown by independent observation: its own cgroup is populated and it still answers get_state |
| K5 | Stop during a tool call | The effect is uncertain and shown |
| K6 | Recover without proof, without confirmation, or with changed pins | Refused |
| K7 | Recover after proof, then launch | New claim ID and incarnation, generation +1, same leaf. The cancelled prompt is not replayed. |
| K8 | The engine loads a different leaf on resume | Refused before admission opens. The engine stays claimed and contained until force stop proves it stopped. |
| K9 | An interrupt that never settles | Stays uncertain. Force stop stays available, and ordinary takeover is refused while fenced. |
| K10 | Controller killed between the TERM and kill phases | Restart checks the invocation ID and re-runs from TERM for the same stop. Nothing is recorded as killed that wasn't observed. |
| K11 | Controller killed after the confirmation is recorded, before TERM | Same as K10 |
| K12 | A member that forks in a loop during enumeration and kill | The freeze stops forking. The enumeration is complete, and populated 0 follows cgroup.kill. |
| K13 | A member writes its own pid to another cgroup's cgroup.procs |
Refused by the namespace, and the kill is complete. If not refused, stopped is unavailable for real cohorts, and K1 expects uncertain. |
| K14 | A unit with the recorded name but a different invocation ID | Evidence unavailable. No signals, uncertain. |
| K15 | engine cgroup path missing or unreadable, or the shim gone |
Evidence unavailable, not empty. uncertain. |
| K16 | Boot proof requested for a claim from a different machine ID | foreign-host. No proof. |
| K17 | Two launcher calls with one eligibility record | One launch. The other refuses, and no second engine starts. |
| K18 | The leaf changes after eligibility, before launch | Launch refused. The reservation stays until released with proof. |
7. Claude permission requests (I4)
- Pi native dialogs are cut from CHAT-03 (lead decision 30). CHAT-01
lines 239–242 and 292 block non-permission forms until CHAT-03D exists,
so a Pi
extension_ui_requestis shown disabled, with a reason, and is never answered. - Claude, in I4 after B1.
can_use_toolmaps to allow-once and deny. Choices that widen permissions are disabled, and there is no invented cancel (lines 294–295). Pending permission requests are re-read from initialize on reconnect, but only if B1 shows that the pinned version supports it.
| # | Fixture | Expected |
|---|---|---|
| P3 | A Pi confirm, select, input or editor dialog |
Disabled, with a reason. No response is sent. |
| P7 | A Claude permission request (I4, recorded fixture) | Allow-once and deny only. Widening choices disabled. |
8. Claude: B1 and the catalogue (carry-forward 2)
What B1 is. docs/plans/chat-00/README.md lines 193–195: prove the
exact Claude CLI and protocol combination, capability negotiation,
transcript branch selection and the image and file input shapes. SDK and
source examples narrow the uncertainty, but they don't certify the
installed binary. CHAT-00 line 66 names the gap for history: Claude's
persisted branch format and leaf selection.
What proves it (I3). A B1 packet under agents/<author>/work/chat-03/b1/,
approved by Filbert and Rocko, with five parts:
-
Pin. The exact version string and the sha256 of the binary the adapter will run. Today
claude --versionreports 2.1.283; the plan saw 2.1.269. The adapter checks both at launch and otherwise refuses withengine-pin-mismatch. It runs the engine with auto-update off, and the author cites the setting from the pinned docs. -
Sources. The official protocol sources for that exact version, by hash. These are the C-* entries in
chat-00/sources.json, re-pinned to that version. C-HEADLESS and C-REF are floating documents (sources.jsonlines 106–119) and can't be tied to a version. The packet fixes them as fetched copies, each with its sha256 and fetch date, and says plainly that they aren't version-bound. -
Recordings. Made from that binary, in a scratch config directory with no repository credentials:
- initialize and its capability answer, including whether
interrupt_cancel_queued_v1andpending_permission_requestsexist; - a text prompt and an image prompt;
- a
can_use_toolallow and a deny; - an interrupt.
They are redacted, hash-pinned and used as the I4 fixtures. The plan rules out model calls for discovery (line 166). Any recording that calls a model or reads Claude auth therefore needs Jason's go first, because it spends money and touches credentials. Without that go, I3 stops after parts 1 and 2, and B1 stays open.
- initialize and its capability answer, including whether
-
Branch selection. Two persisted transcripts from that binary, one of them with a fork (an edit or a rewind). The packet states the leaf rule and shows that
--resumecontinues the leaf the parser picks. -
Negative. The parser refuses a transcript from another version with
unsupported-harness.
The catalogue after B1 (I4).
- Rocko's launcher runs Claude with the default config directory
(
agents/rocko/launch.sh:86, noCLAUDE_CONFIG_DIR). So its transcripts sit in the same~/.claude/projects/<cwd slug>/directory as every other Claude session started in this checkout, including T3 threads and Jason's own sessions. - The catalogue therefore never lists that directory. It opens only session
IDs the repository launcher recorded for that seat
(
.pi/state/rocko/session-idandlaunches/*/session-id). Each ID must resolve to exactly one file under the approved root, opened with the CHAT-02 safe-open rules. - A launcher receipt is a hint, not membership. CHAT-01 lines 62–64 require an approved project and workspace mapping, and say OS readability, cwd and seat labels don't establish it. The access mapping stays under B4 review.
- The CHAT-02 guarantees still apply: read-only, no writes (F17), cursors, byte caps, continuation parts and inert rendering.
| # | Fixture | Expected |
|---|---|---|
| L1 | A Claude transcript at the pinned version | Default leaf shown and other branches readable, as in F12 |
| L2 | A transcript from another version | unsupported-harness |
| L3 | A decoy session file in the same directory, named in no receipt | Never listed, never opened |
| L4 | A receipt naming a file outside the root, a symlink or another project | Refused, as in F7–F9 |
| L5 | Sidechain or subagent entries | Kept separate, not merged into the main branch |
| L6 | No-write check | F17 over the Claude directory |
| L7 | The binary's sha256 or version differs at launch | engine-pin-mismatch. No launch. |
The I4 adapter maps Claude stream-json onto CHAT-01 events. The §5 and §9 fixtures then run again against the recordings.
9. Return flow and events
| # | Fixture | Expected |
|---|---|---|
| E1 | The plan §6 return flow on the fake engine: send, acknowledge, user, toolCall, toolResult, new final answer | Shown once in the same conversation, with no refresh, no duplication, and the draft and reading position kept |
| E2 | U+2028 and U+2029 inside JSON strings, and CRLF | Each parsed as one record |
| E3 | Multipart final, two blocks, null request correlation, duplicate delivery | The CHAT-01 cases at lines 110–112 |
| E4 | A page read just after a message_end but before its entry is persisted, then a subscription; again with a gap or a new epoch |
Replay is advertised unavailable. The client shows a reconcile marker at the seam and re-reads the page after agent_settled. After that, each message appears exactly once. A gap or new epoch also reconciles. If the author shows an atomic cut from the source, overlap is deduplicated instead, and the fixture pins that case. |
| E5 | An unknown native event | No client event. Controller evidence records its type and byte count, and the terminal count goes up. No record fails the schema. |
| E6 | A tool result delayed across a pause and a reconnect | Reconciled without a manual refresh |
| E7 | The mediated terminal as observer, then as controller | Renders the same stream as the library client, and submits only as controller |
The browser leg of the return flow is CHAT-05. I1 proves it at the library and the terminal.
10. Checks and mutation pass
-
Suites. These must pass:
node --test packages/conversation/tests/;- the control-board, webui and seat package tests;
node docs/plans/chat-00/check.mjs,chat-01/check.mjsandchat-01c/check.mjs;- every
scripts/test-*.shsuite at the candidate's base, includingtest-queue.sh, since queue A1 landed in34a72af9.
-
Nested runners. A test that starts
node --testitself clearsNODE_TEST_CONTEXTfor that child, or it can hide failures. The DEFERRED entry on nested runners (Filbert's A1 review6933b885, note N13) was closed at34a72af9, and new tests keep to the same rule. -
Isolation. Every fixture runs in temporary directories. No fixture touches live registrations,
.pi/state,~/.claudeor a live seat (plan line 272). -
Mutation pass. Run in a scratch copy, as for the CHAT-02 Console. Each of these mutants must fail at least one test:
- dispatch skips the generation recheck;
abortis sent beforeclear_queue;- a revision is published by rename or overwrite, so an existing revision name can be replaced;
- the late-event filter ignores the incarnation;
- the text policy checks only the first character, missing leading whitespace;
- an acknowledged SIGTERM promotes a stop to
stopped; - the process-group fallback can reach
stopped; - a confirmation is not consumed on use;
- disconnect releases control;
- the composer isn't cleared on transfer;
- engine output is read with
readline; - a revision is published by exclusive create in place of the
temp-file-then-
link()sequence, so a partial file becomes visible; - a second controller reclassifies a claim whose owner is alive;
- the live-session guard skips the real-path check;
- the pending slot is released at
agent_startinstead of at reconciliation; - a pipe is written again after an unknown write outcome;
- an ack followed by no run marks the item
workingorfinishedinstead ofdelivery-unknown(N5); - admission reopens without the post-settle empty clear;
abortis sent after aclear_queuetimeout;- the incarnation token check is skipped;
- a missing cgroup path counts as empty;
- a unit with a different invocation ID is signalled;
- a restart during force stop records the kill phase as done;
- an eligibility record can be used twice;
- a non-empty
clear_queuereachesreconciled; - the pending slot's in-flight preflight is abandoned at the fence instead of awaited, so a late ack's run is never aborted (N4);
- the fake engine queues a Mosaic prompt while streaming instead of throwing (N10's conformance check must catch it);
- a no-unit observation frees a pair that has a spawn marker (W20);
- a settled slot, an idle engine and empty clears are taken as
interruption evidence, so
turnState: interruptedis written without a finalstopReason: aborted(N14 and N15 must catch it); - a receipt whose run ended
stopis relabelledfailed(reasoninterrupted) because a stop was in progress (N14); turnState: input-reconciledis used to reconcile an Interrupt (N3, N14, N15);- Interrupt with no slot held and no visible run creates a stop record or writes to the engine (N16);
- a run is attributed to the slot with the overlap checks skipped, so
the receipt settles
finishedorfailedafter an overlap signal, or the binding staysactive(N8, N19, N20); - a settle that doesn't close the slot's own run (no
agent_startsince the ack, or anagent_startwith noagent_end) is taken as the slot's settle (N8, N19); - an
agent_startoragent_settledwith no slot held is ignored, so admission stays open (N13, N21); - the binding launches without the
--no-*flags, or with an--extensionargument (N24); - an
ack-without-startrun settlesfailed(N11); - an empty clear is recorded as proof that no external input was removed, instead of naming the seal as the basis (N22).
Mutant 25 (drift) left with the drift check for CHAT-07, and its number isn't reused. I4 adds two more: the pin check is skipped, and the decoy file is listed.
11. Choices in this brief, for review
| Choice | Tradeoff |
|---|---|
No broker queue; a busy engine refuses busy |
No queued follow-ups until CHAT-04, and no native queueing of mediated input |
No streamingBehavior, steer or follow_up, and one pending slot |
Same as above. On pinned Pi a Mosaic prompt then never enters a native queue (lines 860–863). It can still lose a race in preflight to another run (lines 860–949), which the seal and the overlap signals handle (§3). |
Sealed engine: --no-extensions and no --extension (unsealed-engine; lead decision 31) |
Without run IDs, only excluding every other source of turns lets a run be attributed to the slot by order. The cost is that no explicit extension loads in CHAT-03, goal included (Limit 11). They come back in CHAT-06. The overlap signals watch for a seal failure. |
Overlap signals close admission and make the binding uncertain (run-overlap) |
One overlap needs a force stop to move on, and never produces a guessed receipt. Some seal failures give no signal (Limits 7 and 11). |
The slot settles only on evidence tied to its own prompt, never from clear_queue, idle or a settle alone |
An ack with no run, and a run that fails before its user message, stay delivery-unknown: no "unsent" state and no safe-resend button |
| Receipt settlement and the stop outcome are classified separately (§3 rules 4 and 5) | A turn that finishes during an Interrupt is reported as finished, not interrupted. The cost is that the stop can't reconcile: until C-5 it stays uncertain and needs a force stop, even though nothing was lost. |
Interrupt with nothing running refuses no-turn |
Narrows the completion race to the clear and abort window. The controller can refuse just before it reads a new run's agent_start; the run keeps going and the actor retries. |
| Poison the pipe on an unknown write | One transport glitch needs a force stop, with no guessing about half-written lines |
| Live-session guard by path | Fixture-only is enforced in code. The guard needs a reviewed change to lift at CHAT-07. |
| No default claim root | Nothing live can use I1 before the CHAT-07 data-map decision |
Cohort is an engine child cgroup in a delegated user scope, held by a shim, with a cgroup namespace against migration |
Depends on systemd-run --user, delegation, cgroup.freeze, cgroup.kill and user namespaces. Without them, or if K13 fails, stops end at uncertain. The verifier is fixture-grade until B3/B4. |
One actor, local-operator, guarded only by socket-directory mode |
Same-uid agents are not kept out (Limit 1) |
| Dedup lasts one controller incarnation, fenced by a token (deviation V-1, accepted with limits in lead decision 25) | A retry across a restart refuses stale-incarnation instead of returning the receipt. The client shows "outcome unknown, check the transcript" and never resends. CHAT-04's durable receipts must restore the contract behavior. |
| Three increments, with I3/I4 gated on B1 | CHAT-03 can close Pi-only, with that outcome recorded |
| Claude catalogue from launcher receipts only | A Claude session started outside the launcher never appears |
Carry-forwards
| # | Sage's item | Where | Acceptance evidence |
|---|---|---|---|
| 1 | D1 writer-claim record and R3-1 (lead decisions item 8) | §1, §2, §3 | W1–W9, W11–W17, W20, G1–G3, N1–N25 and H18–H23 pass. Mutants 2, 3, 12–20 and 26–39 fail tests. R3-1 binds only sealed engines, settles each receipt only on evidence tied to its own prompt (§3 rule 4), and classifies the stop separately (rule 5). Any overlap signal closes admission. Only an interrupted turn reconciles. The other stop outcomes leave the stop uncertain until C-5 lands in CHAT-04 (lead decision 30). |
| 2 | The Claude catalogue after B1 | §8; increments I3 and I4 | The B1 packet (parts 1–5), approved by Filbert and Rocko. L1–L7 pass. |
| 3 | The board send that can become a Pi slash command | §4 | S1–S7 pass. Mutants 5 and 10 fail tests. The DEFERRED entry stays open for seats still on tmux. |
| 4 | The CHAT-01/01C contracts, by path and hash | "Contracts implemented" | The hash table matches sha256sum at approval. C-3 is a separate reviewed item if B1 needs it, and C-5 moves to CHAT-04. C-1, C-2 and C-4 are cut (lead decision 30). Deviation V-1 is accepted with limits (lead decision 25): H21 and H23 pass, the client never resends, and V-1's closure is a required CHAT-04 item (lead decision 25). CHAT-04's brief must list it. |
| 5 | Path authority | "Files owned" | The overlap table. Every changed path in each candidate is inside "Files owned". |
| 6 | The source author | "Owner and reviewer" | <slot>, named by Darkwing after this brief is approved |
Limits
- Same-uid control. Any process running as the same user can connect to the socket and take control. Agent seats run as that user, so an agent could take over another conversation. CHAT-03 is fixture-only, so no live conversation is exposed. Live use waits for a reviewed local owner-channel design (B4).
- Fake engines. They model the pinned docs. The real-binary smoke covers framing only. I1 makes no model call, so real-engine behavior is unproven until CHAT-06.
- B2 untouched. Tool isolation and trust and settings parity (CHAT-00 lines 196–197) are untouched. Any real engine for a real seat waits on B2.
- Containment. The cgroup covers local processes. Network and remote
effects stay
uncertain. A same-uid process outside the cohort can still move members out. The verifier is the fixture registry. No CHAT-03 stop proof is live authority until B3/B4 and B2. - Sole writer. CHAT-03 has no session drift check; it moves to CHAT-07 (lead decision 30). The claim binds only controllers that use it. Real sessions stay off-limits through the live-session guard.
- Durability. W5 kills processes at each barrier. Real power loss
isn't tested; the fsync-then-
link()order is the stated basis. - What the clear can't see. Under the seal no extension queues input,
so the stop proof rests on the seal. If the seal fails, three things
escape it. Input queued after the final empty clear is outside the
proof (§3 rule 7), and it shows up only if it starts a run (O1) or
queues visibly (O5).
Custom messages an extension queues straight into the agent while it
streams (
agent-session.jslines 1112–1118) are dropped by the clear without being returned (lines 1195–1203) and aren't counted inget_state(line 1206).nextTurnmessages (line 1110) survive clear and abort and attach to the next Mosaic prompt (lines 910–913). The last two give no signal at all (N22, N23). - Takeover has nothing to recover. With no broker queue, a takeover has no queued input to turn into drafts. Q13 is CHAT-04's.
- Interrupt racing completion. A turn can finish, fail, or never start while the clear and abort are in flight. The receipt reports what happened, but until C-5 the stop can't reconcile and needs a force stop. Sessions with short turns will see this more often.
- Restart dedup (V-1). Within one incarnation a retry returns the
receipt. Across a restart it refuses
stale-incarnation, and the actor has to check the transcript. CHAT-04's durable receipts close this. - Sealed engines only. Pinned Pi ties no run to a prompt, so
CHAT-03 attributes runs by order under the seal. A session that loads
any explicit extension, goal included, can't bind (
unsealed-engine, lead decision 31). The five Pi repository seats (Darkwing, Dewey, Filbert, Researcher, Sage) load goal throughscripts/agent-host-dev.shline 137, so none of them can bind with goal loaded. Under lead decision 30, a seat bound to the Console runs without goal, and goal continuation is a CHAT-06 item. The seal rests on the launch argv and the pinned Pi build. A failure there that the overlap signals don't catch goes undetected if it queues straight into the agent (Limit 7). Other failures are caught late. Another run's usermessage_startthat arrives before any signal puts the item inworking, and the signal that follows can only mark it outcome unknown (N8). A spurious settle delayed until after the receipt settled signals only once the receipt is final (N21). CHAT-06 owns explicit and turn-starting extensions, which need run-to-prompt evidence from Pi.
Out of scope
- Durable queues, drafts, uploads, queue edit and cancel, takeover-to-draft (CHAT-04).
- Remote control (CHAT-04R) and the chat UI (CHAT-05).
- The candidate, cutover and any real seat migration (CHAT-06, CHAT-07).
- Changes to
tools/tmux, the board, the seat package, launchers,roles/,contracts/and thedocs/plans/chat-0*contracts. - The CHAT-03I fleet-communications charter (CHAT-01 lines 232–238).
- Authenticated multi-actor control and the local owner-channel design (B4).
- The Claude adapter and history, unless B1 passes.
- Filbert's idsDigest note, which stays in
agents/dewey/work/chat-02/FOLLOWUPS.md. - Explicit extensions, goal continuation among them, and any pinning of them (CHAT-06; Limit 11; lead decisions 30 and 31).
- Cut by lead decision 30: Pi native dialogs and the CHAT-03D companion
(C-2), the unknown-event type (C-4) and the removed-input
turnProof(C-1). C-5 moves to CHAT-04, and the idle session drift check to CHAT-07.
Gate
- Brief.
- Filbert approves the exact
BRIEF.mdhash. - Rocko's pass on §5 and §6 leaves no blocking finding open.
- Sage commits and pins the brief. Darkwing then names the author.
- Filbert approves the exact
- Each increment.
- Filbert approves the exact candidate hashes. Rocko passes the §5 and §6 code in I1 and the B1 packet in I3.
- Every §10 suite is green on an index export, and every mutant is killed.
- Sage commits. A push needs Jason's word.
- Contract items. Each is a separate reviewed change to the CHAT-01
files, never folded into an increment.
- C-3, if B1 shows a gap, is approved before I4 starts.
- C-5 moves to CHAT-04 (lead decision 30). CHAT-03 ships the interim
rule: every stop outcome except Interrupted stays
uncertain. - Deviation V-1: accepted with limits (lead decision 25). I1 is approved only with H21 and H23 passing and the client's "outcome unknown, check the transcript" display. V-1 stays open until CHAT-04's durable receipts restore the contract behavior; CHAT-04's brief must list that as a required item.
- I3. Jason's go comes before any model call.
- CHAT-03 done. All of these:
- I1 is approved and committed;
- I4 is approved and committed, or B1 is recorded as not passed (the triggers are in "What ships"), with Claude still refusing and CHAT-06 blocked.
- No live check. No real seat migrates, so there is none. The live proof belongs to CHAT-07.