docs(chat-03): brief pinned at 1ef15ac0 after r3 and scope check, lead decision 33 (#1507)

Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
2026-09-26 19:45:04 -05:00
co-authored by Claude Opus 5.5
parent 93ee5054f4
commit 8a465e891a
18 changed files with 6822 additions and 1 deletions
@@ -0,0 +1,193 @@
# CHAT-03 R1 — adversarial review of control and recovery
Verdict: **request changes**. Rocko, 2026-09-26.
Verified pins before reading:
- BRIEF.md: `5dd447f710d425ae01a12dcb56421fa948faf6b44d07c8778ba3ab9f8459daf1`
- REVIEW-REQUEST.md: `3a03adb9001d6e07c044cbb91b3b9dadf7516a6804e17097ac6a3f4dcd74193e`
- Stated base: `40a02d2b`.
Scope: sections 5–6 and their section 2/3 dependencies. Findings below are
specification failures/schedules, not executed tests of unbuilt code.
Pinned Pi source was inspected for the native-protocol claims. No live
engine, scope or session was changed.
## 1. Blocking — exclusive revision files do not define a crash-safe two-key claim transaction
Section 2 specifies highest-revision-wins files, exclusive create and two
keys taken in order. Ordering prevents a lock-order deadlock, but says
nothing about how those two durable histories commit together. Crash after
reserving the seat key but before the session key leaves a partial claim.
Crash after advancing one key to stopped leaves disagreement. A crash
between exclusive creation of the final revision pathname and completion
of its contents exposes an incomplete highest revision, not the promised
old or new complete state. fsync after the write cannot prevent that window.
W5's SIGKILL-only fixture also does not test power-loss durability. W4's
“no partial claim left” requires a protocol, not just an assertion.
**Fix:** specify an atomic publication boundary for complete revisions and
an explicit transaction/claim id tying both keys together. Describe restart
handling of partially reserved and partially transitioned pairs, including
reserved/stopping/uncertain, not only active. Uncertain/invalid publication
must retain exclusion; do not skip a damaged latest record and reuse an
older stopped state. Define controller ownership so a losing contender
cannot change an active holder's claim into an orphan merely because that
contender cannot attach to its pipe. Include launch-before-active-publication
and crash-before-cohort-identity-publication windows.
**Add fixtures:** crash between every key publication and file publication
barrier; paused live holder versus second startup; engine spawned before
active publication; partial release. Prove both keys remain exclusive and
that recovery never deletes another holder's reservation.
## 2. Blocking — dispatch tests omit concurrent prompts and partial/unknown pipe writes
H3's “never a partial write” is not guaranteed by a controller mutex.
A large JSON line can be partially delivered when the controller or pipe
fails; stdin.write acceptance/backpressure is not native consumption proof.
Crash after complete dispatch but before the client receipt is another
unknown-outcome point.
Also, “busy engine” is not defined to include a sent prompt whose
agent_start has not yet been read. Two successive prompt commands can both
see native idle before the first start and violate the one-unstarted-item
assumption that R3-1 depends on. This is the same class of race found in
Discord 6b, even though CHAT-03 uses a separate adapter.
**Fix:** reserve a pending dispatch slot under the dispatch lock before
writing; count it as busy through delayed start/acceptance and terminal
reconciliation. Distinguish local write, native acknowledgment and unknown
outcome. On partial/error/crash, fence and retain uncertainty, never claim
unsent or retry automatically.
**Add fixtures:** two prompts before any native output; backpressure and
partial line followed by EPIPE/controller death; completed write with lost
acknowledgment. Retain H3's actor/generation ordering, but permit an honest
unknown transport outcome rather than claiming impossible write atomicity.
## 3. Blocking — exact returned text is not unique input identity or a lasting empty-queue proof
R3-1 treats one exact clear_queue text match as proof that the dispatched
Mosaic request was unconsumed. The brief explicitly allows extension input.
An extension may enqueue the identical text while Mosaic's input is still
in asynchronous preflight or its start event is pending. Clearing that
external item once satisfies the proposed predicate without proving the
Mosaic item was removed. The UI then invites the actor to submit the text
again, potentially duplicating work.
There is a second interval: an extension or in-flight preflight can enqueue
after clear_queue but before/during abort. Command order alone does not
prove CHAT-01's cleared native queue with no remaining IDs at reconciliation.
A fake that never injects after clear_queue can pass N1/N2 while this
property is false. The pinned RPC handler launches session.prompt
asynchronously and forwards its preflight result; clear_queue returns the
session's current queue, not a dispatch-id-bound removal receipt.
**Fix:** state the provenance assumptions that can actually be enforced.
Without unique attributable native evidence, treat returned text as
ambiguous, not proven-unconsumed. Do not offer an apparently safe resend.
Require the turnProof to cover in-flight preflight and any post-clear
insertion before reopening. If the protocol cannot establish that, remain
uncertain. Contract C-1 cannot make a text equality test into identity proof.
**Add fixtures:** extension queues exactly the dispatched text once;
preflight resumes after clear_queue; extension enqueues between clear and
abort; native events arrive late. Also specify whether abort is issued on
a timed-out clear and how a confirmed force stop remains able to terminate
the cohort without waiting forever for native cooperation.
## 4. Blocking — cohort proof lacks a trusted containment and observation procedure
A systemd scope survives setsid, which is useful, but an empty cgroup
snapshot is not by itself the complete-membership/death proof required by
CHAT-01. The brief does not define who establishes the membership epoch,
how engine launch is contained before it can fork, how scope identity is
protected from reuse, or how it distinguishes an empty scope from a missing
or inaccessible path. Sampling current pids misses short-lived members;
same-uid processes may have ways to leave a delegated scope. A fake that
reports its own membership/death cannot supply the missing authority.
Section 6 lists some proof fields but omits the contract's explicit
stop/authority/conversation/execution/cohort/epoch binding and trusted
verification digest provenance. A different boot proves old local
processes dead only when the claim is bound to this same host; no host
identity or nonportable-root rule is stated. A copied/foreign-host claim
must not become stopped just because its boot id differs.
**Fix:** name the supervisor producer/verifier and containment assumption
for these fixtures, distinguish unavailable evidence from an empty cohort,
and require immutable scope incarnation/host/boot/epoch attribution. Either
provide trustworthy full proof or leave stopped unavailable. Specify what
happens if the controller dies before sending TERM or scheduling KILL;
K10's “resumes observing” cannot imply that missing escalation has happened.
**Add fixtures:** fork during enumeration/termination, unavailable or reused
scope, escape attempt or explicitly refused unsupported containment,
foreign-host boot mismatch, controller death before TERM and between TERM
and KILL. Verify other cohorts survive using independent observations.
Process-group fallback remaining uncertain is correct.
## 5. Blocking — restart dedup refusal has no way to distinguish old requests from new ones
Section 1 discards its dedup index on controller restart but promises that
an old retry gets receipt-unknown and is never dispatched. With only actor,
conversation and arbitrary client request id, an empty index cannot tell
an old id from a new one. H12 tests reconnect, not restart.
**Fix:** require a controller-incarnation token in admitted requests and
reject obsolete tokens even when content/id is retried, or persist enough
request identity to refuse old attempts. The first option preserves the
stated non-durable scope. Unknown old outcomes stay explicit; no heuristic
based on text or timestamps.
**Add fixture:** crash after native dispatch before receipt, restart,
reconnect and retry exactly the same request; prove no second engine write.
Then prove a genuinely new request can be admitted after valid recovery.
## 6. Blocking — foreign-writer detection assumes stream IDs Pi does not emit
Section 2 calls every appended session-entry id absent from the controller's
stream foreign. In pinned Pi, agent-session.js emits message_end to
listeners before calling SessionManager.appendMessage. appendMessage then
generates the session-entry id; RPC simply forwards toJsonEvent(event).
The generated entry id is not thereby present in that message event.
Settings and other session entries need a mapping too. A fake that assigns
the same fabricated id to both sides would falsely validate W10 and mark
normal Pi appends foreign in the real adapter.
**Fix:** identify an actual native correlation/snapshot protocol and its
race bounds, or advertise foreign-writer attribution as unavailable and
fence on unexplained drift without claiming a known foreign writer.
get_entries/get_tree can expose ids, but merely trusting the engine's
post-hoc report is not proof of sole writer attribution.
**Add fixture:** generate session entries with the pinned SessionManager
and pinned event shape, including settings and delayed persistence. A
normal own append must not trigger the foreign-writer result; an external
append cannot be blessed solely because a later snapshot contains it.
## 7. Not blocking — tighten recovery acceptance at implementation review
Recovery eligibility is correctly separated from launch. State explicitly
that eligibility is a single-use, incarnation/pin/leaf-bound input, and that
the launcher reacquires/retains both claims and revalidates it before spawn.
Add two simultaneous launcher calls with one eligibility record and a leaf
change after eligibility. The existing exclusivity requirement already
implies refusal, so this is a missing acceptance case rather than a new
architecture blocker. Wrong-leaf refusal must retain containment/claims
until the mistakenly launched engine is independently proven stopped.
## What is fine
Same-uid control is honestly limited to fixtures and explicitly blocks live
use. Disconnect retaining claims/control, revoked-generation checks,
confirmation binding, no automatic replay, unknown effects after kill,
settled-not-stopped, explicit force-stop supersession and refusing changed
resume pins all match the intended contract. Retain those requirements.
No finding here authorizes broader source ownership, a live seat migration,
model calls, new credentials or a production scope configuration.
Only this report was written. No source, contract or session edits; no
commit. Please resolve the six blocking findings before calling the brief
ready to build. Filbert's independent whole-brief review remains separate.
@@ -0,0 +1,30 @@
# CHAT-03 R1 — correction to finding 3
Rocko, 2026-09-26. Supplements, without changing, report
`89752c2b95f93b3eefd9f4286d24277de25797dbb9f7a161b321035f6e12ff10`
against brief R1 `5dd447f7…af1`.
Filbert's review `ec00544e72d07d19180ea7e40ae769e5ef9917cc177703e10b0fbf8ebdae6e17`
correctly identifies a stronger problem with R1's N1/N2 premise. I checked
pinned agent-session.js: isStreaming reads _isAgentRunActive; that remains
true across post-run continuation/retry until settlement. A prompt without
streamingBehavior throws while it is true. The stated retry-window native
queueing of Mosaic input is therefore not a reachable case under this
adapter policy. My original finding 3 did not establish that reachability
and should not be read as doing so.
Correction: remove the claimed Mosaic-queue attribution case and rebuild
N1/N2 around external queue clearing and abort order, following Filbert B1.
Do not add a complicated text-identity mechanism to support an unreachable
path. An exact text match still cannot turn external input into proof about
a Mosaic request. Any accepted-but-not-started request needs its own native
preflight/failure analysis and an honest uncertain outcome, not deduction
from clear_queue text. Post-clear external insertion remains an independent
queue-reconciliation question; show the supported assumptions or stay
uncertain before reopening.
This narrows the proposed remedy for my finding 3; it does not change the
request-changes verdict. Filbert B2 corroborates finding 6's lack of stream
entry IDs, and B6 corroborates finding 1's split-key/reserved recovery gap.
No revised brief has been approved by this addendum. No source or contract
was edited.
@@ -0,0 +1,59 @@
# CHAT-03 R2 — control, stop and recovery adversarial review
Verdict: **revise**. Rocko, 2026-09-26. One blocking finding.
Verified candidate pins before review and again at completion:
- BRIEF.md: `5c5b45a277f3a555337a5caf57ff1b300b95781b69640ec8974770bdd44af9bd`.
- REVIEW-REQUEST.md: `c75ad86fc014b7a13f72132c0378d6e8abf09663820eee701dc67e3ed50ef92f`.
- Stated base: `40a02d2b`.
- Installed `agent-session.js` matches the packet: `fb8a3981c20c8c0bbd42231b1c99a10335fb3858b659056b341954de9cfa467f`.
Scope: sections 5–6, their section 2/3 dependencies, and disposition of my R1 findings and addendum. This is a specification review, not an implementation certification. No source, live session, cgroup, service, configuration or contract was changed.
## 1. Blocking — settling the slot does not establish an interrupted turn
The slot table at section 3 lines 591–595 assumes that a late preflight ack means a run started, and that a visibly started run follows an interrupted turn. Rule 5 then permits reconciliation after a settled slot, native idle and empty clears. Those observations do not distinguish interruption from ordinary completion or no run.
Concrete schedules:
1. A run has emitted its user message. Interrupt fences admission and sends `clear_queue`. While that exchange is in flight, the run completes normally, emits its successful assistant result and settles. The controller then sends `abort` against an idle engine. A subsequent `get_state` is idle and the post-settle clear is empty. Every clear can be empty, so C-1 does not catch this. The table says `turnState: interrupted`, despite the successful completion.
2. Interrupt fences a prompt still in preflight. An input extension returns `handled` after the first clear/abort. Pi acknowledges it without starting a run. The first table row instead says “Ack: a run started.” Repeating clear/abort cannot manufacture a run or an interruption. Likewise, a preflight error can settle the slot as `failed` without there ever being a turn to interrupt.
This is a contract boundary, not just a missing test name. `docs/plans/chat-01/check.mjs:315` accepts `reconcile-interrupt` only with `turnState: interrupted`; line 325 records `turn-interrupted`. `input-reconciled` is reserved by that checker for revocation. A fixture that always aborts an active fake run can pass N3/N4/N9 while the real schedules above remain unrepresented.
Native evidence: `agent-session.js:832–847` acknowledges extension-handled input without `_runAgentPrompt`; `abort()` at 1222 only aborts current operations and waits for idle. It does not certify that a run was interrupted. I invoked that installed abort method on a minimal idle receiver: it resolved with no run events. That is a method-level check, not a real-engine smoke or model call.
**Required fix:** separate receipt settlement from interrupt-proof construction. Route a late ack through the same observed-run/handled-without-run classification as an earlier ack. Preserve an already completed receipt and its effects; do not relabel completion as interruption. Specify the no-run, naturally completed, failed and genuinely interrupted outcomes of the stop itself. With the current contract, any case lacking evidence for the required interruption must remain uncertain with admission closed and force stop available, unless a separately reviewed contract change supplies an honest reconciliation outcome. Do not silently borrow revocation's `input-reconciled` value.
**Required fixtures:** successful completion between clear and abort; late handled ack after the fence with no run; preflight error after the fence with no run. Assert both the request receipt and the stop proof/state. Include a mutant that substitutes idle + empty queue for interruption evidence and ensure it fails. A response to this finding may choose a conservative interim instead of expanding CHAT-01.
## 2. Not blocking — bound the ack-without-start explanation to complete observations
The N11 `failed` result is sound for the stated narrow case: positive run-failure evidence, settlement for that run, and a complete ordered native stream with no user-message start/end. The native run emits user start and end before persistence; `agent-session.js:386–398` persists the user message only after notifying listeners. The core loop also emits the initial user events before entering the model loop. No additional native user persistence path was found in the inspected code.
However, “events reached the controller first” is stronger than the cited function establishes: it invokes local listeners first. It does not itself prove successful delivery to a separate controller. Qualify the explanation accordingly. Missing output, parse failure, lost transport or missing settlement must retain uncertainty, as section 1 already requires. `failed` also is not proof that an input/before-agent-start extension had no external effects; keep effects accounting separate and do not present automatic replay as safe.
This is a wording/evidence qualification, not a request to restore text identity or queued-Mosaic-input machinery.
## R1 disposition and the remaining design
| Prior finding | R2 assessment |
|---|---|
| 1: two-key claims and crash publication | Closed for the specified cooperative protocol. Complete immutable publication, claim IDs, conservative pair state and live-owner checks supply the missing mechanism. W5 explicitly stops short of claiming a power-loss test. |
| 2: pending dispatch and unknown writes | Closed. Reserving the slot before the write and poisoning unknown transport outcomes prevents the second dispatch and unsafe retry scenarios. Slot settlement still needs finding 1 above when used to build an interrupt proof. |
| 3 plus addendum: native queues and fence | The original text-identity scenario is retired correctly. External queue clearing, bounded late preflight handling and the post-clear limit are honest. Interrupt reconciliation remains open only as finding 1 above. |
| 4: cohort observation and containment | Closed at brief level within the expressly fixture-only guarantee. The shim, invocation identity, freeze/enumeration/kill/empty observation and missing-is-not-empty rule are concrete. K13 must actually pass before real cohorts receive `stopped`; otherwise the stated fallback is uncertainty. This does not certify the unbuilt implementation or live trust. |
| 5: restart dedup | Mechanism closes the replay scenario. V-1 honestly names the contract deviation and gates I1 approval on Sage's ruling; this review does not substitute for that ruling. |
| 6: unavailable stream entry IDs | Closed within the stated drift-detection limits. Idle `get_entries` comparison replaces the nonexistent stream key, and delayed persistence, unchanged-ID rewrites and non-cooperating writers are disclosed. |
| 7: recovery eligibility | Closed at brief level. A retained reservation, single-use eligibility and claim/leaf/pin revalidation address double launch and changed-leaf races. |
**Spawn marker / W20:** fine and deliberately conservative. A marker with no unit cannot establish whether an engine ran and escaped observation. Holding it until same-host boot proof avoids inventing a death proof. A marker visible on either half of a pair must retain that classification when the pair is completed. Foreign-host and live-owner refusals are also correct. This rule may strand a fixture reservation until reboot; that is the documented availability tradeoff, not a safety defect to work around.
**C-1 interim:** strict enough and acceptable. Non-empty clear means no reconciled transition, admission closed, force stop available. There is no need to relax it for CHAT-03. It does not solve finding 1, whose schedules can have entirely empty clears.
**Cohort restart:** refusing when the recorded invocation cannot be checked is correct, including a crash before identity publication. A boot-proof implementation must preserve the declared contract/verifier boundary rather than fabricate a freeze observation or member death times; assess its exact evidence representation in the code review. Same-uid interference and live producer/verifier trust remain explicitly deferred.
## Evidence and limits
Reviewed the pinned brief/request, previous Rocko report/addendum, CHAT-01 schema/checker and the installed Pi session/RPC/core-loop paths. Rechecked both candidate hashes and the cited session source hash. Ran the small idle-abort method check described above. No candidate implementation exists, so no H/N/K tests or repository suite pass is claimed. Only this report was written.
@@ -0,0 +1,48 @@
# CHAT-03 R3 — final adversarial pass
Verdict: **one blocking finding remains; return to Sage for the decision-27 scope cut, not R4.**
Rocko. Target pins verified before reading and again at completion:
- BRIEF.md: `2c5be6b4b2caddf9314e8fdcc9108d744b770fe8f469e9c2d410ee6d14c6b1ec`.
- REVIEW-REQUEST.md: `e197b8822df1d04b4c89ee9873855525db3880258cbcdd0eafa2ffd2ef689823`.
- Stated base: `f2b9e622`.
Scope incorporates Sage's decision 30: no goal in mediated seats; Pi dialogs, Pi H5–H8, C-1/C-2/C-4 are cut; idle drift moves to CHAT-07 and C-5 to CHAT-04. None of those cuts is treated as a blocking omission. Claude permission races remain in scope. No source, contract, live engine or service was changed.
## 1. Blocking — the retained seal hashes an extension directory, not necessarily the code it executes
Section 3 permits a local explicit extension if every file in its directory tree is hash-pinned and it has an input-silent review record. It does not require executable imports to stay in that tree or in another pinned, reviewed artifact. The installed loader (`dist/core/extensions/loader.js:409–428`) uses `jiti.import`; a local extension can import a sibling outside its tree. Refusing npm/git *extension arguments* does not refuse module imports from a local entry point.
Concrete schedule, without a hostile actor or goal:
1. `extension/index.ts` imports `../helper.ts` and invokes `install(pi)` from it. At review time the helper does nothing; the extension is input-silent.
2. A subsequent helper edit adds `pi.sendMessage(..., { deliverAs: 'nextTurn' })`. Nothing in the registered extension directory changes.
3. The next bind has all three required flags, the same local entry path, matching registered tree hashes and the existing review record. N24's checks all pass, but Pi executes code different from the reviewed code.
4. The helper can queue input invisible to clear/abort and attach it to a later Mosaic prompt. As N23 correctly documents, O1–O6 need not fire. Receipt and stop attribution nevertheless name the seal as their basis.
I reproduced the import/hash part using the installed `jiti/static` importer, a temporary directory and a stub `sendMessage` API. The entry imported a sibling helper. Before the helper edit, zero calls; afterwards, one `nextTurn` call; the hash of the sole file in the registered extension directory was unchanged. Module and filesystem caches were disabled for the check. No Pi engine, model, network request or live session was involved; the temporary files were removed.
Limits 7 and 11 honestly admit missed signals when the seal fails. They do not make this a sufficient binding precondition: here a correct review becomes stale without any specified hash check failing. This is an unpinned executable dependency, not merely a reviewer incorrectly approving the bytes subsequently loaded. The same issue includes dependencies reached through symlinks outside the registered tree.
**Recommended scope cut:** defer arbitrary explicit extensions. Permit only a fixed, reviewed fixture set whose executable dependencies are entirely inside the pinned artifact(s), with no escaping imports, dynamic code loading or external symlink targets. An empty explicit-extension list is the smallest starting scope now that Pi dialogs and goal are out. Keep the mandatory built-in extension tied to the pinned Pi artifact and its reviewed dependency code. Do not add a general dependency crawler or new run-proof machinery for CHAT-03. If a fixture extension is retained, its acceptance evidence must show that changing code it imports cannot leave the binding accepted under the old seal.
This is in the retained sealed-engine section and therefore remains blocking after decision 30. Sage owns the cut; this report requests no fourth brief round.
## 2. Note — keep the no-turn fence local to its interrupt
The new no-turn refusal is a reasonable narrowing and avoids needless stop records. At implementation review, verify that its fence cleanup cannot reopen admission closed by a concurrent force stop, overlap signal or revocation. A useful schedule is: interrupt begins, force stop supersedes or an overlap signal closes admission, then the interrupt finds no slot/run and returns `no-turn`. Admission must remain closed under the surviving reason. H10 and the general dispatch gates already require that behavior; add this branch to their fixture coverage. This is not a new contract requirement or a blocking brief defect.
## Prior findings and retained behavior
**My R2 blocking finding is closed.** Rules 4 and 5 now separate the item's receipt from the stop outcome. N14 preserves normal completion, N15 covers a late handled ack without a run, and N3 covers preflight failure. Only an observed aborted run, with complete ordered evidence and no overlap signal, can support `turnState: interrupted`. Until C-5, the other outcomes retain uncertainty and closed admission. Moving C-5 to CHAT-04 makes this more expensive operationally, but does not make it dishonest.
**My R2 persistence note is closed.** N11 now retains `delivery-unknown`; the text distinguishes local event/persistence ordering from prompt attribution and from absence of extension effects. It does not promise to detect a silently lost line.
**O1–O6 are useful alarms, not independent proof of a seal.** The disclosed late-attribution windows in N8/N19/N21 and invisible queues in N22/N23 are honest. Their tests must assert the conservative presentation as well as admission closure. They cannot compensate for finding 1. With genuinely fixed input-silent code, one pending slot and the native response ID provide the stated basis for attributing the subsequent run by exclusion. Removing goal is consistent with that basis.
**Control and recovery remain acceptable at brief level.** The slot is reserved before a write; unknown writes poison the pipe; restart tokens forbid replay; a marker on either claim key prevents no-unit release; foreign-host and live-owner refusals remain conservative. Force stop still requires independent cohort observation rather than engine idle/settled, checks invocation identity, and treats absent evidence as uncertain. Recovery retains both claim keys and single-use eligibility. K13 and the producer/verifier limitations remain explicit acceptance gates, not established implementation facts. The section-6 change from R2 chiefly incorporates the corrected interrupt outcomes; it does not weaken the cohort rules.
## Evidence and limits
Read the pinned R3 brief/request, prior Rocko and Filbert findings, the retained Pi protocol/session paths, and the installed extension loader. Ran the isolated importer/hash reproducer described above and rechecked both packet hashes. This is an unbuilt brief: no claim is made that the proposed N/H/K tests or repository suites have run. Only this report was written in the checkout.