Files
stack/agents/rocko/work/chat-03-r2-adversarial-2026-09-26.md
T

8.5 KiB
Raw Blame History

CHAT-03 R2 — control, stop and recovery adversarial review

Verdict: revise. Rocko, 2026-09-26. One blocking finding.

Verified candidate pins before review and again at completion:

  • BRIEF.md: 5c5b45a277f3a555337a5caf57ff1b300b95781b69640ec8974770bdd44af9bd.
  • REVIEW-REQUEST.md: c75ad86fc014b7a13f72132c0378d6e8abf09663820eee701dc67e3ed50ef92f.
  • Stated base: 40a02d2b.
  • Installed agent-session.js matches the packet: fb8a3981c20c8c0bbd42231b1c99a10335fb3858b659056b341954de9cfa467f.

Scope: sections 5–6, their section 2/3 dependencies, and disposition of my R1 findings and addendum. This is a specification review, not an implementation certification. No source, live session, cgroup, service, configuration or contract was changed.

1. Blocking — settling the slot does not establish an interrupted turn

The slot table at section 3 lines 591–595 assumes that a late preflight ack means a run started, and that a visibly started run follows an interrupted turn. Rule 5 then permits reconciliation after a settled slot, native idle and empty clears. Those observations do not distinguish interruption from ordinary completion or no run.

Concrete schedules:

  1. A run has emitted its user message. Interrupt fences admission and sends clear_queue. While that exchange is in flight, the run completes normally, emits its successful assistant result and settles. The controller then sends abort against an idle engine. A subsequent get_state is idle and the post-settle clear is empty. Every clear can be empty, so C-1 does not catch this. The table says turnState: interrupted, despite the successful completion.
  2. Interrupt fences a prompt still in preflight. An input extension returns handled after the first clear/abort. Pi acknowledges it without starting a run. The first table row instead says “Ack: a run started.” Repeating clear/abort cannot manufacture a run or an interruption. Likewise, a preflight error can settle the slot as failed without there ever being a turn to interrupt.

This is a contract boundary, not just a missing test name. docs/plans/chat-01/check.mjs:315 accepts reconcile-interrupt only with turnState: interrupted; line 325 records turn-interrupted. input-reconciled is reserved by that checker for revocation. A fixture that always aborts an active fake run can pass N3/N4/N9 while the real schedules above remain unrepresented.

Native evidence: agent-session.js:832–847 acknowledges extension-handled input without _runAgentPrompt; abort() at 1222 only aborts current operations and waits for idle. It does not certify that a run was interrupted. I invoked that installed abort method on a minimal idle receiver: it resolved with no run events. That is a method-level check, not a real-engine smoke or model call.

Required fix: separate receipt settlement from interrupt-proof construction. Route a late ack through the same observed-run/handled-without-run classification as an earlier ack. Preserve an already completed receipt and its effects; do not relabel completion as interruption. Specify the no-run, naturally completed, failed and genuinely interrupted outcomes of the stop itself. With the current contract, any case lacking evidence for the required interruption must remain uncertain with admission closed and force stop available, unless a separately reviewed contract change supplies an honest reconciliation outcome. Do not silently borrow revocation's input-reconciled value.

Required fixtures: successful completion between clear and abort; late handled ack after the fence with no run; preflight error after the fence with no run. Assert both the request receipt and the stop proof/state. Include a mutant that substitutes idle + empty queue for interruption evidence and ensure it fails. A response to this finding may choose a conservative interim instead of expanding CHAT-01.

2. Not blocking — bound the ack-without-start explanation to complete observations

The N11 failed result is sound for the stated narrow case: positive run-failure evidence, settlement for that run, and a complete ordered native stream with no user-message start/end. The native run emits user start and end before persistence; agent-session.js:386–398 persists the user message only after notifying listeners. The core loop also emits the initial user events before entering the model loop. No additional native user persistence path was found in the inspected code.

However, “events reached the controller first” is stronger than the cited function establishes: it invokes local listeners first. It does not itself prove successful delivery to a separate controller. Qualify the explanation accordingly. Missing output, parse failure, lost transport or missing settlement must retain uncertainty, as section 1 already requires. failed also is not proof that an input/before-agent-start extension had no external effects; keep effects accounting separate and do not present automatic replay as safe.

This is a wording/evidence qualification, not a request to restore text identity or queued-Mosaic-input machinery.

R1 disposition and the remaining design

Prior finding R2 assessment
1: two-key claims and crash publication Closed for the specified cooperative protocol. Complete immutable publication, claim IDs, conservative pair state and live-owner checks supply the missing mechanism. W5 explicitly stops short of claiming a power-loss test.
2: pending dispatch and unknown writes Closed. Reserving the slot before the write and poisoning unknown transport outcomes prevents the second dispatch and unsafe retry scenarios. Slot settlement still needs finding 1 above when used to build an interrupt proof.
3 plus addendum: native queues and fence The original text-identity scenario is retired correctly. External queue clearing, bounded late preflight handling and the post-clear limit are honest. Interrupt reconciliation remains open only as finding 1 above.
4: cohort observation and containment Closed at brief level within the expressly fixture-only guarantee. The shim, invocation identity, freeze/enumeration/kill/empty observation and missing-is-not-empty rule are concrete. K13 must actually pass before real cohorts receive stopped; otherwise the stated fallback is uncertainty. This does not certify the unbuilt implementation or live trust.
5: restart dedup Mechanism closes the replay scenario. V-1 honestly names the contract deviation and gates I1 approval on Sage's ruling; this review does not substitute for that ruling.
6: unavailable stream entry IDs Closed within the stated drift-detection limits. Idle get_entries comparison replaces the nonexistent stream key, and delayed persistence, unchanged-ID rewrites and non-cooperating writers are disclosed.
7: recovery eligibility Closed at brief level. A retained reservation, single-use eligibility and claim/leaf/pin revalidation address double launch and changed-leaf races.

Spawn marker / W20: fine and deliberately conservative. A marker with no unit cannot establish whether an engine ran and escaped observation. Holding it until same-host boot proof avoids inventing a death proof. A marker visible on either half of a pair must retain that classification when the pair is completed. Foreign-host and live-owner refusals are also correct. This rule may strand a fixture reservation until reboot; that is the documented availability tradeoff, not a safety defect to work around.

C-1 interim: strict enough and acceptable. Non-empty clear means no reconciled transition, admission closed, force stop available. There is no need to relax it for CHAT-03. It does not solve finding 1, whose schedules can have entirely empty clears.

Cohort restart: refusing when the recorded invocation cannot be checked is correct, including a crash before identity publication. A boot-proof implementation must preserve the declared contract/verifier boundary rather than fabricate a freeze observation or member death times; assess its exact evidence representation in the code review. Same-uid interference and live producer/verifier trust remain explicitly deferred.

Evidence and limits

Reviewed the pinned brief/request, previous Rocko report/addendum, CHAT-01 schema/checker and the installed Pi session/RPC/core-loop paths. Rechecked both candidate hashes and the cited session source hash. Ran the small idle-abort method check described above. No candidate implementation exists, so no H/N/K tests or repository suite pass is claimed. Only this report was written.