Files
stack/agents/rocko/work/queue-as-data-adversarial-2026-09-26.md
T
jason.woltjeandClaude Opus 5.5 1c5f6bc3a0 docs(queue): queue-as-data plan round 6, Rocko approved (#1508)
Plan 282fabbb (Filbert) and the six adversarial rounds (Rocko, r6 80cde839).
Lead item 15: Gate F first, then A1 and A2 as separate reviewed commits.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-26 16:05:16 -05:00

404 lines
22 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Queue as data: adversarial plan review
VERDICT: revise
Rocko, 2026-09-26, for Sage / #1508. Read-only investigation; this report
is the only file written for this assignment. No commits, issue writes,
credentials read, live launches, or source changes. This is a design review,
not a claim that unimplemented queue code has failed tests.
Reviewed Filbert's plan after the Q5–Q9 decisions and the final section-5
proposal/decision distinction were folded in:
- `agents/filbert/work/queue-as-data-plan-2026-09-26.md`, SHA-256
`59d5a22f1f67426237761341b686c3a9dcf0dffb2d353bab8dd2f9741bcc1fdb`.
- Brief `docs/plans/2026-09-13_queue-as-data.md`, SHA-256
`12d8d39f95f40e7d1e93f3a4420390152bbf52d04a6596b48d443512ce559061`.
- HEAD `21e3e908b6f64d3eba202480fede05e174799e1d` plus current shared
working files: AGENTS, QUEUE, scripts/mosaic, seat, control-board, ledger,
launcher and Gitea helper sources. No `packages/queue` exists yet.
High means a credible loss of state, authority bypass, or false acceptance;
medium means a material workflow or evidence defect. Findings against decided
items request reconsideration or clarification by Sage; they do not silently
reverse those decisions. Q1–Q4 remain with Jason.
## Findings
### R1 — High: the two-file write has no crash/recovery contract
**Where:** plan 2.A and 3.A, atomic write under a lock; test kills only before
rename.
**Scenario:** A reads revision 10, installs revision 11 of queue.json, then
dies before QUEUE.md is rendered. JSON readers and humans now select different
work. Separately, a standalone renderer can read revision 10 before B's write
and replace the Markdown with revision 10 after B has installed revision 11.
A timeout-based stale-lock remover can also admit B while a slow A is alive.
Atomic rename prevents a torn individual file, not these outcomes.
**Fix:** specify one canonical lock identity across the checkout and its
compatibility symlink; acquire it before reading, not merely before writing.
All mutations and standalone render use it. Generate both outputs from one
validated revision; record a durable transaction/revision and define recovery
for every interruption point. Readers must detect an incomplete pair rather
than treating it as current. Prefer kernel-released locking or provable stale
owner detection; do not steal a lock because a deadline elapsed. Bound waits.
Document fsync behavior if surviving a host crash is promised.
**Acceptance:** race add/add, move/note and render/write through both checkout
paths; kill after each file replacement and before/after commit; inject disk
errors and a dead lock holder. Recover one acknowledged operation exactly
once, or refuse with a concrete recovery instruction. Never erase a completed
neighbor's update.
### R2 — High: Q3's path-limited commit still captures another author's work
**Where:** Q3; AGENTS invariant 8.
**Scenario:** someone has an uncommitted change to QUEUE.md's hand-written
header or historical footer. `git commit -- queue.json QUEUE.md` includes
that entire file, not just the generated table. A concurrent ordinary commit
can also capture the queue between its write and auto-commit: it does not
honor the queue lock. An index-lock or hook failure leaves already-mutated
queue state without the promised commit, and an unsafe rollback could erase
a later writer. Git's index lock is not a lock over the whole operation.
**Fix:** default auto-commit off until Q3 is settled. If authorized, require
clean queue paths relative to the intended baseline, no conflicting staged
versions, no merge/rebase, and the authorized branch; serialize *all* relevant
commit writers through an agreed repository protocol. Freeze and validate
the exact prospective file contents, check HEAD/revision before committing,
and define retry/recovery without reset/checkout of others' changes. A queue
lock alone cannot enforce cooperation by arbitrary host git commands.
The statement that a validator means invariant 8 holds is false: schema
validation is not the applicable suites. Identify the required checks and
bind their receipt to the actual prospective content/code version. Absence
of a shell suite wrapper does not exempt existing Node package tests.
**Acceptance:** pre-dirty header, partially staged queue file, unrelated
staged files, competing ordinary commit, hook failure, and changed HEAD.
Verify both the committed tree and preservation of the other author's index
and working changes. No push authority follows from any local commit.
### R3 — High: single-writer checks validate appearance, not provenance
**Where:** 5.9, 3.A, migration and `render --check`.
**Scenario:** hand-edit a row from queued to done, run render, and both schema
validation and `render --check` pass. A Markdown hand edit is silently erased
by the *next successful CLI mutation*, before the suite ever checks it.
A merge can produce valid JSON with illegal transitions or removed rows;
two clones can independently allocate the same new id. A fresh clone has no
local lock installation or uncommitted audit trail and cannot reconstruct
the claimed writer from a matching table.
**Fix:** check for unexpected drift *before* mutating or rendering, refuse
ambiguous markers and refuse to overwrite unexplained manual changes. Define
a reviewed genesis/migration and validate each subsequent operation against
its parent revision, including immutable IDs, deletions and authority rules.
Use a committed operation history or another explicitly authorized receipt
scheme, with merge reconciliation and fresh-clone validation. Keep a high
water mark/tombstones for retired IDs, including removed row 7. If multiple
clone writers are out of scope, refuse their integration until reconciled;
a per-checkout lock does not solve that problem.
Be honest about the boundary: code running as the same host user cannot
prevent that user from rewriting both data and receipts. Without an enforced
integration boundary, this is a cooperative writer with drift detection,
not tamper-proof authority. Hooks alone would not fix fresh clones and are
outside Piece A's brief anyway.
**Acceptance:** valid-but-illegal hand-edited JSON with regenerated Markdown,
deleted highest id, duplicate-id branch additions, conflicting merge,
fresh clone, missing/duplicate render markers, and preexisting Markdown drift.
### R4 — High: identity attribution is still being used as authorization
**Where:** Q5, 3.A authority tests, 5.3 and 5.8; brief iron-clad point 5.
**Scenario:** a seat runs `--by jason` or sets `MOSAIC_AGENT_NAME=sage` and
moves a protected row. If assign lacks the move permission check, it can
instead assign another seat's row to itself and then move it. Even without
spoofing, the proposed rule only restricts *clearing* required, while the
brief says only Jason *moves* required rows. These are different rules.
**Fix:** separate claimed attribution from verified authority. Q5 can map
the new lead's queue role, but does not turn an environment variable into
proof or grant managed-worker policy authority. Specify the permitted actor
for every mutating verb/field, including assign, dependencies, required,
gate ownership, notes and terminal rows. Either retain an explicit reviewed
human gate for protected mutations or have Jason accept the limited
cooperative trust model; weekly reporting is detection after the fact, not
enforcement. Reconcile the exact required-row exception before building.
**Acceptance:** forged identity, reassignment bypass, dependency/gate edits
by an ordinary seat, edits to done/parked rows, and every operation on a
required row. No roles/*.json expansion is implied by this review.
### R5 — High: Gate G can appear to pass with extra human help or no work
**Where:** Gate G setup and pass evidence; ledger.mjs `messageKind` (line 78)
and `readSessions` (line 92); agent-host-dev.sh context and fresh handling.
**Scenario:** Jason supplies a correction through the board or a lead relays
it. The ledger counts it as board/agent, leaving exactly one human message.
`--fresh` creates a new conversation but preserves old session files and
loads shared SOUL/USER/CONTEXT/skills; a task-specific hint there can make
the demo pass without discovering the assignment. Another seat can move
the row using the test seat's claimed identity. Reading the brief and
moving the row alone is also not evidence of actually starting its work.
**Observed read-only probe:** `messageKind` returns human for the initial
instruction, board for `[dragon-lin:control-board -> dragon-lin:rocko]`
followed by a human correction, and agent for a similarly prefixed Sage
relay. The new `[from: sage (...) -> to: rocko (...) class=actionable]`
header is classified human. Table 2 aggregates a seat's sessions over the
date range, not an identified Gate G session. It is unsuitable as the
sole acceptance counter in either direction.
**Fix:** pin the exact fresh session id, initial context/brief/queue
revisions, test interval and command output. Audit *all* input channels in
that interval; distinguish the initial instruction from further human
input regardless of routing. Preserve generic governance context but
exclude task-specific pre-coaching. Require a concrete first authorized
work action tied to the returned row and that session. Jason's observed
pass/fail remains authoritative; do not infer it from the row or git author.
Allow normal prerequisite reads instead of requiring the literal first
filesystem read to be the brief. A commit is not necessary proof of G and
must not bypass Q3 or suite requirements.
**Acceptance:** negative controls with a board correction, relayed hint,
resumed session and row moved by a helper must fail. Use unit tests with
competing eligible rows to test selection: the single-row demo cannot.
### R6 — High: dependencies and `next` do not prevent unauthorized starts
**Where:** Q8, section 1, 3.A and 5.11.
**Scenario:** `next` skips a dependent row, but the seat directly runs
`move ID in-progress` and bypasses the filter. Two sessions of the same
seat both read `next`, then both start the same work; a state of in-progress
does not distinguish resume by the owner session from a second claimant.
An in-review row returned to its author may cause premature continuation;
a reviewer can repeatedly receive a row it already reviewed if its receipt
is not represented.
**Fix:** enforce prerequisite and authority predicates at transition time
under the lock, not only in selection. Define atomic claiming or an explicit
single-active-session assumption with detectable conflict. Return the action
role (implement, resume, review, wait) and track review round/receipt so
review dispatch has a stopping condition.
There are unresolved semantics: Q8's done-only dependencies cannot express
the brief's row-6-*blocked* start exception. State-priority ordering differs
from the existing first-eligible-row cadence. `after` alone cannot express
every arbitrary owner priority, especially between active rows. Decide these
once and update the brief/cadence together. Test that required work cannot
be subordinated to ordinary queued work through dependencies.
### R7 — High: posting then recording cannot promise exactly one review request
**Where:** 2.D, 3.D, 5.14; scripts/gitea-api.sh.
**Scenario:** Gitea accepts a comment but the response is lost, or the CLI
dies before storing its id. The row remains unchanged and retry posts a
second request. Conversely, a posted request survives a failed local commit.
A repository lock cannot roll back the remote side effect. The existing
helper supplies an ordinary POST, not an exactly-once transaction.
**Fix:** define a durable operation id containing row, revision, review
round and candidate digest; persist intent before sending. Include the id
in the request, reconcile uncertain outcomes by lookup/readback, and refuse
blind retries if delivery cannot be determined. Specify outbox/recovery
state and bounded lock behavior. An authorized token path also needs an
expected account check; checking that the path differs from the default is
not proof of the intended poster. Do not read or provision tokens before Q4.
**Acceptance:** accepted POST with dropped response, crash before receipt,
duplicate retry, concurrent moves, stale candidate and local commit failure.
Fake success/failure alone does not test these boundaries.
### R8 — High: missing from the open list does not establish closed
**Where:** Q7 and 2.E.
**Scenario:** a row references a typo/nonexistent issue, a deleted issue, a
PR number or an inaccessible item. It is absent from a complete open-issues
response and the proposed algorithm declares it closed. Even a short page
only shows that that response is short; it does not prove each referenced
number is a known issue. A done row can therefore pass with no closure
evidence.
**Fix:** retain the additional open-list call and full-page refusal, which
are useful, but represent absent as unknown unless there is positive issue
existence/type/state evidence. Fetch/cache positively identified referenced
issues under a revised call budget, or report incomplete evidence instead
of success. Bound cache age and detect changed scope. Specify `--no-issues`
as unknown queue issue checks, not zero violations.
**Acceptance:** nonexistent/deleted/PR references, denied access, malformed
responses, a full page and no-issues mode. This is a requested correction
to decided Q7, not an instruction to silently implement another fetch policy.
### R9 — High: the liveness exemption can turn missing evidence into green
**Where:** Q6, 2.E and 3.E; scan.mjs lines 272–307.
**Scenario:** remove a native seat's registration and it gets the same
non-violation label as a T3 seat. `loadRegistrations` returns records plus
errors; it does not check liveness and loads multiple layouts. Matching a
seat name alone can select a fleet/other-workspace record. A dead or reused
PID is not evidence of the intended active session. Corrupt records must
not become the missing-registration exemption.
**Fix:** preserve Q6's prohibition on T3 probing, but distinguish explicitly
unsupported runtime, missing expected registration, invalid record, stale
record and live matched record. Match repository/layout/session identity.
Propagate scanner errors. Show coverage/unknowns alongside violations and
do not call E fully passed on zero violations when liveness was untested.
Record the owner's explicit acceptance of a temporary reduced gate.
**Acceptance:** missing native registration, corrupt record, duplicate seat
across layouts, wrong workspace, dead PID, unknown PID and an exempt T3 seat.
### R10 — Medium: history needed for E is absent from the schema
**Where:** 5.3, 5.8 and 5.12; brief E's age and same-day remediation rules.
**Scenario:** `queue note` refreshes updatedAt every week, making a
long-neglected required row appear young. A row is changed twice between
ledger runs and updatedBy remembers only the last writer, so E cannot list
every protected change as 5.8 promises. Removing a row and having no delete
verb does not prove an id was never reused. A single latest timestamp also
cannot prove each observed violation was remedied within the day.
**Fix:** define stable createdAt/requiredSince, operation history and audit
retention, with explicit time zone and first-observed violation evidence.
Do not depend on optional Q3 commits for mandatory checks. If a journal is
chosen, make corrections new entries and place it in a dedicated directory;
do not rewrite BUILD-LOG, SESSIONS or existing append-only runtime logs.
Terminal required rows should not age forever as outstanding work.
### R11 — High: some section-5 fixes change owner requirements, not details
**Where:** 5.6, 5.7, 5.10, 5.13 and 5.14; brief A/C/D and iron-clad point 3.
**Scenario:** the builder follows “settle in A's round” and makes brief
optional, reopens parked rows, allows completion without Jason, or changes
the issue/gate semantics. Those are externally visible requirements changes.
The brief explicitly demands an existing brief for every row and says
parked/done rows never change. A queued row can have an existing planning
brief not yet accepted; a missing brief is not technically necessary to
make queued reachable.
**Fix:** put each unresolved behavioral change in a disposition table:
original rule, proposed rule, approving authority, acceptance change. Keep
already-decided changes separate. Do not conflate source review approval
with owner approval of altered gates. Decide the blanket `any→blocked`
versus terminal immutability conflict and block→block semantics too.
The multi-row issue problem is real, but “open iff at least one row not
done” is only sound if rows exhaust the issue's scope. Row 1/22 explicitly
separates broader MVP acceptance from one targeted correction; row 6 has
two issue phases. Model which row gates issue closure, or report mismatch
for disposition without forcing automatic closure. A single reviewer field
also loses row 6's independent Darkwing/Dewey reviews; use a review policy
that can preserve both if that row is migrated intact.
### R12 — Medium: immutable review candidate and receipt lifecycle are missing
**Where:** 5.14 and Piece D pass tests.
**Scenario:** a request posts `/tmp/candidate.sha256`; the file is replaced
or disappears, so a later reviewer cannot identify what was approved. This
takeover already found the historical #1511/#1512 `/tmp` snapshots absent.
With `issues: [int]`, “the piece's issue” also no longer selects an unambiguous
posting target. Posting a request alone does not establish a full review
round or preserve its verdict when the row loops through review again.
**Fix:** resolve and store the candidate content digest and durable commit
or artifact reference, not just a path verbatim. Name the request issue,
review round, required reviewers and receipt links; preserve prior rounds.
Test request→changes requested→new candidate→approval against exact pins.
Keep the brief's ban on new review files, but evaluate the stronger proposed
ban separately: no-files-anywhere is not a substitute for durable receipts.
### R13 — Medium: brief validation and migration are too shallow
**Where:** 3.A migration/C and 5.12; current QUEUE rows and parked table.
**Scenario:** a directory or symlink outside the repository “exists” as a
brief, or an untracked local brief makes the author's tests pass but is absent
in a fresh clone. A nonexistent section passes even though the fresh seat is
instructed to open only the named brief section. Migration shortens a state
cell and drops an acceptance/publication boundary that is not actually
duplicated in CURRENT or an issue. A valid golden Markdown file proves none
of those semantics.
**Fix:** require a readable repository-contained regular brief, a defined
symlink policy, and tracked prospective content for publishable rows. Either
validate section anchors or disallow them for actionable rows until resolved.
Produce a reviewed row-by-row migration map preserving owners, review lanes,
gates, required status, dependencies, historical IDs and each unresolved
boundary. Map parked entries with no briefs explicitly rather than inventing
authority from their free text. Preserve history through exact durable links
or a reviewed archive; do not assume it was copied elsewhere.
### R14 — Medium: package reuse and acceptance wiring need concrete decisions
**Where:** 2.E, 3.B, section 1; package.json and existing module layout.
**Scenario:** ledger imports `@mosaic/queue` but adding a sibling package.json
does not install a workspace package. Root package.json declares no workspaces;
a read-only `import.meta.resolve('@mosaic/seat')` probe returns
`ERR_MODULE_NOT_FOUND` today. Moving the registration loader into seat also
touches a package outside E's listed files and requires its regression tests.
**Fix:** use the existing relative-module convention or explicitly plan
package linking and its installation/lockfile consequences; do not leave
it implicit. Reuse the loader without silently expanding ownership, and
specify liveness/error handling separately (R9). Name actual package test
commands and affected suites. C also needs the promised two owner-accepted
briefs, not just missing-path tests. D needs a full review round and E the
dated ledger receipt/remediation evidence. B's `--check` verifies inputs
exist, not that a running or fresh session received the new cadence; inspect
the launch context and then perform G. Clarify the brief's “A to E in order”
against the proposed C-with-A and D/E parallel delivery before dispatch.
## What is fine
- Repository-local `scripts/mosaic` dispatch, keeping seat argv/env behavior,
and tests for that behavior are appropriate. Q1 correctly remains an owner
wording choice; no global CLI or fleet edits are needed.
- Queue data under docs/plans and code under packages/queue respect
bootstrap-only root. An edit to the existing root AGENTS.md is permitted;
a new root lock/journal/config file would not be. The plan proposes no
contract, secret, run-record or managed-role changes.
- Required as a flag, an explicit review-to-work edge, strict unknown-field
validation, dependency cycle checks, deterministic rendering and a
non-mutating render check are good building blocks. They need the
transaction/history/authority rules above to deliver the claimed result.
- Missing identity refusal, Q4 held for explicit credentials, no automatic
push, no T3-backed product feature, and preservation of the old table log
are sound boundaries. Historical log corrections must remain additions.
- Combining migration with the introduction of the authoritative data file
avoids a deliberate two-master interval. A reviewed migration map still
matters. Reusing one validator is correct once module resolution is explicit.
## Recommended disposition
Revise the write/recovery/commit protocol (R1–R3), settle protected authority
and requirement changes (R4/R11), strengthen Gate G evidence (R5), and specify
transition claims and review delivery recovery (R6/R7) before implementation.
Correct E's positive-evidence rules and audit model (R8–R10); complete the
candidate, migration and packaging details (R12–R14). These are bounded plan
changes, not authorization to build a new policy service or alter live state.
Validation for this review was source inspection and read-only in-process
probes of the existing message classifier and module resolution. No queue
implementation or fault-injection suite was run because none exists yet.
This report records the session outcome in place of editing SESSIONS.md,
as this assignment explicitly permits writes only to the report.