Plan 282fabbb (Filbert) and the six adversarial rounds (Rocko, r6 80cde839). Lead item 15: Gate F first, then A1 and A2 as separate reviewed commits. Co-Authored-By: Claude Opus 5.5 <[email protected]>
22 KiB
Queue as data: adversarial plan review
VERDICT: revise
Rocko, 2026-09-26, for Sage / #1508. Read-only investigation; this report is the only file written for this assignment. No commits, issue writes, credentials read, live launches, or source changes. This is a design review, not a claim that unimplemented queue code has failed tests.
Reviewed Filbert's plan after the Q5–Q9 decisions and the final section-5 proposal/decision distinction were folded in:
agents/filbert/work/queue-as-data-plan-2026-09-26.md, SHA-25659d5a22f1f67426237761341b686c3a9dcf0dffb2d353bab8dd2f9741bcc1fdb.- Brief
docs/plans/2026-09-13_queue-as-data.md, SHA-25612d8d39f95f40e7d1e93f3a4420390152bbf52d04a6596b48d443512ce559061. - HEAD
21e3e908b6f64d3eba202480fede05e174799e1dplus current shared working files: AGENTS, QUEUE, scripts/mosaic, seat, control-board, ledger, launcher and Gitea helper sources. Nopackages/queueexists yet.
High means a credible loss of state, authority bypass, or false acceptance; medium means a material workflow or evidence defect. Findings against decided items request reconsideration or clarification by Sage; they do not silently reverse those decisions. Q1–Q4 remain with Jason.
Findings
R1 — High: the two-file write has no crash/recovery contract
Where: plan 2.A and 3.A, atomic write under a lock; test kills only before rename.
Scenario: A reads revision 10, installs revision 11 of queue.json, then dies before QUEUE.md is rendered. JSON readers and humans now select different work. Separately, a standalone renderer can read revision 10 before B's write and replace the Markdown with revision 10 after B has installed revision 11. A timeout-based stale-lock remover can also admit B while a slow A is alive. Atomic rename prevents a torn individual file, not these outcomes.
Fix: specify one canonical lock identity across the checkout and its compatibility symlink; acquire it before reading, not merely before writing. All mutations and standalone render use it. Generate both outputs from one validated revision; record a durable transaction/revision and define recovery for every interruption point. Readers must detect an incomplete pair rather than treating it as current. Prefer kernel-released locking or provable stale owner detection; do not steal a lock because a deadline elapsed. Bound waits. Document fsync behavior if surviving a host crash is promised.
Acceptance: race add/add, move/note and render/write through both checkout paths; kill after each file replacement and before/after commit; inject disk errors and a dead lock holder. Recover one acknowledged operation exactly once, or refuse with a concrete recovery instruction. Never erase a completed neighbor's update.
R2 — High: Q3's path-limited commit still captures another author's work
Where: Q3; AGENTS invariant 8.
Scenario: someone has an uncommitted change to QUEUE.md's hand-written
header or historical footer. git commit -- queue.json QUEUE.md includes
that entire file, not just the generated table. A concurrent ordinary commit
can also capture the queue between its write and auto-commit: it does not
honor the queue lock. An index-lock or hook failure leaves already-mutated
queue state without the promised commit, and an unsafe rollback could erase
a later writer. Git's index lock is not a lock over the whole operation.
Fix: default auto-commit off until Q3 is settled. If authorized, require clean queue paths relative to the intended baseline, no conflicting staged versions, no merge/rebase, and the authorized branch; serialize all relevant commit writers through an agreed repository protocol. Freeze and validate the exact prospective file contents, check HEAD/revision before committing, and define retry/recovery without reset/checkout of others' changes. A queue lock alone cannot enforce cooperation by arbitrary host git commands.
The statement that a validator means invariant 8 holds is false: schema validation is not the applicable suites. Identify the required checks and bind their receipt to the actual prospective content/code version. Absence of a shell suite wrapper does not exempt existing Node package tests.
Acceptance: pre-dirty header, partially staged queue file, unrelated staged files, competing ordinary commit, hook failure, and changed HEAD. Verify both the committed tree and preservation of the other author's index and working changes. No push authority follows from any local commit.
R3 — High: single-writer checks validate appearance, not provenance
Where: 5.9, 3.A, migration and render --check.
Scenario: hand-edit a row from queued to done, run render, and both schema
validation and render --check pass. A Markdown hand edit is silently erased
by the next successful CLI mutation, before the suite ever checks it.
A merge can produce valid JSON with illegal transitions or removed rows;
two clones can independently allocate the same new id. A fresh clone has no
local lock installation or uncommitted audit trail and cannot reconstruct
the claimed writer from a matching table.
Fix: check for unexpected drift before mutating or rendering, refuse ambiguous markers and refuse to overwrite unexplained manual changes. Define a reviewed genesis/migration and validate each subsequent operation against its parent revision, including immutable IDs, deletions and authority rules. Use a committed operation history or another explicitly authorized receipt scheme, with merge reconciliation and fresh-clone validation. Keep a high water mark/tombstones for retired IDs, including removed row 7. If multiple clone writers are out of scope, refuse their integration until reconciled; a per-checkout lock does not solve that problem.
Be honest about the boundary: code running as the same host user cannot prevent that user from rewriting both data and receipts. Without an enforced integration boundary, this is a cooperative writer with drift detection, not tamper-proof authority. Hooks alone would not fix fresh clones and are outside Piece A's brief anyway.
Acceptance: valid-but-illegal hand-edited JSON with regenerated Markdown, deleted highest id, duplicate-id branch additions, conflicting merge, fresh clone, missing/duplicate render markers, and preexisting Markdown drift.
R4 — High: identity attribution is still being used as authorization
Where: Q5, 3.A authority tests, 5.3 and 5.8; brief iron-clad point 5.
Scenario: a seat runs --by jason or sets MOSAIC_AGENT_NAME=sage and
moves a protected row. If assign lacks the move permission check, it can
instead assign another seat's row to itself and then move it. Even without
spoofing, the proposed rule only restricts clearing required, while the
brief says only Jason moves required rows. These are different rules.
Fix: separate claimed attribution from verified authority. Q5 can map the new lead's queue role, but does not turn an environment variable into proof or grant managed-worker policy authority. Specify the permitted actor for every mutating verb/field, including assign, dependencies, required, gate ownership, notes and terminal rows. Either retain an explicit reviewed human gate for protected mutations or have Jason accept the limited cooperative trust model; weekly reporting is detection after the fact, not enforcement. Reconcile the exact required-row exception before building.
Acceptance: forged identity, reassignment bypass, dependency/gate edits by an ordinary seat, edits to done/parked rows, and every operation on a required row. No roles/*.json expansion is implied by this review.
R5 — High: Gate G can appear to pass with extra human help or no work
Where: Gate G setup and pass evidence; ledger.mjs messageKind (line 78)
and readSessions (line 92); agent-host-dev.sh context and fresh handling.
Scenario: Jason supplies a correction through the board or a lead relays
it. The ledger counts it as board/agent, leaving exactly one human message.
--fresh creates a new conversation but preserves old session files and
loads shared SOUL/USER/CONTEXT/skills; a task-specific hint there can make
the demo pass without discovering the assignment. Another seat can move
the row using the test seat's claimed identity. Reading the brief and
moving the row alone is also not evidence of actually starting its work.
Observed read-only probe: messageKind returns human for the initial
instruction, board for [dragon-lin:control-board -> dragon-lin:rocko]
followed by a human correction, and agent for a similarly prefixed Sage
relay. The new [from: sage (...) -> to: rocko (...) class=actionable]
header is classified human. Table 2 aggregates a seat's sessions over the
date range, not an identified Gate G session. It is unsuitable as the
sole acceptance counter in either direction.
Fix: pin the exact fresh session id, initial context/brief/queue revisions, test interval and command output. Audit all input channels in that interval; distinguish the initial instruction from further human input regardless of routing. Preserve generic governance context but exclude task-specific pre-coaching. Require a concrete first authorized work action tied to the returned row and that session. Jason's observed pass/fail remains authoritative; do not infer it from the row or git author. Allow normal prerequisite reads instead of requiring the literal first filesystem read to be the brief. A commit is not necessary proof of G and must not bypass Q3 or suite requirements.
Acceptance: negative controls with a board correction, relayed hint, resumed session and row moved by a helper must fail. Use unit tests with competing eligible rows to test selection: the single-row demo cannot.
R6 — High: dependencies and next do not prevent unauthorized starts
Where: Q8, section 1, 3.A and 5.11.
Scenario: next skips a dependent row, but the seat directly runs
move ID in-progress and bypasses the filter. Two sessions of the same
seat both read next, then both start the same work; a state of in-progress
does not distinguish resume by the owner session from a second claimant.
An in-review row returned to its author may cause premature continuation;
a reviewer can repeatedly receive a row it already reviewed if its receipt
is not represented.
Fix: enforce prerequisite and authority predicates at transition time under the lock, not only in selection. Define atomic claiming or an explicit single-active-session assumption with detectable conflict. Return the action role (implement, resume, review, wait) and track review round/receipt so review dispatch has a stopping condition.
There are unresolved semantics: Q8's done-only dependencies cannot express
the brief's row-6-blocked start exception. State-priority ordering differs
from the existing first-eligible-row cadence. after alone cannot express
every arbitrary owner priority, especially between active rows. Decide these
once and update the brief/cadence together. Test that required work cannot
be subordinated to ordinary queued work through dependencies.
R7 — High: posting then recording cannot promise exactly one review request
Where: 2.D, 3.D, 5.14; scripts/gitea-api.sh.
Scenario: Gitea accepts a comment but the response is lost, or the CLI dies before storing its id. The row remains unchanged and retry posts a second request. Conversely, a posted request survives a failed local commit. A repository lock cannot roll back the remote side effect. The existing helper supplies an ordinary POST, not an exactly-once transaction.
Fix: define a durable operation id containing row, revision, review round and candidate digest; persist intent before sending. Include the id in the request, reconcile uncertain outcomes by lookup/readback, and refuse blind retries if delivery cannot be determined. Specify outbox/recovery state and bounded lock behavior. An authorized token path also needs an expected account check; checking that the path differs from the default is not proof of the intended poster. Do not read or provision tokens before Q4.
Acceptance: accepted POST with dropped response, crash before receipt, duplicate retry, concurrent moves, stale candidate and local commit failure. Fake success/failure alone does not test these boundaries.
R8 — High: missing from the open list does not establish closed
Where: Q7 and 2.E.
Scenario: a row references a typo/nonexistent issue, a deleted issue, a PR number or an inaccessible item. It is absent from a complete open-issues response and the proposed algorithm declares it closed. Even a short page only shows that that response is short; it does not prove each referenced number is a known issue. A done row can therefore pass with no closure evidence.
Fix: retain the additional open-list call and full-page refusal, which
are useful, but represent absent as unknown unless there is positive issue
existence/type/state evidence. Fetch/cache positively identified referenced
issues under a revised call budget, or report incomplete evidence instead
of success. Bound cache age and detect changed scope. Specify --no-issues
as unknown queue issue checks, not zero violations.
Acceptance: nonexistent/deleted/PR references, denied access, malformed responses, a full page and no-issues mode. This is a requested correction to decided Q7, not an instruction to silently implement another fetch policy.
R9 — High: the liveness exemption can turn missing evidence into green
Where: Q6, 2.E and 3.E; scan.mjs lines 272–307.
Scenario: remove a native seat's registration and it gets the same
non-violation label as a T3 seat. loadRegistrations returns records plus
errors; it does not check liveness and loads multiple layouts. Matching a
seat name alone can select a fleet/other-workspace record. A dead or reused
PID is not evidence of the intended active session. Corrupt records must
not become the missing-registration exemption.
Fix: preserve Q6's prohibition on T3 probing, but distinguish explicitly unsupported runtime, missing expected registration, invalid record, stale record and live matched record. Match repository/layout/session identity. Propagate scanner errors. Show coverage/unknowns alongside violations and do not call E fully passed on zero violations when liveness was untested. Record the owner's explicit acceptance of a temporary reduced gate.
Acceptance: missing native registration, corrupt record, duplicate seat across layouts, wrong workspace, dead PID, unknown PID and an exempt T3 seat.
R10 — Medium: history needed for E is absent from the schema
Where: 5.3, 5.8 and 5.12; brief E's age and same-day remediation rules.
Scenario: queue note refreshes updatedAt every week, making a
long-neglected required row appear young. A row is changed twice between
ledger runs and updatedBy remembers only the last writer, so E cannot list
every protected change as 5.8 promises. Removing a row and having no delete
verb does not prove an id was never reused. A single latest timestamp also
cannot prove each observed violation was remedied within the day.
Fix: define stable createdAt/requiredSince, operation history and audit retention, with explicit time zone and first-observed violation evidence. Do not depend on optional Q3 commits for mandatory checks. If a journal is chosen, make corrections new entries and place it in a dedicated directory; do not rewrite BUILD-LOG, SESSIONS or existing append-only runtime logs. Terminal required rows should not age forever as outstanding work.
R11 — High: some section-5 fixes change owner requirements, not details
Where: 5.6, 5.7, 5.10, 5.13 and 5.14; brief A/C/D and iron-clad point 3.
Scenario: the builder follows “settle in A's round” and makes brief optional, reopens parked rows, allows completion without Jason, or changes the issue/gate semantics. Those are externally visible requirements changes. The brief explicitly demands an existing brief for every row and says parked/done rows never change. A queued row can have an existing planning brief not yet accepted; a missing brief is not technically necessary to make queued reachable.
Fix: put each unresolved behavioral change in a disposition table:
original rule, proposed rule, approving authority, acceptance change. Keep
already-decided changes separate. Do not conflate source review approval
with owner approval of altered gates. Decide the blanket any→blocked
versus terminal immutability conflict and block→block semantics too.
The multi-row issue problem is real, but “open iff at least one row not done” is only sound if rows exhaust the issue's scope. Row 1/22 explicitly separates broader MVP acceptance from one targeted correction; row 6 has two issue phases. Model which row gates issue closure, or report mismatch for disposition without forcing automatic closure. A single reviewer field also loses row 6's independent Darkwing/Dewey reviews; use a review policy that can preserve both if that row is migrated intact.
R12 — Medium: immutable review candidate and receipt lifecycle are missing
Where: 5.14 and Piece D pass tests.
Scenario: a request posts /tmp/candidate.sha256; the file is replaced
or disappears, so a later reviewer cannot identify what was approved. This
takeover already found the historical #1511/#1512 /tmp snapshots absent.
With issues: [int], “the piece's issue” also no longer selects an unambiguous
posting target. Posting a request alone does not establish a full review
round or preserve its verdict when the row loops through review again.
Fix: resolve and store the candidate content digest and durable commit or artifact reference, not just a path verbatim. Name the request issue, review round, required reviewers and receipt links; preserve prior rounds. Test request→changes requested→new candidate→approval against exact pins. Keep the brief's ban on new review files, but evaluate the stronger proposed ban separately: no-files-anywhere is not a substitute for durable receipts.
R13 — Medium: brief validation and migration are too shallow
Where: 3.A migration/C and 5.12; current QUEUE rows and parked table.
Scenario: a directory or symlink outside the repository “exists” as a brief, or an untracked local brief makes the author's tests pass but is absent in a fresh clone. A nonexistent section passes even though the fresh seat is instructed to open only the named brief section. Migration shortens a state cell and drops an acceptance/publication boundary that is not actually duplicated in CURRENT or an issue. A valid golden Markdown file proves none of those semantics.
Fix: require a readable repository-contained regular brief, a defined symlink policy, and tracked prospective content for publishable rows. Either validate section anchors or disallow them for actionable rows until resolved. Produce a reviewed row-by-row migration map preserving owners, review lanes, gates, required status, dependencies, historical IDs and each unresolved boundary. Map parked entries with no briefs explicitly rather than inventing authority from their free text. Preserve history through exact durable links or a reviewed archive; do not assume it was copied elsewhere.
R14 — Medium: package reuse and acceptance wiring need concrete decisions
Where: 2.E, 3.B, section 1; package.json and existing module layout.
Scenario: ledger imports @mosaic/queue but adding a sibling package.json
does not install a workspace package. Root package.json declares no workspaces;
a read-only import.meta.resolve('@mosaic/seat') probe returns
ERR_MODULE_NOT_FOUND today. Moving the registration loader into seat also
touches a package outside E's listed files and requires its regression tests.
Fix: use the existing relative-module convention or explicitly plan
package linking and its installation/lockfile consequences; do not leave
it implicit. Reuse the loader without silently expanding ownership, and
specify liveness/error handling separately (R9). Name actual package test
commands and affected suites. C also needs the promised two owner-accepted
briefs, not just missing-path tests. D needs a full review round and E the
dated ledger receipt/remediation evidence. B's --check verifies inputs
exist, not that a running or fresh session received the new cadence; inspect
the launch context and then perform G. Clarify the brief's “A to E in order”
against the proposed C-with-A and D/E parallel delivery before dispatch.
What is fine
- Repository-local
scripts/mosaicdispatch, keeping seat argv/env behavior, and tests for that behavior are appropriate. Q1 correctly remains an owner wording choice; no global CLI or fleet edits are needed. - Queue data under docs/plans and code under packages/queue respect bootstrap-only root. An edit to the existing root AGENTS.md is permitted; a new root lock/journal/config file would not be. The plan proposes no contract, secret, run-record or managed-role changes.
- Required as a flag, an explicit review-to-work edge, strict unknown-field validation, dependency cycle checks, deterministic rendering and a non-mutating render check are good building blocks. They need the transaction/history/authority rules above to deliver the claimed result.
- Missing identity refusal, Q4 held for explicit credentials, no automatic push, no T3-backed product feature, and preservation of the old table log are sound boundaries. Historical log corrections must remain additions.
- Combining migration with the introduction of the authoritative data file avoids a deliberate two-master interval. A reviewed migration map still matters. Reusing one validator is correct once module resolution is explicit.
Recommended disposition
Revise the write/recovery/commit protocol (R1–R3), settle protected authority and requirement changes (R4/R11), strengthen Gate G evidence (R5), and specify transition claims and review delivery recovery (R6/R7) before implementation. Correct E's positive-evidence rules and audit model (R8–R10); complete the candidate, migration and packaging details (R12–R14). These are bounded plan changes, not authorization to build a new policy service or alter live state.
Validation for this review was source inspection and read-only in-process probes of the existing message classifier and module resolution. No queue implementation or fault-injection suite was run because none exists yet. This report records the session outcome in place of editing SESSIONS.md, as this assignment explicitly permits writes only to the report.